For AI agents: visit https://docs.cast.ai/llms.txt for an index of all pages formatted in Markdown and endpoints in OpenAPI.
Jump to Content
Cast AI
GuidesAPI ReferenceRelease Notes
Log InCast AI
Guides
Log In
GuidesAPI ReferenceRelease Notes
  • Getting started
    • About the read-only agent
  • Connecting your cluster
    • Connect using the castctl CLI
    • Connect using the Cast AI console
    • Cloud Connect
    • GCP Private Service Connect
    • AWS PrivateLink
  • Enable automation
    • Autoscaler preparation checklist
    • Troubleshooting cluster onboarding
  • Platform permissions & data privacy
    • Kubernetes permissions
    • Cloud permissions
      • AKS workload identity impersonation
      • GKE service account impersonation
    • Data collection and storage
    • Communication requirements
  • Cast AI Anywhere
    • Overview
    • Getting started
  • API access
  • Component management
    • Hosted components
      • Cluster controller
      • Spot Handler
    • Helm charts
    • Terraform provider
      • GKE via GitOps
      • EKS via GitOps
      • AKS via GitOps
      • Terraform troubleshooting
    • Component control
    • Cast AI Operator
    • Open source components
      • Audit log exporter
      • egressd (deprecated)
      • GPU metrics exporter (deprecated)
    • Troubleshooting Cast AI components
  • Disconnect your cluster
  • Overview
  • Getting started
  • Feature reference
  • Kentroller
  • Scheduled rebalancing for Karpenter clusters
  • Continuous rebalancing
  • Overview
  • Getting started
  • Runbooks
    • Fix container image vulnerabilities
    • Synchronize Workload Autoscaler recommendations
  • Autoscaling
    • Node templates
    • Node configuration
    • Spot Instances
      • Spot interruption prediction API
    • GPU Instances
    • Storage-optimized nodes
    • TPU Instances (GKE)
    • AWS Neuron Instances (EKS)
    • GPU sharing
      • Time-slicing
      • Multi-Instance GPU (MIG)
      • Multi-Process Service (MPS)
      • Fractional GPUs (AWS)
    • Dynamic Resource Allocation (DRA)
    • Pod placement
    • Pod Pinner
    • Subnets
    • Network bandwidth
    • Commitments
      • Enterprise commitments
      • AWS capacity reservations
    • Autoscaler settings
    • Autoscaler Node Labels and Taints
    • Managing DaemonSets with Cast AI
    • Troubleshooting node autoscaling
  • Downscaling
    • Evictor
    • Evictor vs. Rebalancer
  • Rebalancing
    • Workload preparation
    • Scheduled rebalancing
    • Paused drain configuration
  • Cluster hibernation
    • Cluster hibernation (Legacy)
  • Migration from Karpenter
  • Upgrading Kubernetes version
  • Cluster certificate rotation
  • Container Live Migration
    • Concept
      • Overview
      • Probe and lifecycle behavior
    • Reference
      • Requirements and limitations
      • Labels, Annotations, and Events
    • Tutorials
      • Using Container Live Migration with Evictor and Rebalancer
  • Pod mutations
    • Quickstart
    • Overview
    • Tutorials
      • Enable Workload Autoscaler with pod mutations
    • Reference
  • Using ARM nodes with Cast AI
  • Business continuity
  • Watchdog
  • Overview
  • Getting started
  • Custom edge locations
  • Edge configuration
  • Manual edge provisioning
  • Overview
  • Workload Autoscaler configuration
    • Available settings
    • Annotations reference
      • Legacy annotations reference (deprecated)
  • Scaling policies
    • How-to: Create a scaling policy
    • How-to: Manage scaling policies
  • Custom metrics
  • JVM workload optimization
  • In-Place Pod Resizing
  • Node-aware DaemonSet sizing
  • Pod startup recommendations
  • Horizontal Pod Autoscaling
    • Tutorials
      • How-to: Configure HPA on a workload
      • How-to: HPA in scaling policies
      • How-to: Migrate from legacy horizontal scaling to HPA
    • Reference
      • KEDA compatibility
      • Vertical & horizontal workload autoscaling
      • Legacy horizontal scaling (v1) (deprecated)
  • Event log
  • Overview
  • Available savings
  • OpsPilot
  • Dashboard
  • Cluster score
  • Organization-level reports
    • Organizational cluster cost report
    • Organizational allocation groups
    • Idle resources report
  • Cluster-level reports
    • Efficiency
    • Workloads
    • Namespaces
    • Allocation groups
    • Cost comparison
    • Reliability Metrics
      • Reliability Metrics Reference
  • GPU utilization
  • Network cost
  • Storage cost
  • CPU vs. memory cost calculation
  • Metrics
    • Tutorials
      • Integrating Prometheus Metrics with New Relic
  • Introduction
  • How it works
  • Installation
  • Configuration