For AI agents: visit https://docs.cast.ai/llms.txt for an index of all pages formatted in Markdown and endpoints in OpenAPI.
Jump to Content
Guides
API Reference
Release Notes
Log In
Guides
Search
Log In
English
Guides
API Reference
Release Notes
Migration from Karpenter
Introducing Cast AI
Getting started
About the read-only agent
Connecting your cluster
Connect using the castctl CLI
Connect using the Cast AI console
Cloud Connect
GCP Private Service Connect
AWS PrivateLink
Enable automation
Autoscaler preparation checklist
Troubleshooting cluster onboarding
Platform permissions & data privacy
Kubernetes permissions
Cloud permissions
AKS workload identity impersonation
GKE service account impersonation
Data collection and storage
Communication requirements
Cast AI Anywhere
Overview
Getting started
API access
Component management
Hosted components
Cluster controller
Spot Handler
Helm charts
Terraform provider
GKE via GitOps
EKS via GitOps
AKS via GitOps
Terraform troubleshooting
Component control
Cast AI Operator
Open source components
Audit log exporter
egressd (deprecated)
GPU metrics exporter (deprecated)
Troubleshooting Cast AI components
Disconnect your cluster
Karpenter Enterprise suite
Overview
Getting started
Feature reference
Kentroller
Scheduled rebalancing for Karpenter clusters
Continuous rebalancing
Application Performance Automation (APA)
Overview
Getting started
Runbooks
Fix container image vulnerabilities
Synchronize Workload Autoscaler recommendations
Node Autoscaling
Autoscaling
Node templates
Node configuration
Spot Instances
Spot interruption prediction API
GPU Instances
Storage-optimized nodes
TPU Instances (GKE)
AWS Neuron Instances (EKS)
GPU sharing
Time-slicing
Multi-Instance GPU (MIG)
Multi-Process Service (MPS)
Fractional GPUs (AWS)
Dynamic Resource Allocation (DRA)
Pod placement
Pod Pinner
Subnets
Network bandwidth
Commitments
Enterprise commitments
AWS capacity reservations
Autoscaler settings
Autoscaler Node Labels and Taints
Managing DaemonSets with Cast AI
Troubleshooting node autoscaling
Downscaling
Evictor
Evictor vs. Rebalancer
Rebalancing
Workload preparation
Scheduled rebalancing
Paused drain configuration
Cluster hibernation
Cluster hibernation (Legacy)
Migration from Karpenter
Upgrading Kubernetes version
Cluster certificate rotation
Container Live Migration
Concept
Overview
Probe and lifecycle behavior
Reference
Requirements and limitations
Labels, Annotations, and Events
Tutorials
Using Container Live Migration with Evictor and Rebalancer
Pod mutations
Quickstart
Overview
Tutorials
Enable Workload Autoscaler with pod mutations
Reference
Using ARM nodes with Cast AI
Business continuity
Watchdog
OMNI Edge Provisioning
Overview
Getting started
Custom edge locations
Edge configuration
Manual edge provisioning
Workload Autoscaling
Overview
Workload Autoscaler configuration
Available settings
Annotations reference
Legacy annotations reference (deprecated)
Scaling policies
How-to: Create a scaling policy
How-to: Manage scaling policies
Custom metrics
JVM workload optimization
In-Place Pod Resizing
Node-aware DaemonSet sizing
Pod startup recommendations
Horizontal Pod Autoscaling
Tutorials
How-to: Configure HPA on a workload
How-to: HPA in scaling policies
How-to: Migrate from legacy horizontal scaling to HPA
Reference
KEDA compatibility
Vertical & horizontal workload autoscaling
Legacy horizontal scaling (v1) (deprecated)
Event log
Monitoring & Reporting
Overview
Available savings
OpsPilot
Dashboard
Cluster score
Organization-level reports
Organizational cluster cost report
Organizational allocation groups
Idle resources report
Cluster-level reports
Efficiency
Workloads
Namespaces
Allocation groups
Cost comparison
Reliability Metrics
Reliability Metrics Reference
GPU utilization
Network cost
Storage cost
CPU vs. memory cost calculation
Metrics
Tutorials
Integrating Prometheus Metrics with New Relic
Storage Autoscaling
Introduction
How it works
Installation
Configuration