Requirements and limitations
Before enabling container live migration in your cluster, ensure your infrastructure meets the specific requirements for successful pod migration between nodes. This capability requires particular node configurations, compatible instance types, and specific Kubernetes versions to function reliably.
System requirements
Cloud platform support
For a cross-cloud comparison of node images, TCP preservation paths, ARM64 support, and PVC zone rules, see Cloud providers.
Kubernetes version
Container live migration requires Kubernetes 1.30 or later. This minimum version ensures compatibility with the container runtime enhancements and custom resource definitions that enable live migration functionality.
Earlier Kubernetes versions lack the necessary API stability and container runtime features required for reliable checkpoint and restore operations. In Kubernetes 1.30 (2024), the checkpoint/restore support graduated to the Beta phase.
Node infrastructure requirements
Cast AI management: Both source and destination nodes must be managed by Cast AI. Live migration cannot occur between Cast AI-managed nodes and nodes managed by other provisioners or cloud provider native node groups.
Node image requirements: Supported node images vary by cloud provider. Node images are configured in the cluster's node configuration.
| Cloud | Supported node images | Notes |
|---|---|---|
| EKS | Amazon Linux 2023 (standard and NVIDIA/Neuron) | Bottlerocket and Amazon Linux 2 are not supported |
| GKE | Container-Optimized OS (COS), Ubuntu | COS recommended for TC (kernel 6.6+) |
| AKS | Ubuntu 22.04, Ubuntu 24.04, Azure Linux 3.0 | Ubuntu 24.04 or Azure Linux 3.0 needed for TC |
For kernel requirements and TCP preservation path availability per image, see the cloud provider guide.
Container runtime: Nodes must use containerd v2+ as the container runtime engine. Other runtimes are not supported due to specific integration requirements for checkpoint and restore functionality.
Network topology: Networking requirements depend on the cloud provider and the TCP preservation path in use. The most constrained path is the AWS VPC CNI path on EKS, which requires source and destination nodes to be in the same subnet. The Traffic Control (TC) path has no subnet or AZ constraint on any cloud. For full details, see Cloud providers.
When using the AWS VPC CNI path on EKS, the following networking constraints apply:
Standard configuration (ENABLE_PREFIX_DELEGATION=false):
- Single subnet per Node and Pod
- No cross-subnet migrations (validated by checking node
topology.cast.ai/subnet-idlabel)
Dedicated subnet for Pods:
- ENIConfig CRD must be configured in the cluster
- ENIConfig is per availability zone only
- No cross-AZ migration (validated by
topology.kubernetes.io/zone) - Source node must have
live.cast.ai/custom-networklabel set
Subnet discovery must not be configured:
kubernetes.io/role/cnitag must not be configured on any subnets- Subnet IP capacity: Migrations occur within the same subnet. If a subnet's available IP addresses are exhausted, migrations may fail or become stuck waiting for IP allocation. Monitor subnet IP utilization in clusters with frequent migration activity, particularly in smaller subnets or during large-scale rebalancing operations, where many migrations would be triggered at once.
The TC path and the Calico-as-full-CNI path have no subnet or AZ constraint. See AWS EKS for the full comparison of TCP preservation paths.
Instance type and CPU compatibility
CPU architecture consistency: CPU architecture consistency is critical for successful live migration. Node templates used for live migration must be configured for a single processor architecture—either AMD64 or ARM64, but not both (any). Attempting to migrate workloads between nodes with different architectures (e.g., from AMD64 to ARM64) will fail. Cloud providers may also provision different CPU models within the same instance type, creating compatibility issues with migration.
Node template configuration requirements: Instance families in your node templates must belong to the same generation set for migration compatibility:
- Compatible sets: m3 with c3, or c5 with r5, m5, etc.
- Incompatible: c3 with c5, or mixing different generation families
Configure your node templates to include only instance types from the same generation family via instance constraints. Migration between nodes with different CPU architectures or generation sets will fail.
Architecture support: Container live migration supports both AMD64 and ARM64 architectures, but node templates must be configured for a single architecture. You cannot mix architectures within a live migration-enabled node template.
When configuring your node template:
- Set processor architecture to either AMD64 or ARM64
- Do not select "Any"
All nodes provisioned from the template will use the selected architecture
There are also functional differences between architectures:
- AMD64: Full incremental memory transfer
- ARM64: Non-iterative memory transfer
The architectural difference on ARM64 affects the approach to memory dumping during the checkpoint process. This limitation cannot be overcome.
Container Network Interface (CNI)
TCP preservation during migration is handled differently depending on the cloud provider and CNI in use. There are three TCP preservation paths:
| Path | Pod IP after migration | Available on | Kernel requirement |
|---|---|---|---|
| VPC CNI (forked AWS VPC CNI) | Same IP (preserved) | EKS only | None |
| Traffic Control (TC) | New IP (peers rewrite packets) | EKS, GKE, AKS | Linux kernel 6.6+ |
| Calico as full CNI (VXLAN overlay) | Same IP (annotation pinned) | EKS, AKS | None |
On EKS, a forked version of the AWS VPC CNI is automatically installed on Cast AI-managed, live-enabled nodes when the VPC CNI path is active. It enables IP address preservation across nodes and TCP session continuity during migration. This fork maintains compatibility with standard AWS VPC networking while adding the necessary features for live migration.
On GKE and AKS (and optionally on EKS), the Traffic Control (TC) path uses a privileged DaemonSet that rewrites packets on peer nodes via eBPF so established TCP connections survive migration even though the pod gets a new IP. TC requires Linux kernel 6.6 or later.
On EKS and AKS, Calico as the full CNI with VXLAN overlay preserves the pod IP via annotation pinning, without the subnet constraint of the VPC CNI path and without the kernel requirement of TC.
For per-cloud setup instructions and CNI configuration details, see:
For detailed configuration reference for each TCP preservation path, see:
Supported workload types
Container live migration supports the following Kubernetes workload types:
| Workload type | Support status | Notes |
|---|---|---|
| StatefulSets | Supported | |
| Deployments | Supported | |
| Bare pods | Supported | |
| Jobs | Supported | |
| CronJobs | Supported | |
| Custom controllers | Limited support | Compatibility assessment required. Contact Cast AI. |
| DaemonSets | Not supported | Cannot be migrated by design. |
Multi-container support: Pods with multiple containers are fully supported. All containers within a pod are migrated together, except Init containers: Init containers are skipped during migration by default and will not be rerun on the restored pod.
Storage compatibility
Supported storage types
| Storage type | Support scope |
|---|---|
| Persistent Volume Claims (PVCs) | Zonal PVs: same zone required; Regional PVs: any zone in PV's allowed list (all clouds) |
| Network File System (NFS) | Cross-node |
| EmptyDir volumes | Supported |
| ConfigMap volumes | Supported |
| Secret volumes | Supported |
| Host path volumes | Experimental support only. Contact Cast AI for specific use cases. |
Host path volumes: Host path volumes present unique challenges because they access node-local storage. Limited support exists for some use cases.
PVC reattachment timing: When migrating workloads with Persistent Volume Claims, the PVC must be detached from the source node and reattached to the destination node. On EKS, the average reattachment duration is approximately 9 seconds, though this varies by volume type, size, and current cloud API latency. Factor this additional time into migration planning for PVC-backed workloads and test in your environment to establish baseline expectations.
Current limitations
Hardware and performance constraints
GPU workloads: Not currently supported due to the complexity of checkpointing and restoring GPU memory and compute.
For maximizing the efficiency of GPU workloads on a given node, Cast AI offers sophisticated alternative approaches in the forms of MIG and GPU time-slicing.
Memory transfer performance: Migration time increases with workload memory usage. Large memory footprints may experience naturally longer migration times.
Spot Instance compatibility: When used with Spot Instances, container live migration cannot guarantee successful migration before the CSP terminates the node. Spot interruption notices provide limited time windows that may be insufficient for completing the migration.
Before enabling container live migration on Node templates configured for Spot Instances, contact Cast AI support to assess your workloads and determine the appropriate configuration.
Problematic workloads
Certain workload types may have an increased rate of failure during migration, even when labeled as eligible. The live controller cannot detect all incompatible configurations, so test these workload types thoroughly before production use.
| Workload type | Issue | Mitigation |
|---|---|---|
Applications using io_uring | The io_uring async I/O interface state cannot be checkpointed | Disable at kernel level via node configuration init script (see below) |
| Applications using MPTCP | Multipath TCP connection state cannot be checkpointed | Disable MPTCP in your application (see below) |
| eBPF programs | eBPF program state is not checkpointed | Not currently supported for live migration |
| Privileged containers | Multiple checkpoint incompatibilities | Avoid privileged containers for CLM-enabled workloads |
Disabling io_uring for CLM compatibility:
Add the following to your node configuration init script to disable io_uring at the kernel level:
sudo sysctl -w kernel.io_uring_disabled=2This setting prevents applications from using the io_uring interface, which cannot be checkpointed by CRIU.
Disabling MPTCP for CLM compatibility:
For Go applications, disable MPTCP by setting the following environment variable:
GODEBUG=multipathtcp=0Add this to your container's environment variables in the pod spec. Other languages and frameworks may have similar configuration options to disable MPTCP.
Eligibility vs. successThe
live.cast.ai/migration-enabled=truelabel indicates a workload meets basic eligibility criteria. It does not guarantee successful migration. Always validate specific workload types in a test environment before production rollout.
Automatic workload assessment
Cast AI automatically evaluates workloads for migration eligibility:
Automatic labeling: The live controller scans your cluster and applies migration-eligible labels (live.cast.ai/migration-enabled=true) to workloads that meet all requirements.
Continuous assessment: As workloads change (storage additions, security context modifications, etc.), the controller updates eligibility labels accordingly.
Checking workload eligibility
The simplest way to check if workloads are eligible for live migration is to look for the migration label:
# Check all pods with live migration enabled
kubectl get pods -A -l live.cast.ai/migration-enabled=trueDo note that this will only work once container live migration is enabled and the live controller is operating in the cluster.
See also
Per-cloud setup guides for EKS, GKE, and AKS — node images, TCP preservation, ARM64, and PVC support.
How Kubernetes probes and lifecycle hooks behave during migration.
Reference for all CLM labels, annotations, and migration events.
Updated 29 days ago
