Annotations reference
All Workload Autoscaler settings are available by adding annotations to the workload controller. Annotations are the only way to override vertical scaling settings for individual workloads — the Cast AI console does not support per-workload setting overrides. When the workloads.cast.ai/configuration annotation is detected on a workload, it will be considered as configured by annotations. This allows for flexible configuration, combining annotations and scaling policies.
Changes to the settings via the API/UI are no longer permitted for workloads with annotations. The default or scaling policy value is used when a workload does not have an annotation for a specific setting.
Annotation values take precedence over what is defined in a scaling policy. This means that if a scaling policy is defined in the workload configuration under annotations, all of the individual configuration options defined under the annotation will override the respective policy values. Those that are not defined under the annotation will use system defaults or what is defined in the scaling policy.
Example
apiVersion: apps/v1
kind: Deployment
metadata:
name: my-app
labels:
app: my-app
annotations:
workloads.cast.ai/configuration: |
scalingPolicyName: custom
vertical:
optimization: on
applyType: immediate
anomalyDetection:
cpuPressure:
cpuStallThresholdPercentage: 30.0
minPressuredPodPercentage: 20.0
excludedContainers:
- istio-proxy
antiAffinity:
considerAntiAffinity: false
startup:
period: 5m
confidence:
threshold: 0.5
cpu:
target: p81
lookBackPeriod: 25h
min: 1000m
max: 2500m
applyThresholdStrategy:
type: defaultAdaptive
overhead: 0.15
limit:
type: multiplier
multiplier: 2.0
memory:
target: max
lookBackPeriod: 30h
min: 2Gi
max: 10Gi
applyThresholdStrategy:
type: defaultAdaptive
overhead: 0.35
limit:
type: noLimit
downscaling:
applyType: immediate
memoryEvent:
applyType: immediate
containers:
{container_name}:
cpu:
min: 10m
max: 1000m
memory:
min: 10Mi
max: 2048Mi
jvm:
enabled: true
autoInstrument: true
hpaConverters:
- type: AverageValueFromOriginalRequests
rolloutBehavior:
type: NoDisruption
preferOneByOne: true
delaySeconds: 120
horizontal:
optimization: on
minReplicas: 5
maxReplicas: 10
scaleDown:
stabilizationWindow: 5m
containersGrouping:
- key: name
operator: contains
values: ["data-processor"]
into: data-processorLabels
In addition to the workloads.cast.ai/configuration annotation, Workload Autoscaler recognizes the following label on workload controllers and pods:
| Label | Values | Where to apply | Description |
|---|---|---|---|
workload-autoscaler.cast.ai/enabled | "true" | Controller metadata.labels and spec.template.metadata.labels | Required when whitelisting mode is enabled. Only workloads with this label receive applied recommendations. The label must be present on both the controller (Deployment, StatefulSet, DaemonSet) and the pod template. |
Configuration Structure
Below is a configuration structure reference for setting up a workload to be controlled by annotations.
Note
workloads.cast.ai/configurationhas to be a valid YAML string. In cases where the annotation contains an invalid YAML string, the entire configuration will be ignored.
scalingPolicyName
If not set, the system assigns the workload to system default policies based on workload type: resiliency for StatefulSets and balanced for all other workloads.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
scalingPolicyName | string | No | System default (see below) | Specifies the scaling policy name to use. When set, this annotation allows the workload to be managed by both annotations and the specified scaling policy. |
scalingPolicyName: custom-policyFallback for invalid policy names
If the scalingPolicyName value does not match any existing scaling policy, the workload falls back to system default policies:
- StatefulSets:
resiliency - All other workloads:
balanced
When this fallback occurs, the Cast AI Console displays a warning status on the workload in the Optimization page, such as: Could not find scaling policy by name '<policy-name>'. To resolve this, update the annotation to reference a valid scaling policy name.
vertical
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
vertical | object | No | - | Vertical scaling configuration. |
vertical:
optimization: on
applyType: immediate
antiAffinity:
considerAntiAffinity: false
startup:
period: 5m
confidence:
threshold: 0.5
excludedContainers:
- istio-proxyvertical.optimization
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
optimization | string |
| Enable vertical scaling ("on"/"off"). If using the |
vertical:
optimization: onvertical.applyType
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
applyType | string | No | "immediate" | Allows configuring the autoscaler operating mode to apply the recommendations. Use immediate to apply recommendations as soon as the thresholds are passed.
|
vertical:
applyType: immediatevertical.antiAffinity
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
antiAffinity | object | No | - | Configuration for handling pod anti-affinity scheduling constraints. |
vertical:
antiAffinity:
considerAntiAffinity: falsevertical.antiAffinity.considerAntiAffinity
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
considerAntiAffinity | boolean |
| false | When true, workload autoscaler will respect pod anti-affinity rules on hostname or host port and, as a result, issue recommendations for these pods in a deferred manner.When false (default), recommendations for pods containing one of the constraints above will be applied immediately instead.
|
vertical:
antiAffinity:
considerAntiAffinity: falsevertical.startup
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
startup | object | No | Configuration for how Workload Autoscaler handles resource usage observed during the startup period. Supports three modes: include startup metrics in calculations (default), apply original CPU requests during startup then switch to the optimized recommendation (startup recommendations), or exclude startup metrics from calculations (legacy ignore mode). See Startup metrics. |
vertical:
startup:
period: 5mvertical.startup.period
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
period | duration |
| "2m" | Duration of the startup period. Applies to all startup metrics modes: determines how long to apply original CPU requests (startup recommendations mode), or how long to exclude metrics from calculations (ignore mode). Valid values range from 2m to 60m. Set to 0s to disable. Required when vertical.startup is set.
|
vertical:
startup:
period: 5m # Example: ignore first 5 minutes of metricsvertical:
startup:
period: 0s # Disable startup metrics ignore periodvertical.startup.twoPhaseRecommendations
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
twoPhaseRecommendations | object | No | Configure two-phase startup recommendations. When enabled, newly created pods receive the workload's original resource requests during the startup period, then transition to the optimized recommendation via in-place resize once the startup period ends. |
vertical:
startup:
period: 30s
twoPhaseRecommendations:
enabled: truevertical.startup.twoPhaseRecommendations.enabled
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
enabled | boolean |
| Set to true to enable two-phase startup recommendations for this workload.
|
vertical:
startup:
period: 30s
twoPhaseRecommendations:
enabled: truevertical.startup.twoPhaseRecommendations.requestsOnStartup
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
requestsOnStartup | object | No | Override the resource requests used during the startup phase. Each field (cpu, memory) is optional and falls back to the workload's original pod request for that resource when omitted. Set this when the original requests are unknown or should be different from the workload spec. |
vertical:
startup:
period: 30s
twoPhaseRecommendations:
enabled: true
requestsOnStartup:
cpu: 500m
memory: 256Mivertical.startup.twoPhaseRecommendations.requestsOnStartup.cpu
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
cpu | string | No | Original pod request | CPU request to use during the startup phase. Uses Kubernetes CPU notation (for example, 500m, 2). When omitted, the workload's original CPU request is used. |
vertical:
startup:
period: 30s
twoPhaseRecommendations:
enabled: true
requestsOnStartup:
cpu: 500m # memory falls back to original pod requestvertical.startup.twoPhaseRecommendations.requestsOnStartup.memory
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
memory | string | No | Original pod request | Memory request to use during the startup phase. Uses Kubernetes memory notation (for example, 256Mi, 1Gi). When omitted, the workload's original memory request is used. |
vertical:
startup:
period: 30s
twoPhaseRecommendations:
enabled: true
requestsOnStartup:
memory: 256Mi # cpu falls back to original pod requestvertical.confidence
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
confidence | object | No | - | Configuration for recommendation confidence thresholds. |
vertical:
confidence:
threshold: 0.5
required: falsevertical.confidence.required
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
required | bool |
| When set to:
|
vertical:
confidence:
required: truevertical.confidence.threshold
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
threshold | float |
| 0.9 | Minimum confidence score required to apply recommendations (0.0-1.0). Higher values require more data points for recommendations.
|
vertical:
confidence:
threshold: 0.5vertical.anomalyDetection
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
anomalyDetection | object | No | - | Configuration for anomaly detection features, like CPU pressure-based stall detection. |
vertical:
anomalyDetection:
cpuPressure:
cpuStallThresholdPercentage: 10.0
minPressuredPodPercentage: 50.0vertical.anomalyDetection.cpuPressure
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
cpuPressure | object | No | - | Configuration for CPU pressure-based stall detection using PSI metrics. |
vertical:
anomalyDetection:
cpuPressure:
cpuStallThresholdPercentage: 10.0
minPressuredPodPercentage: 50.0vertical.anomalyDetection.cpuPressure.cpuStallThresholdPercentage
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
cpuStallThresholdPercentage | float | No | Varies by policy | The CPU stall percentage that triggers resource increases. When workloads can't get CPU time and stall exceeds this threshold, they'll be scaled up to reduce contention. Measured over a 5-minute window. Value range: 1.0–100.0. See Stall detection. |
vertical:
anomalyDetection:
cpuPressure:
cpuStallThresholdPercentage: 30.0 # 30% stall thresholdvertical.anomalyDetection.cpuPressure.minPressuredPodPercentage
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
minPressuredPodPercentage | float | No | Varies by policy | The percentage of pods that must exceed the stall threshold before scaling triggers. This prevents scaling the entire workload when only a single pod or small subset is experiencing pressure. Value range: 1.0–100.0. See Stall detection. |
vertical:
anomalyDetection:
cpuPressure:
minPressuredPodPercentage: 20.0 # Scale when 20% of pods are stallingvertical.resourceStrategy
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
resourceStrategy | object | No | - | Sizes the workload using a strategy other than the standard resource recommendation. At pod admission, Workload Autoscaler resolves it per container and writes the resulting requests and limits to the pod, overriding the standard recommendation for that resource. You can also set it per container via vertical.containers.{container_name}.resourceStrategy, which overrides the workload-level setting. See Node-aware DaemonSet sizing. |
vertical:
resourceStrategy:
type: nodeAllocatablePercentage
nodeAllocatablePercentage:
cpuPercent: 5
memoryPercent: 2At this (workload) level, the strategy applies to every container in the pod that already has CPU or memory requests in its spec. To size a container that has no requests defined, or to use different percentages per container, set it at the container level.
vertical.resourceStrategy.type
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type | string | Yes | Which sizing strategy to apply to the workload. Supported values:
|
vertical:
resourceStrategy:
type: nodeAllocatablePercentagevertical.resourceStrategy.nodeAllocatablePercentage
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
nodeAllocatablePercentage | object | Yes when type: nodeAllocatablePercentage | Percentages used by the Cannot be combined with the |
vertical:
resourceStrategy:
type: nodeAllocatablePercentage
nodeAllocatablePercentage:
cpuPercent: 5
memoryPercent: 2vertical.resourceStrategy.nodeAllocatablePercentage.cpuPercent
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
cpuPercent | float | One of cpuPercent or memoryPercent is required | Percentage of the node's allocatable CPU to request, resolved per container (each container gets its own request of At pod admission time, Workload Autoscaler computes the CPU request as |
vertical:
resourceStrategy:
type: nodeAllocatablePercentage
nodeAllocatablePercentage:
cpuPercent: 5vertical.resourceStrategy.nodeAllocatablePercentage.memoryPercent
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
memoryPercent | float | One of cpuPercent or memoryPercent is required | Percentage of the node's allocatable memory to request, resolved per container (each container gets its own request of At pod admission time, Workload Autoscaler computes the memory request as |
vertical:
resourceStrategy:
type: nodeAllocatablePercentage
nodeAllocatablePercentage:
memoryPercent: 2vertical.cpu
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
cpu | object | No | - | CPU-specific scaling configuration. |
vertical:
cpu:
target: p80
lookBackPeriod: 24h
min: 100m
max: 1000m
applyThresholdStrategy:
type: defaultAdaptive
overhead: 0.1vertical.cpu.target
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
target | string | No | "p80" | Resource usage target:
|
vertical:
cpu:
target: p80vertical.cpu.lookBackPeriod
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
lookBackPeriod | duration | No | "24h" | Historical resource usage data window to consider for recommendations (3h-168h). See Look-back Period. |
vertical:
cpu:
lookBackPeriod: 24hvertical.cpu.constraints
Defines minimum and maximum CPU bounds for the autoscaler. Two formats are supported:
constraintsobject (required when usingpercentageOfOriginal): specifies a typed constraint withtypeandvaluefields.- Legacy flat format (
min/maxas strings): still supported for constant values and not planned for deprecation.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
constraints | object | No | - | Container for min and max CPU constraint objects. Required when using percentageOfOriginal. |
vertical.cpu.constraints.min
| Field | Type | Required | Description |
|---|---|---|---|
type | string | Yes | constant for an absolute value; percentageOfOriginal for a percentage of the workload's original request. |
value | string | number | Yes | For constant: Kubernetes CPU notation (e.g., "100m", "2"). For percentageOfOriginal: a number (e.g., 90 = 90%). |
vertical.cpu.constraints.max
Same fields as vertical.cpu.constraints.min.
vertical:
cpu:
constraints:
min:
type: percentageOfOriginal
value: 90
max:
type: constant
value: 2000mvertical:
cpu:
min: 100m
max: 2000m
Note
percentageOfOriginalrequires the workload to have defined resource requests. Not supported for custom workloads or injected containers. If requests are not defined, the constraint is treated as having no limit set.
vertical.cpu.applyThreshold
Deprecation NoticeThe
applyThresholdconfiguration option is deprecated but still supported for backward compatibility. We strongly recommend migrating to the newapplyThresholdStrategyconfiguration format for future compatibility and access to the latest features. See applyThresholdStrategy.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
applyThreshold | float | No | 0.1 | The relative difference required between current and recommended resource values to apply a change immediately:
|
vertical:
cpu:
applyThreshold: 0.1vertical.cpu.applyThresholdStrategy
Warning
applyThresholdandapplyThresholdStrategycannot be used simultaneously in a configuration as that will result in an error.applyThresholdStrategyis the latest and recommended configuration option.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
applyThresholdStrategy | object | No | - | Configuration for the strategy used to determine when recommendations should be applied. The strategy determines how the threshold percentage is calculated based on current resource requests. |
vertical:
cpu:
applyThresholdStrategy:
type: defaultAdaptivevertical.cpu.applyThresholdStrategy.type
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type | string |
| "defaultAdaptive" | The type of threshold strategy to use:
|
vertical:
cpu:
applyThresholdStrategy:
type: defaultAdaptive # Using default adaptive thresholdvertical.cpu.applyThresholdStrategy.percentage
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
percentage | float |
| The fixed percentage threshold to use. Value range: 0.01-2.5.
|
vertical:
cpu:
applyThresholdStrategy:
type: percentage # Using fixed percentage threshold
percentage: 0.3 # 30% thresholdvertical.cpu.applyThresholdStrategy.numerator
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
numerator | float |
| 0.5 | Affects the vertical stretch of the threshold function. Lower values create smaller thresholds.
|
vertical:
cpu:
applyThresholdStrategy:
type: customAdaptive # Using custom adaptive threshold
numerator: 0.1
exponent: 0.1
denominator: 2vertical.cpu.applyThresholdStrategy.denominator
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
denominator | float |
| 1 | Affects threshold sensitivity for small workloads. Values close to 0 result in larger thresholds for small workloads. For example, when numerator is 1, exponent is 1 and denominator is 0 the threshold for 0.5 req. CPU will be 200%.
|
vertical:
cpu:
applyThresholdStrategy:
type: customAdaptive
numerator: 0.1
exponent: 0.1
denominator: 2vertical.cpu.applyThresholdStrategy.exponent
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
exponent | float |
| 1 | Controls how quickly the threshold decreases for larger workloads. Lower values prevent extremely small thresholds for large resources.
|
vertical:
cpu:
applyThresholdStrategy:
type: customAdaptive
numerator: 0.1
exponent: 0.1
denominator: 2vertical.cpu.overhead
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
overhead | float | No | 0.1 | Additional resource buffer when applying recommendations (0.0-2.5, e.g., 0.1 = 10%). If a 10% buffer is configured, the issued recommendation will have +10% added to it, so that the workload can handle further increased resource demand. |
vertical:
cpu:
overhead: 0.1vertical.cpu.limit
Relationship to UI optionsThe annotation fields in this section correspond to the Workload resource limits options in the CAST AI console. See Workload resource limits for the UI-level description of each mode, including how the Semi-automatic (CPU) and Automatic (memory) shortcuts map to specific
type,multiplier,onlyIfOriginalExist, andonlyIfOriginalLowercombinations.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
limit | object | No | Configuration for container CPU limit scaling. Default behaviour when not specified:
|
vertical:
cpu:
limit:
type: multiplier
multiplier: 2.0vertical.cpu.limit.type
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type | string | *Yes | Type of CPU limit scaling to apply:
vertical.cpu.limit configuration option, this field becomes required. |
vertical:
cpu:
limit:
type: keepLimitsvertical.cpu.limit.multiplier
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
multiplier | float | *Yes | Value to multiply the requests by to set the limit. The calculation is: *Required when type is set to |
vertical:
cpu:
limit:
type: multiplier
multiplier: 2.0vertical.cpu.limit.onlyIfOriginalExist
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
onlyIfOriginalExist | boolean | No | false | When set to This flag allows conditional limit management based on the original workload configuration. Only applicable when the type is set to |
vertical:
cpu:
limit:
type: multiplier
multiplier: 2.0
onlyIfOriginalExist: truevertical.cpu.limit.onlyIfOriginalLower
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
onlyIfOriginalLower | boolean | No | false | When set to This flag prevents reducing existing limits and ensures limits only increase when beneficial. Only applicable when the type is set to |
vertical:
cpu:
limit:
type: multiplier
multiplier: 2.0
onlyIfOriginalLower: trueCombining both flags:
When both onlyIfOriginalExist and onlyIfOriginalLower are set to true, the behavior matches Memory's Automatic mode: limits are only set when the workload originally had limits defined AND only when those original limits are lower than the calculated value.
vertical:
cpu:
limit:
type: multiplier
multiplier: 1.5
onlyIfOriginalExist: true
onlyIfOriginalLower: truevertical.cpu.optimization
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
optimization | string | no | This configuration option can be used to disable CPU management for workloads that benefit from memory management only. The workload will then use CPU requests/limits configured in the template.
|
vertical:
cpu:
optimization: offvertical.memory
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
memory | object | No | - | Memory-specific scaling configuration. |
vertical:
memory:
target: max
lookBackPeriod: 24h
min: 128Mi
max: 2Gi
applyThresholdStrategy:
type: defaultAdaptive
overhead: 0.1vertical.memory.target
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
target | string | No | "max" | Resource usage target:
|
vertical:
memory:
target: maxvertical.memory.lookBackPeriod
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
lookBackPeriod | duration | No | "24h" | Historical resource usage data window to consider for recommendations (3h-168h). See Look-back Period. |
vertical:
memory:
lookBackPeriod: 24hvertical.memory.constraints
Defines minimum and maximum memory bounds for the autoscaler. Two formats are supported:
constraintsobject (required when usingpercentageOfOriginal): specifies a typed constraint withtypeandvaluefields.- Legacy flat format (
min/maxas strings): still supported for constant values and not planned for deprecation.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
constraints | object | No | - | Container for min and max memory constraint objects. Required when using percentageOfOriginal. |
vertical.memory.constraints.min
| Field | Type | Required | Description |
|---|---|---|---|
type | string | Yes | constant for an absolute value; percentageOfOriginal for a percentage of the workload's original request. |
value | string | number | Yes | For constant: Kubernetes memory notation (e.g., "128Mi", "2Gi"). For percentageOfOriginal: a number (e.g., 90 = 90%). To control memory limits, see vertical.memory.limit. |
vertical.memory.constraints.max
Same fields as vertical.memory.constraints.min. Note: This does not cap the memory limit directly. To control memory limits, see vertical.memory.limit.
vertical:
memory:
constraints:
min:
type: constant
value: 128Mi
max:
type: percentageOfOriginal
value: 150vertical:
memory:
min: 128Mi
max: 2Gi
Note
percentageOfOriginalrequires the workload to have defined resource requests. Not supported for custom workloads or injected containers. If requests are not defined, the constraint is treated as having no limit set.
vertical.memory.applyThreshold
Deprecation NoticeThe
applyThresholdconfiguration option is deprecated but still supported for backward compatibility. We strongly recommend migrating to the newapplyThresholdStrategyconfiguration format for future compatibility and access to the latest features. See applyThresholdStrategy.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
applyThreshold | float | No | 0.1 | The relative difference required between current and recommended resource values to apply a change immediately:
|
vertical:
memory:
applyThreshold: 0.1vertical.memory.applyThresholdStrategy
Warning
applyThresholdandapplyThresholdStrategycannot be used simultaneously in a configuration as that will result in an error.applyThresholdStrategyis the latest and recommended configuration option.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
applyThresholdStrategy | object | No | - | Configuration for the strategy used to determine when recommendations should be applied. The strategy determines how the threshold percentage is calculated based on current resource requests. |
vertical:
memory:
applyThresholdStrategy:
type: customAdaptive # Using custom adaptive threshold
numerator: 0.1
exponent: 0.1
denominator: 2vertical.memory.applyThresholdStrategy.type
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type | string |
| "defaultAdaptive" | The type of threshold strategy to use:
|
vertical:
memory:
applyThresholdStrategy:
type: defaultAdaptive # Using default adaptive thresholdvertical.memory.applyThresholdStrategy.percentage
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
percentage | float |
| The fixed percentage threshold to use. Value range: 0.01-2.5.
|
vertical:
memory:
applyThresholdStrategy:
type: percentage # Using fixed percentage threshold
percentage: 0.3 # 30% thresholdvertical.memory.applyThresholdStrategy.numerator
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
numerator | float |
| 0.5 | Affects the vertical stretch of the threshold function. Lower values create smaller thresholds.
|
vertical:
memory:
applyThresholdStrategy:
type: customAdaptive # Using custom adaptive threshold
numerator: 0.1
exponent: 0.1
denominator: 2vertical.memory.applyThresholdStrategy.denominator
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
denominator | float |
| 1 | Affects threshold sensitivity for small workloads. Values close to 0 result in larger thresholds for small workloads. For example, when numerator is 1, exponent is 1 and denominator is 0 the threshold for 0.5 req. Memory will be 200%.
|
vertical:
memory:
applyThresholdStrategy:
type: customAdaptive
numerator: 0.1
exponent: 0.1
denominator: 2vertical.cpu.applyThresholdStrategy.exponent
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
exponent | float |
| 1 | Controls how quickly the threshold decreases for larger workloads. Lower values prevent extremely small thresholds for large resources.
|
vertical:
memory:
applyThresholdStrategy:
type: customAdaptive
numerator: 0.1
exponent: 0.1
denominator: 2vertical.memory.overhead
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
overhead | float | No | 0.1 | Additional resource buffer when applying recommendations (0.0-2.5, e.g., 0.1 = 10%). If a 10% buffer is configured, the issued recommendation will have +10% added to it so that the workload can handle further increased resource demand. |
vertical:
memory:
overhead: 0.1vertical.memory.limit
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
limit | object | No | Configuration for container memory limit scaling. The default behavior when not specified:
|
vertical:
memory:
limit:
type: multiplier
multiplier: 1.5vertical.memory.limit.type
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type | string | *Yes | Type of limit scaling to apply:
vertical.memory.limit configuration option, this field (type) becomes required. |
vertical:
memory:
limit:
type: keepLimitsvertical.memory.limit.multiplier
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
multiplier | float |
| Value to multiply the requests by to set the limit on the workload. The calculation is: *Required when type is set to |
vertical:
memory:
limit:
type: multiplier
multiplier: 1.5vertical.memory.optimization
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
optimization | string | no | This configuration option can be used to disable memory management for workloads that benefit from CPU management only (e.g., Java workloads with a fixed heap size). The workload will then use memory requests/limits configured in the template.
|
vertical:
memory:
optimization: offvertical.excludedContainers
Specifies container names to exclude from vertical autoscaling optimization. Excluded containers retain their current resource settings and are not scaled by the Workload Autoscaler. Recommendations are still generated and visible in event logs and the workload detail view.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
excludedContainers | array | No | - | List of container names to exclude from optimization. Names must match exactly as defined in the pod specification. Regex or glob patterns are not supported. |
vertical:
excludedContainers:
- istio-proxy
- logging-agent
Note
excludedContainerstargets containers explicitly defined in your pod spec, such as application containers or native Kubernetes sidecar containers.
vertical.predictiveScaling
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
predictiveScaling | object | No | - | Predictive scaling configuration for CPU. |
vertical:
predictiveScaling:
cpu:
enabled: truevertical.predictiveScaling.cpu
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
cpu | object | No | - | CPU-specific predictive scaling settings. |
vertical:
predictiveScaling:
cpu:
enabled: truevertical.predictiveScaling.cpu.enabled
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
enabled | boolean | No | false | Enable predictive scaling for CPU resources. When enabled, the system forecasts CPU usage based on historical patterns and generates proactive recommendations. Requires Predictive scaling can be enabled on any workload via annotations, even if it is not currently eligible for scaling in this manner. The system will automatically activate predictive scaling once the workload becomes predictable and will seamlessly revert to standard scaling if the patterns are lost. This allows preemptive enablement without monitoring for eligibility. See Predictive workload scaling for more information. |
vertical:
predictiveScaling:
cpu:
enabled: truevertical.downscaling
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
downscaling | object | No | - | Downscaling behavior override. |
vertical:
downscaling:
applyType: immediatevertical.downscaling.applyType
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
applyType | string | No | Default is taken from the vertical scaling policy controlling the workload. | Override application mode:
|
vertical:
downscaling:
applyType: immediatevertical.memoryEvent
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
memoryEvent | object | No | - | Memory event behavior override. |
vertical:
memoryEvent:
applyType: immediatevertical.memoryEvent.applyType
This configuration option is fully compatible with other applyType options and is meant to be used in combination with them. This allows for fine-grained control over both upscaling and downscaling. Here's how they interact:
- If both configuration options are set to the same value (both
immediateor bothdeferred), the behavior remains unchanged. - If
vertical.downscaling.applyTypeis set todeferredandvertical.memoryEvent.applyTypeis set toimmediate:- Upscaling operations will be applied immediately.
- Downscaling operations will be deferred to natural pod restarts.
- If
vertical.downscaling.applyTypeis set toimmediateandvertical.memoryEvent.applyTypeis set todeferred:- Upscaling operations will be deferred to natural pod restarts.
- Downscaling operations will be applied immediately.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
applyType | string |
| Default is taken from the vertical scaling policy controlling the workload. | Override application mode for memory-related events (OOM kills, pressure):
|
vertical:
memoryEvent:
applyType: immediatevertical.containers
Configuration object that contains container-specific settings.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| containers | object | No | - | Container configuration mapping. |
vertical:
containers:vertical.containers.{container_name}
Configuration object that contains resource constraints for a specific container. Replace {container_name} with the name of your container.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| {container_name} | object | No | - | Container resource configuration object. Has to be the name of the container for which resources are being configured. |
vertical:
containers:
? { container_name }
NoteContainer constraints apply to application containers (
spec.containers[]) and native sidecar containers (spec.initContainers[]withrestartPolicy: Always). Traditional initContainers that run during pod startup are not optimized by Workload Autoscaler.
vertical.containers.{container_name}.resourceStrategy
Container-level sizing strategy. Sets a resource strategy for a specific container, overriding vertical.resourceStrategy for that container. Supports the same type and nodeAllocatablePercentage fields as the workload-level strategy. See Node-aware DaemonSet sizing.
Setting the strategy at the container level sizes only that container. It can also add requests to a container that has none in its pod spec. The workload-level strategy, by contrast, only applies to containers that already have requests.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
resourceStrategy.type | string | Yes | - | nodeAllocatablePercentage: size requests as a percentage of the node's allocatable CPU and memory. |
resourceStrategy.nodeAllocatablePercentage.cpuPercent | float | One of cpuPercent or memoryPercent is required | - | Percentage of node-allocatable CPU to request. Valid range: (0, 100]. |
resourceStrategy.nodeAllocatablePercentage.memoryPercent | float | One of cpuPercent or memoryPercent is required | - | Percentage of node-allocatable memory to request. Valid range: (0, 100]. |
vertical:
containers:
app:
resourceStrategy:
type: nodeAllocatablePercentage
nodeAllocatablePercentage:
cpuPercent: 30
memoryPercent: 40
NoteContainer
min/maxconstraints and the resource strategy are per container, but limits andlookBackPeriodare workload-level only. Per-containerlimitandlookBackPeriodsettings are not supported and are silently ignored.
vertical.containers.{container_name}.cpu
Container CPU constraints. Set minimum and maximum CPU limits for a specific container to define the workload autoscaler's scaling range. With the nodeAllocatablePercentage strategy, these constrain (clamp) the resolved percentage for this container. The percentageOfOriginal type is not yet supported at container level — only constant is. The constraints.min and constraints.max blocks are optional, but if specified, both type and value are required.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
constraints.min.type | string | Yes | - | Must be constant. |
constraints.min.value | string | Yes | - | Kubernetes CPU notation (e.g., "10m", "1"). Min cannot be greater than max. |
constraints.max.type | string | Yes | - | Must be constant. |
constraints.max.value | string | Yes | - | Kubernetes CPU notation (e.g., "1000m", "2"). Recommendations won't exceed this value. |
vertical:
containers:
{ container_name }:
cpu:
constraints:
min:
type: constant
value: 10m
max:
type: constant
value: 1000mvertical.containers.{container_name}.memory
Container memory request constraints. Set minimum and maximum memory requests for a specific container to define the workload autoscaler's scaling range. With the nodeAllocatablePercentage strategy, these constrain (clamp) the resolved percentage for this container. These values control requests only; memory limits are managed separately via vertical.memory.limit. The percentageOfOriginal type is not yet supported at container level — only constant is. The constraints.min and constraints.max blocks are optional, but if specified, both type and value are required.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
constraints.min.type | string | Yes | - | Must be constant. |
constraints.min.value | string | Yes | - | Kubernetes memory notation (e.g., "10Mi", "2Gi"). Recommendations will not go below this value. |
constraints.max.type | string | Yes | - | Must be constant. |
constraints.max.value | string | Yes | - | Kubernetes memory notation (e.g., "2048Mi", "4Gi"). Recommendations will not exceed this value. |
vertical:
containers:
{ container_name }:
memory:
constraints:
min:
type: constant
value: 10Mi
max:
type: constant
value: 2048Mivertical.containers.{container_name}.jvm
Container-level JVM optimization settings. These override the workload-level vertical.jvm settings for a specific container, allowing you to enable or disable JVM optimization and auto-instrumentation on individual containers within a multi-container pod. For details on JVM optimization, see JVM workload optimization.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
jvm | object | No | - | Container-level JVM heap optimization and auto-instrumentation configuration. |
vertical:
containers:
{ container_name }:
jvm:
enabled: true
autoInstrument: truevertical.containers.{container_name}.jvm.enabled
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
enabled | boolean | No | false | Enable JVM heap-based memory optimization for this specific container. Overrides the workload-level vertical.jvm.enabled setting for the named container. |
vertical:
containers:
{ container_name }:
jvm:
enabled: truevertical.containers.{container_name}.jvm.autoInstrument
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
autoInstrument | boolean | No | false | Enable automatic JMX agent injection for this specific container. Overrides the workload-level vertical.jvm.autoInstrument setting. Pods must be restarted for injection to take effect. |
vertical:
containers:
{ container_name }:
jvm:
autoInstrument: truevertical.excludedContainers
Specifies container names to exclude from vertical autoscaling. Excluded containers retain their current resource settings and are not scaled by the Workload Autoscaler. Recommendations are still generated and visible, but not applied.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
excludedContainers | array | No | - | List of container names to exclude from optimization. Names must match exactly as defined in the pod specification. |
vertical:
excludedContainers:
- istio-proxy
- logging-agent
Note
excludedContainerstargets containers explicitly defined in your pod spec, such as application containers or native Kubernetes sidecar containers.
vertical.jvm
JVM-specific optimization settings for Java workloads. When enabled, the Workload Autoscaler uses actual JVM runtime data (heap usage, GC behavior, thread count) to generate memory recommendations instead of relying on standard Kubernetes container memory metrics. For a full explanation of JVM optimization, auto-instrumentation, prerequisites, and troubleshooting, see JVM workload optimization.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
jvm | object | No | - | JVM heap-based memory optimization and auto-instrumentation. |
vertical:
jvm:
enabled: true
autoInstrument: truevertical.jvm.enabled
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
enabled | boolean | No | false | Enable JVM heap-based memory optimization for this workload. When enabled, the Workload Autoscaler uses JVM heap metrics to compute memory recommendations and injects -Xms/-Xmx flags. Requires JVM metrics to be available — either through autoInstrument or an existing Prometheus pipeline. See prerequisites. |
vertical:
jvm:
enabled: truevertical.jvm.autoInstrument
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
autoInstrument | boolean | No | false | Enable automatic JMX agent injection for this workload. When enabled, the Workload Autoscaler injects a JMX Prometheus Java agent into Java containers at pod admission time. Pods must be restarted for injection to take effect. See auto-instrumentation for details. |
vertical:
jvm:
autoInstrument: truevertical.hpaConverters
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
hpaConverters | array | No | - | List of HPA conversion strategies to apply when vertical optimization changes resource requests. Prevents percentage-based HPA utilization targets from drifting after Cast AI adjusts requests. See HPA converter for details. |
vertical:
hpaConverters:
- type: AverageValueFromOriginalRequestsvertical.hpaConverters[].type
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type | string | Yes | The conversion strategy to use. Supported values:
|
vertical:
hpaConverters:
- type: AverageValueFromOriginalRequests
NoteThe HPA converter applies to workloads that have their own HPA while Cast AI handles only vertical optimization. If Cast AI's horizontal autoscaling is also enabled for the workload, Cast AI already manages the HPA directly and the converter is not needed. Requires Workload Autoscaler version v1.0.11 or later.
rolloutBehavior
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
rolloutBehavior | object | No | - | Configuration for controlling how recommendations are rolled out. |
rolloutBehavior:
type: NoDisruption
preferOneByOne: true
delaySeconds: 120rolloutBehavior.type
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type | string | No | - | Controls how recommendation updates are rolled out.
For comprehensive workload requirements, see Configuring zero-downtime updates. |
rolloutBehavior:
type: NoDisruptionrolloutBehavior.preferOneByOne
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
preferOneByOne | boolean | No | false | When See Rollout behavior for more information. |
rolloutBehavior:
preferOneByOne: truerolloutBehavior.delaySeconds
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
delaySeconds | integer | No | - | Minimum number of seconds Workload Autoscaler waits between evictions when rolling out a recommendation. This delay is enforced regardless of whether preferOneByOne is enabled. Accepted range: 0–3600. |
rolloutBehavior:
delaySeconds: 120horizontal
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
horizontal | object | No | - | Horizontal autoscaling configuration. |
When horizontal autoscaling is configured via annotations, Cast AI creates and manages a native Kubernetes HorizontalPodAutoscaler (autoscaling/v2) resource in the cluster. For concepts, supported workload types, and compatibility details, see Horizontal autoscaling.
horizontal:
optimization: true
useNative: true
minReplicas: 2
maxReplicas: 10
metrics:
- type: resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
behavior:
scaleDown:
stabilizationWindowSeconds: 300horizontal.optimization
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
optimization | boolean | Yes* | Controls whether Cast AI actively manages horizontal scaling for this workload. Set to *Required when the |
horizontal:
optimization: truehorizontal.useNative
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
useNative | boolean | Yes* | false | Enables native Kubernetes HPA mode. When *Required to enable horizontal autoscaling. |
horizontal:
useNative: truehorizontal.takeOwnership
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
takeOwnership | boolean | No | false | When true, Cast AI takes ownership of an existing native HPA on the workload and replaces its configuration with Cast AI-managed settings. The existing HPA must not be managed by a third party such as KEDA. Only HPAs with CPU or memory utilization triggers are currently eligible for ownership. |
horizontal:
takeOwnership: truehorizontal.minReplicas
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
minReplicas | integer | Yes* | - | Minimum number of pod replicas. Must be greater than 0. |
horizontal:
minReplicas: 2horizontal.maxReplicas
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
maxReplicas | integer | Yes* | - | Maximum number of pod replicas. Must be greater than 0 and ≥ minReplicas. |
horizontal:
maxReplicas: 10horizontal.metrics
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
metrics | array | Yes* | - | Array of metric objects that define scaling triggers. At least one metric is required. Only resource metric type (CPU or memory) is supported via annotations. |
horizontal:
metrics:
- type: resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70metrics[].type
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type | string | Yes | - | The metric source type. Only resource is supported via annotations. |
metrics[].resource.name
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
name | string | Yes | - | The resource to monitor: cpu or memory. |
metrics[].resource.target.type
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type | string | Yes | - | How the target value is interpreted. Accepted values: Utilization (percentage of resource requests averaged across all pods), AverageValue (average raw value across all pods), or Value (raw value). |
metrics[].resource.target.averageUtilization
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
averageUtilization | integer | Conditional | - | Target average utilization percentage across all pods. Required when target.type is Utilization. |
metrics:
- type: resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70metrics[].resource.target.averageValue
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
averageValue | string | Conditional | - | Target average value as a Kubernetes quantity string (e.g., "100m"). Required when target.type is AverageValue. |
metrics[].resource.target.value
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
value | string | Conditional | - | Target value as a Kubernetes quantity string (e.g., "100m"). Required when target.type is Value. |
horizontal.behavior
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
behavior | object | No | - | Custom scaling behavior for scale-up and scale-down directions. Maps directly to the Kubernetes HPA v2 behavior spec. |
horizontal:
behavior:
scaleUp:
stabilizationWindowSeconds: 60
selectPolicy: Max
tolerance: "0.1"
policies:
- type: Pods
value: 4
periodSeconds: 60
scaleDown:
stabilizationWindowSeconds: 300
selectPolicy: Min
policies:
- type: Percent
value: 10
periodSeconds: 60horizontal.behavior.scaleUp / horizontal.behavior.scaleDown
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
scaleUp | object | No | - | Scale-up behavior configuration. |
scaleDown | object | No | - | Scale-down behavior configuration. |
The scaleUp and scaleDown objects share the same structure. The following fields apply to both.
behavior.direction.stabilizationWindowSeconds
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
stabilizationWindowSeconds | integer | No | - | The number of seconds the autoscaler looks back at previous scaling recommendations before acting. This prevents rapid fluctuations in replica count. Valid range: 0–3600. |
behavior:
scaleDown:
stabilizationWindowSeconds: 300behavior.direction.selectPolicy
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
selectPolicy | string | No | - | When multiple scaling policies are defined, determines which one to use. Max selects the policy allowing the most change, Min selects the most conservative, and Disabled prevents scaling in this direction entirely. |
behavior.direction.tolerance
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
tolerance | string | No | - | A non-negative quantity value. Scaling is triggered only when the metric deviation exceeds this tolerance. Must be ≥ 0. |
behavior:
scaleUp:
tolerance: "0.1"behavior.direction.policies
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
policies | array | No | - | Array of scaling policy rules that control the rate of change. |
policies[].type
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type | string | Yes | - | The unit for the scaling rate limit: Pods (absolute count) or Percent (percentage of current replicas). |
policies[].value
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
value | integer | Yes | - | Maximum amount of change allowed per period. Must be > 0. |
policies[].periodSeconds
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
periodSeconds | integer | Yes | - | Time window in seconds for the scaling rate limit. Valid range: 1–1800. |
policies:
- type: Pods
value: 4
periodSeconds: 60
- type: Percent
value: 100
periodSeconds: 60(Deprecated) legacy horizontal fields
DeprecatedThe following fields apply only to legacy horizontal scaling. They are not compatible with native HPA mode (
useNative: true). New configurations should use the fields documented in the horizontal section above. To migrate existing workloads, see Migrate from legacy horizontal scaling.
When using legacy Cast AI horizontal scaling (without useNative: true), the horizontal block uses the following structure:
horizontal:
optimization: on
minReplicas: 1
maxReplicas: 10
scaleDown:
stabilizationWindow: 5m
shortAverage: 3mThe minReplicas and maxReplicas fields work identically to the native HPA configuration documented above. The fields below are specific to legacy mode.
horizontal.optimization (legacy)
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
optimization | string | Yes* | Enable legacy horizontal scaling. Set to In native HPA mode, this field is a boolean ( *Required when the |
horizontal:
optimization: onhorizontal.scaleDown.stabilizationWindow (legacy)
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
stabilizationWindow | duration | *Yes | "5m" | Cooldown period between scale-downs. The duration needs to be parsable (e.g., *Required if the |
horizontal:
scaleDown:
stabilizationWindow: 5mhorizontal.shortAverage (legacy)
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
shortAverage | duration | No | "3m" | Time period to average CPU metrics over before making scaling decisions. Valid range: 1–10 minutes. Not applicable to native HPA mode — this field is ignored when useNative: true. |
horizontal:
shortAverage: 3mcontainersGrouping
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
containersGrouping | array | No | - | Rules for grouping dynamically generated containers with similar naming patterns. |
containersGrouping:
- key: name
operator: contains
values: ["data-processor"]
into: data-processorcontainersGrouping.[].key
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
key | string | *Yes | - | The attribute used to match containers. Currently, only supports name which refers to the container name property. |
containersGrouping.[].operator
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
operator | string | *Yes | - | Defines how the key is evaluated against the values list. Currently, only supports contains. |
containersGrouping.[].values
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
values | array | *Yes | - | A list of string values used for matching against the key with the specified operator. Must contain at least one item. |
containersGrouping.[].into
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
into | string | *Yes | - | The target container name into which matching containers should be grouped. |
Example usage
See Container Grouping for Dynamic Containers.
*Required if the parent object is present in the configuration.
schedule
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
schedule | object | No | - | Controls whether a custom workload is treated as job-like (sporadic execution) or continuous (always running) for optimization purposes. |
schedule:
type: jobLikeschedule.type
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
type | string | No | Explicitly sets whether the custom workload should be treated as job-like or continuous:
|
schedule:
type: jobLike # For sporadic workloads like batch jobs
ImportantThe
schedule.typeconfiguration only applies to custom job-like workloads identified by theworkloads.cast.ai/custom-workloadlabel. Native Kubernetes Jobs with this label are always treated as job-like, and standard workload types (Deployments, StatefulSets, etc.) are always treated as continuous.
Legacy Annotation SupportFor documentation on the legacy annotation format, which is now deprecated, see the Legacy Annotations Reference page .
Migration Guide
NoteThe annotations V2 structure cannot be combined with deprecated annotations V1. When the annotation
workloads.cast.ai/configurationis detected, the workload is considered to be configured by using that annotation and all other annotations starting withworkloads.cast.aiwill be ignored.
To migrate from v1 to v2 annotations:
- Remove all individual legacy
workloads.cast.ai/*annotations - Add the new
workloads.cast.ai/configurationannotation - Move all settings into the YAML structure under the new annotation
For example, these v1 annotations:
workloads.cast.ai/vertical-autoscaling: "on"
workloads.cast.ai/cpu-target: "p80"
workloads.cast.ai/memory-max: "2Gi"Would become:
workloads.cast.ai/configuration: |
vertical:
optimization: on
cpu:
target: p80
memory:
max: 2GiUpdated 3 days ago
