Cluster-level savings report
Drill into savings, capacity, and workload optimization metrics for a single cluster.
The cluster-level savings report drills into a single cluster, showing how much Cast AI has saved on that cluster and where those savings come from. It adapts its sections based on which optimization components are active.
For the shared concepts behind this report, see the Savings Report overview.
You can reach the cluster-level view by clicking a cluster in the organization-level savings report breakdown table.
Overview cards
The top of the report shows the same summary cards as the organization-level view, scoped to this cluster:
- Baseline spend — the modeled cost of running this cluster without Cast AI.
- Realized savings — savings realized through node autoscaler including workload autoscaler impact if both are enabled (shown only if this cluster has node autoscaler managing at least 20% of nodes).
- Workload autoscaler savings — modeled savings from the workload autoscaler (shown whenever workload autoscaler is actively managing at least 20% of workloads in VPA mode on this cluster).
- Actual spend — the real spending for this cluster over the selected period.
Below the cards, the savings over time chart shows realized savings and workload autoscaler savings trends, depending on which autoscalers are adopted on this cluster.
Node autoscaler impact
This section appears when the node autoscaler is active on the cluster, managing at least 20% of the nodes.
Provisioned capacity
Shows baseline, actual, and reduction in provisioned resources, for both CPU and memory. Baseline is what the cluster would have provisioned without Cast AI; actual is what's currently provisioned; reduction is the difference.
Average node count
Compares the average number of nodes the cluster would have run without Cast AI against the actual average node count, with the resulting savings.
Average cost per node
Compares the average cost per node without Cast AI against the actual average cost per node, with the resulting savings.
Workload autoscaler impact
This section appears when the workload autoscaler is active on the cluster, managing at least 20% of workloads in VPA mode.
Workload resource requests
Shows average original requests, average current requests, and resources freed over time — for both CPU and memory. This visualizes how rightsizing has reduced resource demand.
Workload autoscaler adoption
Shows the percentage of workloads managed by the workload autoscaler over time.
Workload utilization rate
Shows CPU and memory utilization (switchable between the two), split into:
- Workloads managed by the workload autoscaler
- Unmanaged workloads
Workload optimization breakdown
A table listing every workload in the cluster. For each workload, you can see:
- Namespace
- Which optimizations are turned on (VPA, HPA)
- Workload autoscaler savings
- Percentage share of total savings
Related resources
Updated 5 hours ago
