Cluster-level savings report

Drill into savings, capacity, and workload optimization metrics for a single cluster.

The cluster-level savings report drills into a single cluster, showing how much Cast AI has saved on that cluster and where those savings come from. It adapts its sections based on which optimization components are active.

For the shared concepts behind this report, see the Savings Report overview.

You can reach the cluster-level view by clicking a cluster in the organization-level savings report breakdown table.

Overview cards

The top of the report shows the same summary cards as the organization-level view, scoped to this cluster:

  • Baseline spend — the modeled cost of running this cluster without Cast AI.
  • Realized savings — savings realized through node autoscaler including workload autoscaler impact if both are enabled (shown only if this cluster has node autoscaler managing at least 20% of nodes).
  • Workload autoscaler savings — modeled savings from the workload autoscaler (shown whenever workload autoscaler is actively managing at least 20% of workloads in VPA mode on this cluster).
  • Actual spend — the real spending for this cluster over the selected period.

Below the cards, the savings over time chart shows realized savings and workload autoscaler savings trends, depending on which autoscalers are adopted on this cluster.

Node autoscaler impact

This section appears when the node autoscaler is active on the cluster, managing at least 20% of the nodes.

Provisioned capacity

Shows baseline, actual, and reduction in provisioned resources, for both CPU and memory. Baseline is what the cluster would have provisioned without Cast AI; actual is what's currently provisioned; reduction is the difference.

Average node count

Compares the average number of nodes the cluster would have run without Cast AI against the actual average node count, with the resulting savings.

Average cost per node

Compares the average cost per node without Cast AI against the actual average cost per node, with the resulting savings.

Workload autoscaler impact

This section appears when the workload autoscaler is active on the cluster, managing at least 20% of workloads in VPA mode.

Workload resource requests

Shows average original requests, average current requests, and resources freed over time — for both CPU and memory. This visualizes how rightsizing has reduced resource demand.

Workload autoscaler adoption

Shows the percentage of workloads managed by the workload autoscaler over time.

Workload utilization rate

Shows CPU and memory utilization (switchable between the two), split into:

  • Workloads managed by the workload autoscaler
  • Unmanaged workloads

Workload optimization breakdown

A table listing every workload in the cluster. For each workload, you can see:

  • Namespace
  • Which optimizations are turned on (VPA, HPA)
  • Workload autoscaler savings
  • Percentage share of total savings

Related resources


What’s Next

Did this page help you?