Glossary
Kubernetes cost terms, defined
The words KubeOn uses on every screen, with the rule behind each one.
A
- Abandoned workload
- A workload that uses almost none of what it requests over a long window. KubeOn flags CPU at or below 2% of request over 14 days.
- Amortized cost
- Cost with upfront and recurring commitment fees (Savings Plans, reservations) spread across the hours they cover, so each day carries its share.
- Anomaly
- A day of spend well above its recent baseline. KubeOn compares each day with the median of the prior 14 days and requires both a dollar and a percentage increase.
C
- Cached tokens
- Prompt tokens a provider serves from its prompt cache at a lower price. KubeOn AI prices them separately from input and output tokens.
- Chargeback
- Billing internal teams or cost centers for the infrastructure they use, usually through the finance system.
- Cost allocation
- Assigning spend to the teams, cost centers, applications or labels responsible for it.
- Cost per 1M tokens
- A unit cost for language models: allocated model spend divided by tokens processed, times one million. KubeOn AI includes PTU and reservation cost in it.
- CUR 2.0
- The current format of the AWS Cost and Usage Report, delivered as a Data Export. It is the only AWS export with split cost allocation data.
D
- Direct cost
- The part of a node's cost charged to a specific pod, before shared and idle cost are applied.
E
- Effective-dated mapping
- An ownership mapping that records when it started and ended, so changing an owner does not rewrite past allocations.
- Efficiency
- The share of paid-for capacity that was used. KubeOn weights each resource's efficiency by its cost.
F
- Fallback
- Sending a request to another target when the first returns a throttling error, a server error or times out.
- FinOps
- The practice of bringing financial accountability to variable cloud spend, so engineering, finance and product make cost trade-offs together.
- FOCUS
- The FinOps Open Cost and Usage Specification, a common schema for billing data across providers.
H
- Headroom
- Capacity added on top of measured usage when sizing a request or node. KubeOn adds 20%.
I
- Idle cost
- The cost of capacity no workload used. Infrastructure idle is node capacity nobody requested; workload idle is requested capacity that went unused.
- IRSA
- IAM Roles for Service Accounts: lets a pod on EKS assume an IAM role through the cluster's OIDC provider, without static keys.
K
- Karpenter
- An open-source Kubernetes node autoscaler. Its NodePool resource defines which instance types a cluster may launch.
L
- LLM gateway
- A single endpoint in front of model providers that routes requests, applies keys, budgets and limits, and records each request's tokens and cost.
M
- max(request, usage)
- The allocation rule that charges each pod for the larger of what it reserved and what it used, per resource and per hour.
N
- Net unblended cost
- The cost of a line item after discounts such as EDP and credits are applied, before amortization.
O
- Orphaned volume
- A cloud disk that still bills after the Kubernetes PersistentVolume that used it is gone.
P
- p95
- The 95th percentile: the value usage stays at or below 95% of the time. Used for sizing because it covers busy hours without chasing rare spikes.
- Pay-as-you-go
- Model usage billed per token, with no commitment.
- Provisioned throughput
- Reserved model capacity billed by the hour or by term, such as Azure OpenAI PTUs, Bedrock model units or Vertex AI GSUs.
- PTU
- Provisioned Throughput Unit: Azure OpenAI capacity bought in fixed units, deployed in Global, Data Zone or Regional pools.
R
- Rate card
- A price list for capacity you own, such as cost per vCPU-hour or GiB-hour, used to price on-premises clusters.
- Reconciliation
- Checking that allocated cost adds up to the billed cost, and listing what could not be attributed.
- Request
- The CPU and memory a container asks the scheduler to reserve for it. Requests decide how many pods fit on a node.
- Rightsizing
- Adjusting requests, node sizes or volume sizes to match measured usage plus headroom.
- Route
- In the KubeOn AI gateway, a named list of targets and rules that apps call instead of a vendor model.
S
- Shared cost
- Cost of platform services many teams use, such as kube-system, monitoring and ingress controllers.
- Showback
- Reporting each team's costs to it without moving budget, as a step before or instead of chargeback.
- Spillover
- Sending traffic beyond a provisioned deployment's capacity to pay-as-you-go instead of throttling it.
- Split cost allocation data
- AWS's per-pod cost for EKS, included in CUR 2.0, computed from requests and usage.
T
- Tokens per minute (TPM)
- The rate of tokens a deployment or key may process each minute; used for capacity sizing and rate limits.
U
- Unallocated cost
- Billed cost that cannot be joined to a cluster resource or an owner. KubeOn lists it with resource IDs instead of hiding it in idle.
See your own clusters in KubeOn.
A 30-minute walkthrough on your billing data, with an engineer who has run Kubernetes cost programs.