Skip to content
NewKubeOn AI: a managed LLM gateway with cost per 1M tokens for every model

Glossary

Kubernetes cost terms, defined

The words KubeOn uses on every screen, with the rule behind each one.

A

Abandoned workload
A workload that uses almost none of what it requests over a long window. KubeOn flags CPU at or below 2% of request over 14 days.
Amortized cost
Cost with upfront and recurring commitment fees (Savings Plans, reservations) spread across the hours they cover, so each day carries its share.
Anomaly
A day of spend well above its recent baseline. KubeOn compares each day with the median of the prior 14 days and requires both a dollar and a percentage increase.

C

Cached tokens
Prompt tokens a provider serves from its prompt cache at a lower price. KubeOn AI prices them separately from input and output tokens.
Chargeback
Billing internal teams or cost centers for the infrastructure they use, usually through the finance system.
Cost allocation
Assigning spend to the teams, cost centers, applications or labels responsible for it.
Cost per 1M tokens
A unit cost for language models: allocated model spend divided by tokens processed, times one million. KubeOn AI includes PTU and reservation cost in it.
CUR 2.0
The current format of the AWS Cost and Usage Report, delivered as a Data Export. It is the only AWS export with split cost allocation data.

D

Direct cost
The part of a node's cost charged to a specific pod, before shared and idle cost are applied.

E

Effective-dated mapping
An ownership mapping that records when it started and ended, so changing an owner does not rewrite past allocations.
Efficiency
The share of paid-for capacity that was used. KubeOn weights each resource's efficiency by its cost.

F

Fallback
Sending a request to another target when the first returns a throttling error, a server error or times out.
FinOps
The practice of bringing financial accountability to variable cloud spend, so engineering, finance and product make cost trade-offs together.
FOCUS
The FinOps Open Cost and Usage Specification, a common schema for billing data across providers.

H

Headroom
Capacity added on top of measured usage when sizing a request or node. KubeOn adds 20%.

I

Idle cost
The cost of capacity no workload used. Infrastructure idle is node capacity nobody requested; workload idle is requested capacity that went unused.
IRSA
IAM Roles for Service Accounts: lets a pod on EKS assume an IAM role through the cluster's OIDC provider, without static keys.

K

Karpenter
An open-source Kubernetes node autoscaler. Its NodePool resource defines which instance types a cluster may launch.

L

LLM gateway
A single endpoint in front of model providers that routes requests, applies keys, budgets and limits, and records each request's tokens and cost.

M

max(request, usage)
The allocation rule that charges each pod for the larger of what it reserved and what it used, per resource and per hour.

N

Net unblended cost
The cost of a line item after discounts such as EDP and credits are applied, before amortization.

O

Orphaned volume
A cloud disk that still bills after the Kubernetes PersistentVolume that used it is gone.

P

p95
The 95th percentile: the value usage stays at or below 95% of the time. Used for sizing because it covers busy hours without chasing rare spikes.
Pay-as-you-go
Model usage billed per token, with no commitment.
Provisioned throughput
Reserved model capacity billed by the hour or by term, such as Azure OpenAI PTUs, Bedrock model units or Vertex AI GSUs.
PTU
Provisioned Throughput Unit: Azure OpenAI capacity bought in fixed units, deployed in Global, Data Zone or Regional pools.

R

Rate card
A price list for capacity you own, such as cost per vCPU-hour or GiB-hour, used to price on-premises clusters.
Reconciliation
Checking that allocated cost adds up to the billed cost, and listing what could not be attributed.
Request
The CPU and memory a container asks the scheduler to reserve for it. Requests decide how many pods fit on a node.
Rightsizing
Adjusting requests, node sizes or volume sizes to match measured usage plus headroom.
Route
In the KubeOn AI gateway, a named list of targets and rules that apps call instead of a vendor model.

S

Shared cost
Cost of platform services many teams use, such as kube-system, monitoring and ingress controllers.
Showback
Reporting each team's costs to it without moving budget, as a step before or instead of chargeback.
Spillover
Sending traffic beyond a provisioned deployment's capacity to pay-as-you-go instead of throttling it.
Split cost allocation data
AWS's per-pod cost for EKS, included in CUR 2.0, computed from requests and usage.

T

Tokens per minute (TPM)
The rate of tokens a deployment or key may process each minute; used for capacity sizing and rate limits.

U

Unallocated cost
Billed cost that cannot be joined to a cluster resource or an owner. KubeOn lists it with resource IDs instead of hiding it in idle.

See your own clusters in KubeOn.

A 30-minute walkthrough on your billing data, with an engineer who has run Kubernetes cost programs.