KubeOn AI
Know what every model call costs, and route it well
KubeOn AI measures cost per 1M tokens across every provider, sizes your provisioned throughput, and runs a managed LLM gateway so your teams do not have to build routing, fallback or budgets themselves.
The problem
AI spend grows faster than anyone can explain it
The bill stops at the deployment
Invoices show a cost per deployment or model, not which product, team or customer spent it.
Three ways to pay for one model
Pay-as-you-go, PTU by the hour and reservations land in different meters, and some are recorded in the wrong one.
Provisioned capacity sits idle
PTUs are paid every hour, while traffic still goes to pay-as-you-go because nothing routes it.
Every team builds its own client
Retries, fallback, keys and rate limits are reimplemented in each service, and none of them know the budget.
Two ways to start
Read-only on day one, the gateway when you are ready
Read-only
Connect billing and telemetry. Nothing in your traffic changes.
- Cost, tokens and requests per deployment
- PTU sizing and reservation coverage
- Budgets and alerts by product, team and tenant
Through the gateway
Change the base URL. KubeOn AI routes and meters every request.
- Exact tokens and cost per request, key and route
- PTU-first spillover and fallback across providers
- Budgets that stop at 100%, rate limits, cache and PII redaction
Product
Four areas, one AI cost model
LLM gateway
Change one URL and stop managing routing
Apps call a route such as chat-default. The gateway sends it to your PTU deployment until it is busy, spills over to pay-as-you-go, falls back to another provider when one throttles, and records the cost of every request against the team's key.
- OpenAI and Anthropic compatible
- Budgets and rate limits per key
- Managed for you, or self-hosted in your cluster
import osfrom openai import OpenAIclient = OpenAI( base_url="https://gateway.kubeon.io/v1", api_key=os.environ["KUBEON_AI_KEY"], # a team key, not a provider key)resp = client.chat.completions.create( model="chat-default", # a route name, not a vendor model messages=[{"role": "user", "content": "Summarize this ticket in two lines."}],)print(resp.choices[0].message.content)Tokenomics
Cost per 1M tokens that includes your commitments
PTU hours and reservations are allocated to the deployments and models that used them, so pay-as-you-go and provisioned capacity can be compared on the same number, by model, team and cloud.
- Input, output and cached tokens
- Past PTU billing corrected in place
- Model-mix savings from observed cost
Providers
Every model provider your teams use
Hosted APIs, cloud model platforms and models you run yourself, in one cost model.
- Azure OpenAI
- Azure AI Foundry
- Amazon Bedrock
- Amazon SageMaker
- Google Vertex AI
- OpenAI API
- Anthropic API
- Mistral
- Self-hosted models on your GPUs
Data handling
Built for prompts that must not leak
No prompt storage by default
Only metadata is recorded unless you turn logging on for a route.
PII redaction
Emails, phone numbers and card numbers masked before a request leaves the gateway.
Your provider credentials
Read from your secrets manager or workload identity, never from app code.
Region pinning
Routes can be held to US or EU regions for data residency.
FAQ
KubeOn AI, answered
Shares owners, budgets and sign-in with KubeOn.
No. KubeOn AI works read-only from billing exports and provider metrics. The gateway is optional and adds per-request metering, routing, fallback and enforcement.
See your AI spend per team, model and request.
A 30-minute walkthrough on your own Azure OpenAI, Bedrock or Vertex AI usage.