Skip to content
NewKubeOn AI: a managed LLM gateway with cost per 1M tokens for every model

KubeOn AI

Know what every model call costs, and route it well

KubeOn AI measures cost per 1M tokens across every provider, sizes your provisioned throughput, and runs a managed LLM gateway so your teams do not have to build routing, fallback or budgets themselves.

The problem

AI spend grows faster than anyone can explain it

The bill stops at the deployment

Invoices show a cost per deployment or model, not which product, team or customer spent it.

Three ways to pay for one model

Pay-as-you-go, PTU by the hour and reservations land in different meters, and some are recorded in the wrong one.

Provisioned capacity sits idle

PTUs are paid every hour, while traffic still goes to pay-as-you-go because nothing routes it.

Every team builds its own client

Retries, fallback, keys and rate limits are reimplemented in each service, and none of them know the budget.

Two ways to start

Read-only on day one, the gateway when you are ready

Read-only

Connect billing and telemetry. Nothing in your traffic changes.

  • Cost, tokens and requests per deployment
  • PTU sizing and reservation coverage
  • Budgets and alerts by product, team and tenant
Connect a provider

Through the gateway

Change the base URL. KubeOn AI routes and meters every request.

  • Exact tokens and cost per request, key and route
  • PTU-first spillover and fallback across providers
  • Budgets that stop at 100%, rate limits, cache and PII redaction
Gateway quickstart

LLM gateway

Change one URL and stop managing routing

Apps call a route such as chat-default. The gateway sends it to your PTU deployment until it is busy, spills over to pay-as-you-go, falls back to another provider when one throttles, and records the cost of every request against the team's key.

  • OpenAI and Anthropic compatible
  • Budgets and rate limits per key
  • Managed for you, or self-hosted in your cluster
Python
import osfrom openai import OpenAIclient = OpenAI(    base_url="https://gateway.kubeon.io/v1",    api_key=os.environ["KUBEON_AI_KEY"],   # a team key, not a provider key)resp = client.chat.completions.create(    model="chat-default",                  # a route name, not a vendor model    messages=[{"role": "user", "content": "Summarize this ticket in two lines."}],)print(resp.choices[0].message.content)
Explore the gateway

Tokenomics

Cost per 1M tokens that includes your commitments

PTU hours and reservations are allocated to the deployments and models that used them, so pay-as-you-go and provisioned capacity can be compared on the same number, by model, team and cloud.

  • Input, output and cached tokens
  • Past PTU billing corrected in place
  • Model-mix savings from observed cost
Explore tokenomics

Providers

Every model provider your teams use

Hosted APIs, cloud model platforms and models you run yourself, in one cost model.

All integrations
  • Azure OpenAI
  • Azure AI Foundry
  • Amazon Bedrock
  • Amazon SageMaker
  • Google Vertex AI
  • OpenAI API
  • Anthropic API
  • Mistral
  • Self-hosted models on your GPUs

Data handling

Built for prompts that must not leak

No prompt storage by default

Only metadata is recorded unless you turn logging on for a route.

PII redaction

Emails, phone numbers and card numbers masked before a request leaves the gateway.

Your provider credentials

Read from your secrets manager or workload identity, never from app code.

Region pinning

Routes can be held to US or EU regions for data residency.

FAQ

KubeOn AI, answered

Shares owners, budgets and sign-in with KubeOn.

No. KubeOn AI works read-only from billing exports and provider metrics. The gateway is optional and adds per-request metering, routing, fallback and enforcement.

See your AI spend per team, model and request.

A 30-minute walkthrough on your own Azure OpenAI, Bedrock or Vertex AI usage.