KubeOn AI pricing
Editions for every stage of your AI program
Start read-only with tokenomics, add the gateway when you want routing and control, and move to Enterprise for a self-hosted gateway and approval-gated reservations.
Insights
Read-only tokenomics for teams that want visibility first.
Get a quote
Quoted on providers and volume, annually
- Cost per 1M tokens by model, team and cloud
- PayGo, PTU hourly and reservation buckets
- Owner attribution from tags
- Budgets and alerts
- Report explorer and exports
- All providers, read-only
Gateway
For platform teams that want routing, metering and control in one endpoint.
Get a quote
Quoted on providers and volume, annually
- Everything in Insights
- Managed LLM gateway, US or EU
- PTU-first routing, spillover and fallback
- Team keys with budgets and rate limits
- Response cache and PII redaction
- PTU sizing and reservation coverage
- Per-request cost and OpenTelemetry export
Enterprise
For regulated or high-volume AI programs across many providers.
Get a quote
Quoted on providers and volume, annually
- Everything in Gateway
- Self-hosted gateway in your network
- Approval-gated reservation automation
- Custom routes, guardrails and retention
- SSO, custom roles and audit export
- Named technical account manager
- 24/7 support
Model usage stays on your provider bills at your negotiated rates.
Compare
What each KubeOn AI edition includes
| Feature | Insights | Gateway | Enterprise |
|---|---|---|---|
| Tokenomics | |||
| Cost per 1M tokens, all-in | Included | Included | Included |
| Input, output and cached tokens | Included | Included | Included |
| Pricing buckets and billing correction | Included | Included | Included |
| Model-mix savings | Included | Included | Included |
| Per-request tokens and cost | Not included | Included | Included |
| Gateway | |||
| OpenAI and Anthropic compatible endpoint | Not included | Included | Included |
| Routing strategies and fallback | Not included | Included | Included |
| PTU-first spillover | Not included | Included | Included |
| Team keys, budgets and rate limits | Not included | Included | Included |
| Response cache | Not included | Included | Included |
| PII redaction and region pinning | Not included | Included | Included |
| Self-hosted gateway | Not included | Not included | Included |
| Provisioned capacity | |||
| Right-sizing from per-minute telemetry | Not included | Included | Included |
| Reservation coverage by pool | Not included | Included | Included |
| Approval-gated purchase and renewal | Not included | Not included | Included |
| Spend control | |||
| Budgets and alerts | Included | Included | Included |
| Hard stops at budget | Not included | Included | Included |
| Report explorer and dashboards | Included | Included | Included |
| Security and support | |||
| SSO with SAML or OIDC | Not included | Included | Included |
| Custom roles and audit export | Not included | Not included | Included |
| Support hours | Business hours | Extended | 24/7 |
How we count
The terms in a KubeOn AI quote
Tracked AI spend
Model and provisioned capacity spend that KubeOn AI reads from your billing exports, across every provider you connect.
Gateway requests
Requests routed through the KubeOn AI gateway, counted per month. Cache hits are counted but never call a provider.
Provider
A connected source such as an Azure tenant, an AWS organization, a Google Cloud billing account or an API organization.
User
A person who signs in. Apps that call the gateway with keys are not users.
FAQ
KubeOn AI pricing questions
Billing
Your providers do, at your negotiated rates. KubeOn AI routes requests through your own Azure, AWS, Google Cloud and API accounts, so token charges stay on those bills.
Gateway
It is part of the Enterprise edition. It runs on your infrastructure, so you pay your provider for the compute it uses.
Tell us which providers you use. We will send a quote.
Providers, monthly AI spend and whether you want the gateway. No billing data needed.