Skip to content
NewKubeOn AI: a managed LLM gateway with cost per 1M tokens for every model

KubeOn AI pricing

Editions for every stage of your AI program

Start read-only with tokenomics, add the gateway when you want routing and control, and move to Enterprise for a self-hosted gateway and approval-gated reservations.

Insights

Read-only tokenomics for teams that want visibility first.

Get a quote

Quoted on providers and volume, annually

Get a quote
  • Cost per 1M tokens by model, team and cloud
  • PayGo, PTU hourly and reservation buckets
  • Owner attribution from tags
  • Budgets and alerts
  • Report explorer and exports
  • All providers, read-only
Most complete

Gateway

For platform teams that want routing, metering and control in one endpoint.

Get a quote

Quoted on providers and volume, annually

Get a quote
  • Everything in Insights
  • Managed LLM gateway, US or EU
  • PTU-first routing, spillover and fallback
  • Team keys with budgets and rate limits
  • Response cache and PII redaction
  • PTU sizing and reservation coverage
  • Per-request cost and OpenTelemetry export

Enterprise

For regulated or high-volume AI programs across many providers.

Get a quote

Quoted on providers and volume, annually

Talk to sales
  • Everything in Gateway
  • Self-hosted gateway in your network
  • Approval-gated reservation automation
  • Custom routes, guardrails and retention
  • SSO, custom roles and audit export
  • Named technical account manager
  • 24/7 support

Model usage stays on your provider bills at your negotiated rates.

Compare

What each KubeOn AI edition includes

Features included in each KubeOn edition
FeatureInsightsGatewayEnterprise
Tokenomics
Cost per 1M tokens, all-inIncludedIncludedIncluded
Input, output and cached tokensIncludedIncludedIncluded
Pricing buckets and billing correctionIncludedIncludedIncluded
Model-mix savingsIncludedIncludedIncluded
Per-request tokens and costNot includedIncludedIncluded
Gateway
OpenAI and Anthropic compatible endpointNot includedIncludedIncluded
Routing strategies and fallbackNot includedIncludedIncluded
PTU-first spilloverNot includedIncludedIncluded
Team keys, budgets and rate limitsNot includedIncludedIncluded
Response cacheNot includedIncludedIncluded
PII redaction and region pinningNot includedIncludedIncluded
Self-hosted gatewayNot includedNot includedIncluded
Provisioned capacity
Right-sizing from per-minute telemetryNot includedIncludedIncluded
Reservation coverage by poolNot includedIncludedIncluded
Approval-gated purchase and renewalNot includedNot includedIncluded
Spend control
Budgets and alertsIncludedIncludedIncluded
Hard stops at budgetNot includedIncludedIncluded
Report explorer and dashboardsIncludedIncludedIncluded
Security and support
SSO with SAML or OIDCNot includedIncludedIncluded
Custom roles and audit exportNot includedNot includedIncluded
Support hoursBusiness hoursExtended24/7

How we count

The terms in a KubeOn AI quote

Tracked AI spend

Model and provisioned capacity spend that KubeOn AI reads from your billing exports, across every provider you connect.

Gateway requests

Requests routed through the KubeOn AI gateway, counted per month. Cache hits are counted but never call a provider.

Provider

A connected source such as an Azure tenant, an AWS organization, a Google Cloud billing account or an API organization.

User

A person who signs in. Apps that call the gateway with keys are not users.

FAQ

KubeOn AI pricing questions

Billing

Your providers do, at your negotiated rates. KubeOn AI routes requests through your own Azure, AWS, Google Cloud and API accounts, so token charges stay on those bills.

Gateway

It is part of the Enterprise edition. It runs on your infrastructure, so you pay your provider for the compute it uses.

Tell us which providers you use. We will send a quote.

Providers, monthly AI spend and whether you want the gateway. No billing data needed.