Skip to content
NewKubeOn AI: a managed LLM gateway with cost per 1M tokens for every model

Changelog

What shipped, release by release

Agent and hub releases share a version number. Subscribe to get release notes by email.

  1. v1.9KubeOn AI

    KubeOn AI: the LLM gateway

    • One OpenAI and Anthropic compatible endpoint for Azure OpenAI, Bedrock, Vertex AI, OpenAI and Anthropic
    • Routes with PTU-first spillover, fallback on 429, 5xx and timeouts, and retries with backoff
    • Team keys with budgets, hard stops and token and request rate limits
    • Exact-match response cache, PII redaction and region pinning
  2. v1.8Savings

    Node sizing on CPU and memory, with family changes

    • Node groups are sized on both CPU and memory from 14 days of peak requests
    • New finding: change node family when one resource is mostly idle, with the saving capped at the real price gap
    • Node savings are scaled to a full month for every cluster
    • Possible orphaned volumes with a check-first cleanup script, shown as an upper bound
  3. v1.7.2KubeOn AI

    KubeOn AI: report explorer and AI budgets

    • Drill-down reports by product, zone, team and tenant across Azure and AWS
    • Alerts at an amount or a share of budget, including a deployment's share of its tenant
    • Month-close reload with unmatched meters flagged
  4. v1.7PlatformAllocation

    Sorting and filtering on every table

    • Every table sorts and filters, server-side for paged tables
    • Allocation efficiency sorts by cost-weighted efficiency
    • Savings scope excludes the empty namespace; sub-cent and unbilled estimates are labelled
  5. v1.6.1KubeOn AI

    KubeOn AI: PTU sizing, reservations and billing correction

    • Right-sizing from per-minute telemetry, with break-even volume for moves off PTU
    • Reservation coverage by Global, Data Zone and Regional pool, after right-sizing
    • Approval-gated purchase and renewal policies
    • Hourly PTU charges moved out of pay-as-you-go in place, with reservation terms set
  6. v1.6SavingsGovernance

    A change proposal for every saving

    • Orphaned volume detector and proposals for every savings bucket, including node and volume changes
    • Bundled proposals with per-item evidence for the long tail
    • Released file system volumes reported as orphans, priced as an even share of their file system
    • Sortable proposals list with coverage against the savings total
  7. v1.5AllocationPlatform

    Assets with capacity, and provider badges

    • Assets aggregated like other reports, with node capacity, billing detail and drill-in
    • Time-weighted capacity: average vCPU and memory, and nodes running at once
    • Cloud provider detection with badges, and grouping by provider on Assets and Cloud costs
  8. v1.4AssistantGovernance

    Ask KubeOn and reviewed change proposals

    • Ask KubeOn: questions answered from read-only tools, with every figure traced
    • Explain buttons on budgets, anomalies and proposals
    • Proposal summaries written from the evidence, never lowering the computed risk
    • Usage ledger for assistant calls with tokens, cost and latency
  9. v1.3GovernancePlatform

    Governance and per-cluster diagnostics

    • Budgets with run-rate forecasts, alert rules, anomaly detection and notification channels
    • Slack, Microsoft Teams, email and webhook delivery with per-channel status
    • Per-cluster diagnostics with a 48-hour heartbeat and 10 checks
  10. v1.2AllocationPlatform

    Faster cloud cost views

    • Daily cloud cost rollups behind Overview and Cloud costs
    • Faster pod allocation queries and pre-warmed 7, 30 and 90-day views
    • Network cost by transfer class, with namespace flows from VPC flow logs