Deployment
Install once per cluster, run the hub anywhere
KubeOn has two parts. A read-only agent runs as a CronJob in each cluster. A hub joins its snapshots with your billing data, on KubeOn Cloud, in your cloud account or in your data center.
helm repo add kubeon https://charts.kubeon.iohelm repo updatehelm upgrade --install kubeon-agent kubeon/kubeon-agent \ --namespace kubeon --create-namespace \ -f kubeon-agent.yamlThe agent
A CronJob with read-only RBAC and nothing else
The agent reads the Kubernetes API, writes a snapshot and exits. It never reads Secrets or ConfigMaps and has no write verbs.
| Kind | CronJob, hourly (0 * * * *), concurrencyPolicy Forbid |
|---|---|
| Run time | Typically under a minute; hard timeout 15 minutes |
| Resources | Requests 100m CPU and 256Mi; limits 500m and 1Gi |
| Security context | Non-root (uid 1000), read-only root filesystem, all capabilities dropped |
| Network | Outbound HTTPS only. No Service, no Ingress, no inbound port |
| Storage | No PersistentVolume. One gzipped snapshot per run |
apiVersion: rbac.authorization.k8s.io/v1kind: ClusterRolemetadata: name: kubeon-agent-readerrules: - apiGroups: [""] resources: [nodes, pods, namespaces, persistentvolumes, persistentvolumeclaims, services, resourcequotas, limitranges] verbs: [get, list, watch] - apiGroups: [apps] resources: [deployments, statefulsets, daemonsets, replicasets] verbs: [get, list, watch] - apiGroups: [batch] resources: [jobs, cronjobs] verbs: [get, list, watch] - apiGroups: [storage.k8s.io] resources: [storageclasses] verbs: [get, list, watch] - apiGroups: [metrics.k8s.io] resources: [nodes, pods] verbs: [get, list]Install
Helm, Flux, Argo CD or Terraform
One chart, the same values everywhere. GitOps users commit a HelmRelease or an Application and let the controller roll it out to every cluster, with the cluster name substituted per cluster.
helm repo add kubeon https://charts.kubeon.iohelm repo updatehelm upgrade --install kubeon-agent kubeon/kubeon-agent \ --namespace kubeon --create-namespace \ -f kubeon-agent.yamlThe hub
Three ways to run the hub
KubeOn Cloud
We run, patch and back up the hub. You install agents and grant read-only billing access.
- Agents send snapshots over HTTPS to an ingest endpoint in your region
- Billing read through a read-only role, managed identity or service account
- Hosting in the US or the EU
- Upgrades without downtime
What you do
- 1. Create a workspace and choose a region.
- 2. Connect billing with the read-only setup for your cloud.
- 3. Install the agent with the token from your workspace.
Self-hosted in your cloud
The hub runs in your account. Billing data and snapshots never leave it.
- AWS: a Terraform module for ECS on Fargate with an internal Application Load Balancer
- RDS PostgreSQL 16, Multi-AZ in production; ElastiCache Redis 7.1 with TLS
- A KMS key with rotation for the database, cache, buckets, secrets and logs
- Azure, Google Cloud or any Kubernetes: the Helm chart in hub mode
# env/prod.tfvarsenvironment = "prod"region = "us-east-1"vpc_id = "vpc-0abc1234def567890"private_subnet_ids = ["subnet-0aaa1111bbb22222c", "subnet-0ddd3333eee44444f"]alb_internal = truealb_internal_allowed_cidr_blocks = ["10.20.0.0/16"] # your VPN or corporate rangesalb_certificate_arn = "arn:aws:acm:us-east-1:444455556666:certificate/EXAMPLE"dns_zone_id = "Z0123456789EXAMPLE"dns_hostname = "kubeon.example.com"cur_bucket_name = "acme-cur-exports"athena_database = "cur_database"athena_table = "cur2"athena_workgroup = "primary"cur_assume_role_arn = "arn:aws:iam::111122223333:role/KubeOnBillingAccessRole"cur_external_id = "<the ExternalId you passed to CloudFormation>"agent_account_ids = ["111122223333"]agent_role_name_pattern = "*-kubeon-agent"hub_image_registry = "ghcr.io/kubeon"image_tag = "1.8.2"hub_desired_count = { api = 2, web = 2, worker = 2 }db_instance_class = "db.r7g.large"redis_node_type = "cache.t4g.medium"admin_email = "finops-admin@example.com"ai_enabled = falseOn-premises and air-gapped
For data centers and networks without internet egress.
- Helm chart in hub mode on OpenShift, Rancher, Tanzu or kubeadm
- API, web, worker and a single scheduler; autoscaling from 2 to 6 API replicas
- PodDisruptionBudget, NetworkPolicies, ServiceMonitor and External Secrets support
- Images and charts mirrored to your private registry
kubectl create namespace kubeon-hubkubectl -n kubeon-hub create secret generic kubeon-hub-secrets \ --from-literal=DATABASE_URL="postgresql://kubeon:<password>@postgres.internal:5432/kubeon" \ --from-literal=JWT_SECRET="$(openssl rand -base64 48)"helm upgrade --install kubeon-hub kubeon/kubeon \ --namespace kubeon-hub \ --set mode=hub \ --set secrets.existingSecret=kubeon-hub-secrets \ --set ingress.enabled=true \ --set ingress.host=kubeon.example.com \ --set autoscaling.api.maxReplicas=6| Hub option | KubeOn Cloud | Self-hosted in your cloud | On-premises and air-gapped |
|---|---|---|---|
| Who runs the hub | KubeOn | You, in your cloud | You, in your data center |
| Where billing data is stored | KubeOn Cloud region | Your account | Your network |
| Snapshot destination | HTTPS ingest | Your bucket | S3-compatible store or HTTPS |
| Encryption keys | KubeOn-managed | Your KMS key | Your KMS or HSM |
| Internet egress required | Included | For billing APIs | Not included |
| Upgrades | Automatic | Terraform or Helm, your schedule | Helm, your schedule |
Data flow
From a pod to a line on a team's bill
Times are UTC and configurable. Snapshots run every hour; billing and allocation run daily.
KubeOn hub
KubeOn Cloud or self-hosted in your account
- 1Ingest snapshots
- 2Normalize labels
- 3Allocate
- 4Reconcile
- 5Recommend
The agent lists nodes, namespaces, workloads, pods, volumes and services through its read-only ClusterRole, 500 objects at a time.
It reads current usage from metrics-server and, if configured, the last full hour of history from Prometheus.
It writes one gzipped snapshot with a schema version and row counts per kind, encrypted with your KMS key or sent over HTTPS.
The hub ingests new snapshots, rejects any whose row counts do not match, and retries failures up to 3 times.
Daily, the hub reads billing from your export. The last 3 days are re-read so late line items are not missed.
Costs are allocated for every cluster-day with both usage and billing, and each cluster is reconciled to its bill.
Change proposals are regenerated from the latest savings. Reviewed decisions are kept.
Budgets, alert rules and anomalies are evaluated and notifications delivered.
Verify
Check the first run
Trigger a run by hand, read its log, and open Diagnostics in KubeOn to see the cluster's heartbeat.
kubectl -n kubeon get cronjob,jobkubectl -n kubeon create job --from=cronjob/kubeon-agent kubeon-agent-first-runkubectl -n kubeon logs job/kubeon-agent-first-run# uploaded snapshot clusters/account=111122223333/cluster=prod-use1-a/...Upgrades
Pin, upgrade, roll back
Chart versions are pinned. Upgrade on your schedule, and roll back with Helm or by reverting the version in Git.
helm repo updatehelm upgrade kubeon-agent kubeon/kubeon-agent \ --namespace kubeon --version 1.8.2 \ -f kubeon-agent.yaml# roll back to the previous releasehelm rollback kubeon-agent --namespace kubeonSee your own clusters in KubeOn.
A 30-minute walkthrough on your billing data, with an engineer who has run Kubernetes cost programs.