The problem
Standing up a managed cluster takes an hour. Running production workloads on it for two years, surviving upgrades, passing an audit, and handling a 3am incident is a different job.
We build and migrate Kubernetes platforms for healthcare teams running clinical workloads, defense subcontractors with CUI in their clusters, and B2B SaaS under SOC 2. The pattern repeats: the cluster is fine, the workloads run, and the platform is one upgrade away from a bad weekend.
Control mapping
Kubernetes gets assessed twice in regulated environments, and the questions differ. HIPAA asks which safeguard a control satisfies. FedRAMP and DoD assessors ask which NIST SP 800-53 family it belongs to. The same cluster decisions answer both, which is why the mapping is worth drawing once instead of rebuilding per framework.
- §164.312(a)(1) · AC, SCNamespace isolation with default-deny NetworkPolicy, so tenancy is enforced by the network rather than by convention.
- §164.312(d) · IAWorkload identity federated to the cloud IAM plane. No static service-account keys anywhere in the cluster.
- §164.312(c)(1) · CM, SIAdmission control that rejects unsigned images, with provenance verified at admission rather than only at build.
- §164.312(b) · AUCluster audit policy shipped to an immutable store, covering control-plane calls and not just workload logs.
- §164.312(e)(1) · SCPrivate control-plane endpoints and private service connectivity, so data paths never traverse a public address.
- §164.308(a)(1)(ii)(D) · SI, RAContinuous posture and drift detection against the baseline, with the rejection itself logged as evidence.
Both regimes have been in scope on real clusters. Under HIPAA, a cross-organization GCP landing zone in Terraform with GKE Autopilot in the destination and Cloud SQL reached over Private Service Connect. Under FedRAMP, a migration from AKS to EKS spanning AWS GovCloud and GCC High, where the partition boundary and identity federation were the hard part, not the workloads.
Migration approach
We build the new platform alongside the existing one, replay production traffic against it, and cut over when the response surface is provably identical. Not before.
What you get
The baseline, shipped as code in every engagement.
- Namespace isolation. One namespace per team or workload domain, with quotas, limits, and network policy enforced at the boundary.
- Default-deny network policy. Explicit allow rules per dependency, on Calico, Cilium, or platform-native policy.
- Pod Security Standards. Restricted profile cluster-wide, with documented exceptions where a workload requires one.
- Image admission control. Kyverno or OPA Gatekeeper rejecting unsigned images and untrusted registries.
- Secrets from cloud KMS. External Secrets Operator against Secrets Manager, Secret Manager, or Key Vault. No long-lived secrets in cluster YAML.
- GitOps as the only path to production. Argo CD or Flux. Every cluster state derives from a commit.
- Observability. Prometheus, OpenTelemetry, and Loki or equivalent, with SLO-aligned alerting.
Who runs what
Worth stating plainly, because it is the question every platform engagement leaves unanswered.
We build it
- Architecture baseline, decided with you and shipped as code
- Cluster, networking, admission control, secrets, and GitOps wiring
- Observability stack and the alerts that matter
- Runbooks, on-call training, and a walkthrough of every decision
You run it
- Your team owns the platform after handoff. That is the point.
- The application layer stays yours throughout
- Upgrades and incident response follow the runbooks we hand over
- An operational retainer is available if you would rather we stayed on it
Engagement models
Most clients start with an audit, then move to a migration or a build. Some come straight to migration when a compliance date is already fixed.
Platform audit
2 weeks · fixed fee
Your existing platform assessed against security, reliability, and compliance baselines. A written report with a prioritised remediation roadmap.
Migration
8 weeks · fixed fee
Production workloads moved with zero customer-visible downtime. Parallel environments, traffic replay, documented rollback, 30 days of post-cutover support.
Platform build
10 weeks · fixed fee
A production platform from scratch. Architecture baseline, GitOps, admission control, observability, runbooks, and on-call training.
Every engagement is fixed fee with written acceptance criteria, so you know the deliverables and the cost before work starts. Thirty minutes on your current cluster is enough to scope which one fits.