Production Kubernetes, built for regulated workloads

We architect, migrate, and operate production Kubernetes platforms for healthcare, defense, federal, and B2B SaaS teams. Namespace isolation, network policies, pod security, signed-artifact admission, and zero-downtime migrations across EKS, GKE, AKS, and OKE, including their regulated variants.

The problem

Standing up a managed cluster takes an hour. Running production workloads on it for two years, surviving upgrades, passing an audit, and handling a 3am incident is a different job.

We build and migrate Kubernetes platforms for healthcare teams running clinical workloads, defense subcontractors with CUI in their clusters, and B2B SaaS under SOC 2. The pattern repeats: the cluster is fine, the workloads run, and the platform is one upgrade away from a bad weekend.

Control mapping

Kubernetes gets assessed twice in regulated environments, and the questions differ. HIPAA asks which safeguard a control satisfies. FedRAMP and DoD assessors ask which NIST SP 800-53 family it belongs to. The same cluster decisions answer both, which is why the mapping is worth drawing once instead of rebuilding per framework.

  • §164.312(a)(1) · AC, SCNamespace isolation with default-deny NetworkPolicy, so tenancy is enforced by the network rather than by convention.
  • §164.312(d) · IAWorkload identity federated to the cloud IAM plane. No static service-account keys anywhere in the cluster.
  • §164.312(c)(1) · CM, SIAdmission control that rejects unsigned images, with provenance verified at admission rather than only at build.
  • §164.312(b) · AUCluster audit policy shipped to an immutable store, covering control-plane calls and not just workload logs.
  • §164.312(e)(1) · SCPrivate control-plane endpoints and private service connectivity, so data paths never traverse a public address.
  • §164.308(a)(1)(ii)(D) · SI, RAContinuous posture and drift detection against the baseline, with the rejection itself logged as evidence.

Both regimes have been in scope on real clusters. Under HIPAA, a cross-organization GCP landing zone in Terraform with GKE Autopilot in the destination and Cloud SQL reached over Private Service Connect. Under FedRAMP, a migration from AKS to EKS spanning AWS GovCloud and GCC High, where the partition boundary and identity federation were the hard part, not the workloads.

Migration approach

We build the new platform alongside the existing one, replay production traffic against it, and cut over when the response surface is provably identical. Not before.

Parallel-environment migration: assess, build the new platform alongside the existing one, replay production traffic, run workloads in shadow, cut over via DNS or load balancer, then decommission after burn-in.
Cutover is a DNS or load balancer change, not a forklift event. Rollback is one command.

What you get

The baseline, shipped as code in every engagement.

  • Namespace isolation. One namespace per team or workload domain, with quotas, limits, and network policy enforced at the boundary.
  • Default-deny network policy. Explicit allow rules per dependency, on Calico, Cilium, or platform-native policy.
  • Pod Security Standards. Restricted profile cluster-wide, with documented exceptions where a workload requires one.
  • Image admission control. Kyverno or OPA Gatekeeper rejecting unsigned images and untrusted registries.
  • Secrets from cloud KMS. External Secrets Operator against Secrets Manager, Secret Manager, or Key Vault. No long-lived secrets in cluster YAML.
  • GitOps as the only path to production. Argo CD or Flux. Every cluster state derives from a commit.
  • Observability. Prometheus, OpenTelemetry, and Loki or equivalent, with SLO-aligned alerting.

Who runs what

Worth stating plainly, because it is the question every platform engagement leaves unanswered.

We build it

  • Architecture baseline, decided with you and shipped as code
  • Cluster, networking, admission control, secrets, and GitOps wiring
  • Observability stack and the alerts that matter
  • Runbooks, on-call training, and a walkthrough of every decision

You run it

  • Your team owns the platform after handoff. That is the point.
  • The application layer stays yours throughout
  • Upgrades and incident response follow the runbooks we hand over
  • An operational retainer is available if you would rather we stayed on it

Engagement models

Most clients start with an audit, then move to a migration or a build. Some come straight to migration when a compliance date is already fixed.

Platform audit

2 weeks · fixed fee

Your existing platform assessed against security, reliability, and compliance baselines. A written report with a prioritised remediation roadmap.

Migration

8 weeks · fixed fee

Production workloads moved with zero customer-visible downtime. Parallel environments, traffic replay, documented rollback, 30 days of post-cutover support.

Platform build

10 weeks · fixed fee

A production platform from scratch. Architecture baseline, GitOps, admission control, observability, runbooks, and on-call training.

Every engagement is fixed fee with written acceptance criteria, so you know the deliverables and the cost before work starts. Thirty minutes on your current cluster is enough to scope which one fits.

Kubernetes Migration · Second Engagement
“This was our second engagement with Lucas and he led a full Kubernetes migration for us. He scoped the work clearly, built out the manifests, set up proper namespace isolation, network policies, and resource limits, and migrated our workloads without any production downtime. He also took the time to document the architecture and walk our internal team through the new setup so we weren't left in the dark. Solid Kubernetes expertise, clean execution, and zero drama.”
Ryan S.
CTO, AI SaaS
Verified review
Lucas Jones, Founder and Principal Engineer at Stonebridge Tech Solutions

Lucas Jones, founder

Principal Engineer · Stonebridge Tech Solutions
Cloud Infrastructure · Data Platforms · Software Engineering

Six years building cloud infrastructure and CI/CD pipelines in regulated environments. HIPAA, FedRAMP, and SOC 2 engagement work for healthcare and defense engineering teams across AWS, GCP, Azure, and OCI. The senior engineer who scopes your engagement leads it and stays accountable through handoff. Senior engineers only, all US citizens. No offshore delivery.

The same patterns documented in Field Notes are what get applied during real client engagements.

Or, book directly

Pick a time. Skip the back-and-forth.

30-minute discovery call. We walk your current cluster, talk about the engagement that fits, and you get a written proposal within 48 hours.