I designed and built most of the platform I run: GitOps delivery across five GKE clusters that all 38 production services ship through, a family of three Helm charts, a Terraform landing zone and an API gateway. Engineers ship through it every day.
The trouble showed up during the migration. To move those 38 services across, I wrote a tool that diffed the rendered config against what was live, and for the length of the migration that diff was our only definition of correct. Correct meant "matches what's running now", bugs included. It got us through with zero downtime (that story is One render, one verdict), and it left me unable to answer a basic question afterwards: what is the platform supposed to do?
Nothing written down answers that. The charts, the Terraform and the operators describe how the platform works today. The rest lives in a few people's heads, mostly mine, and in what we remember from past incidents. So every change to the shared charts waits for my review, because the context to judge it sits with me.
What a missing spec costs#
The review queue is only the part I feel every day:
| You can't... | Because... |
|---|---|
| Test it | There's nothing to test against |
| Replace any part of it | "Same behaviour" has no definition, so every swap is a risk nobody can size |
| Hand it over | The real spec is in a few people's heads |
| Say no | There's no stated boundary, so every request is arguable |
| Sell it internally | There's no product surface, only a list of tools we run |
| Let a machine work on it | An agent has no way to know if it succeeded |
The migration was an example of the second row. The diff tool could tell us that something had changed, but never whether the change was allowed.
Components#
A component is the smallest piece of the platform you can install, upgrade and break on its own. Envoy Gateway is one. So are External Secrets Operator, Argo CD, Kargo, Alloy, our service chart and our GKE Terraform module. The check is whether I can upgrade it by itself, and whether it can break when nothing around it changed.
I write a component down like a function signature: what goes in, what it needs, what comes out and what it promises. The definition is a small file next to the code it describes, and the code stays where it is. This is a trimmed example for Envoy Gateway:
# components/envoy-gateway/component.yaml
name: envoy-gateway
capability: gateway
implementation:
kind: helm
paths: [gateway/controller, gateway/shared]
inputs: # what a platform engineer configures
- name: listenerHostnames
required: true
- name: tlsSecretRef
required: true
- name: rateLimitBackend
default: redis
dependencies:
runtime: [gateway-api-crds] # must exist before it can run
test: [redis] # only needed to test rate limiting
outputs: # what other components may rely on
- name: sharedGatewayName
- name: gatewayClassName
contract: ./contract.yamlEach field is a commitment. An input that isn't listed isn't supported, and an output that isn't listed can change without warning. The test dependencies are what the local test setup installs, and the length of that list turns out to be the most useful number in the whole exercise.
Capabilities#
Nobody outside the platform team cares about components. An app team doesn't want Argo CD. It wants merged code to reach production. A capability is that kind of promise, delivered by one or more components, and you can tell you have one if you can state it without naming a tool. "Delivery" passes. "Argo CD" doesn't.
This is our map:
| Capability | The promise | Components today |
|---|---|---|
| Runtime | Workloads run | GKE module, cluster policies, service chart |
| Gateway | HTTP exposure and traffic policy | Envoy Gateway, shared Gateway config, Redis |
| Identity | A workload can prove who it is | Workload Identity config, RBAC |
| Secrets | Secrets get delivered and rotated | External Secrets Operator, Secret Manager config |
| Telemetry | Metrics, logs and traces reach the backend | Alloy, Grafana Cloud config |
| Alerting | Defined conditions produce a page | Alert generation, rule definitions |
| Delivery | Merged commits reach production | Argo CD, Kargo, repo conventions |
| Cloud landing | A correctly configured project | Project factory, GKE module, VPC, DNS |
The first draft had Observability as one capability, and that was a mistake. A broken Alloy config and a bad PromQL threshold fail in different ways and have different owners, so they became Telemetry and Alerting. Identity got its own row because it's the seam between GCP and Kubernetes, where most cross-boundary failures happen and where local testing is hardest. Folding it into Runtime would have hidden the part that most needs attention.
Why the lines don't match#
A capability follows what a consumer notices, and a component follows what breaks and upgrades together. I cut them on different lines on purpose.
Force one line and you lose either way. Cut only by what consumers notice and the units get huge: if Delivery is one installable thing, any delivery test needs Argo, Kargo, a Git server and a cluster, so in practice nobody runs one. Cut only by what upgrades together and the promises come out shaped like tools, and nobody can build anything on "Argo CD is healthy".
Contracts#
Capabilities and components both get contracts. A contract is a short list of guarantees written so you can check them. Next to the guarantees it lists invariants (things that stay true all the time, not just after a deploy), the things we explicitly don't promise, and what the local tests can't reach.
This is the contract template filled in for the Gateway capability, shortened. The status values are illustrative, not our real status:
# capabilities/gateway/contract.yaml
capability: gateway
contractVersion: v1
guarantees:
- id: hostname-routing
statement: >
A Service declaring a hostname MUST receive HTTP requests
for that hostname at its configured port.
status: met # met | partial | unmet
target: keep # keep | commit | drop
- id: unmatched-not-routed
statement: >
Requests for an unconfigured hostname MUST NOT reach any Service.
status: met
target: keep
- id: zero-downtime-config-change
statement: >
A Gateway configuration change MUST NOT drop in-flight requests.
status: unmet
target: commit
invariants:
- id: controller-ready
statement: The Gateway controller MUST reach Ready and stay Ready.
nonGuarantees:
- Ordering of requests across backends.
- Any specific Envoy version or configuration format.
notValidatedLocally:
- Cloud load balancer
- Edge security policies
- Public DNS resolution
- Real certificate issuanceA component contract is narrower. The Gateway capability promises that a Service with a hostname receives traffic on it, while the Envoy Gateway component only promises that an HTTPRoute on the shared Gateway becomes Accepted. You need both, because a chart change can break a Service's selector or port name while the HTTPRoute still reports Accepted, and traffic returns 503. Neither a Gateway-only test nor a Helm render check catches that. Only a test of the capability does.
To decide what goes in, I ask whether anyone would file a ticket if it stopped being true. Nobody files one about a replica count, so a replica count is a description and stays out. Everybody files one when requests to their hostname stop arriving.
A contract also has to survive swapping the tool underneath. If the Gateway contract mentions Envoy anywhere, it's wrong, because the point is to let you replace the implementation and prove the replacement was safe. That's why the contract has its own version number. Contract v1 should survive an Envoy Gateway minor upgrade, and an upgrade that forces a contract change is worth stopping to look at.
The rule that matters most for testing is that a component can only promise what it can deliver alone. "Access logs reach Grafana" looks like it belongs in the Envoy Gateway contract, but Envoy can't make that true by itself. Put it there and you can never test Envoy without the whole observability stack, which is how local test setups become unusable. The guarantee goes in the Telemetry contract, with the gateway's log format as the handoff: one access log line per request, with the status, path, duration, request ID and route name.
Whoever gets paged for a guarantee writes it, in MUST and MUST NOT. When someone else writes it, you get a summary of code they just read. Each guarantee has a stable ID, like hostname-routing, so tests and reports can point at it after the wording changes.
One list, with the gap computed#
The obvious way to record what we do today, what we want and the difference is three documents. They rot, and the gap document rots first. So there's one list, the contract itself, and each guarantee carries the two fields you saw above: status for what we do today and target for what we intend. The gap is computed whenever someone looks, so it can't go stale. Putting the two fields side by side also exposes a case nobody thinks to check:
| We intend to promise it | We don't | |
|---|---|---|
| We do it today | Working. Keep it tested. | Accidental promise |
| We don't | Gap. This is the roadmap. | Not a promise. Delete the line. |
The bottom-left box is the roadmap, in terms app teams care about rather than tool upgrades. zero-downtime-config-change above is one of those. The top-right box is where the trouble is. It holds behaviour teams built on that we never meant to support, the thing you remove in a clean-up and then spend a week apologising for. Writing it down turns that future incident into a scheduled deprecation.
To fill in the list I start from the last year of incidents, support tickets and chat questions, not the code. "Is X supposed to work?" points at a guarantee nobody declared. "Why did Y break?" points at one we didn't know we had. Then I ask two app teams what they assume, and I read the code last, only to check that what I wrote is true.
Testing across the GCP boundary#
Half the capabilities span GCP and Kubernetes. Secrets, for example, is a Secret Manager secret, an IAM binding, a Workload Identity link and an ExternalSecret. GCP's official emulators only cover data services like Pub/Sub and Spanner. There's nothing local for GKE, IAM, load balancing or VPC.
But our bugs are almost never in GCP. They're in our wiring: a binding on the wrong service account, a template that renders the wrong key, a pod that doesn't restart when its secret rotates. All of that can be tested locally if you fake GCP's side.
So the handoff is a small seam file. Terraform produces a named set of outputs, the cluster consumes them, and both sides are tested against the same file. This is a simplified example:
# capabilities/secrets/seam.yaml
seam: gcp-to-cluster
producer:
system: terraform
component: project-factory
outputs:
- id: eso-service-account-email
format: "*.iam.gserviceaccount.com"
- id: secret-prefix
consumer:
system: gitops
component: external-secrets-operator
requires: [eso-service-account-email, secret-prefix]On the Terraform side, terraform test proves the outputs get produced, and Conftest over terraform plan -json checks that the IAM bindings have the right shape. The format field exists because terraform test can't check formats: its mocked values are random eight-character strings, so it can't tell you whether a service account email looks like one. On the cluster side, the local setup provides a stub of the seam, either a pre-created secret or External Secrets Operator pointed at a fake provider. That tests our wiring, RBAC, templating and restart-on-change, and deliberately not Secret Manager itself.
The rule is to test the part you wrote against a declared stub of the part you didn't. Some capabilities get much further with it than others:
| Capability | Provable locally | Needs real GCP |
|---|---|---|
| Gateway | Routing, timeouts, rate limits, failure modes | Load balancer, edge security, DNS, certificates |
| Secrets | External Secrets wiring, RBAC, templating, restart-on-change | Workload Identity binding, Secret Manager IAM, real rotation |
| Identity | The shape of bindings, through policy checks | Whether a binding grants the access it should |
| Cloud landing | Terraform logic, policy compliance | Project creation, org policy, real IAM |
Everything in the right-hand column goes into the contract's notValidatedLocally list and gets a scheduled test against a sandbox project instead of a test on every pull request. That test destroys what it creates, runs in one project with a hard budget alert, and applies twice to check the second plan is empty. Without it, "not validated locally" quietly turns into "never validated anywhere".
The test output says out loud what it didn't test. This run is illustrative, not our real results:
Gateway capability contract
PASS hostname-routing
PASS unmatched-not-routed
PART https-redirect - one listener of two tested
GAP zero-downtime-config-change - committed, no test
NOT VALIDATED LOCALLY: cloud load balancer, edge security, DNS, cert issuanceWithout that last line the run looks green and hides its blind spots, which is worse than having no test. There's no coverage percentage either, because a percentage becomes a target.
The local loop#
The loop is Kind plus Chainsaw. I picked Kind over k3d because Kind runs upstream Kubernetes and k3d runs k3s. For the Gateway, the cluster holds only what the tests need:
The prober pod is there because port-forwarding is the main source of flakiness in this kind of test. Traffic starts inside the cluster instead, and a Host header lets a test reach a service "on a domain" with no DNS.
Chainsaw tests are declarative YAML, and their assertions poll until a timeout, which gets rid of sleep, the biggest source of flaky infrastructure tests. Their catch blocks dump resource status and controller logs when a step fails, and I write those first. "Expected 200, got 000" leaves a human stuck and an agent worse off, while the same line followed by the route status and the proxy logs gives either of them something to act on. This is a sketch of the first Gateway test, and it hasn't run yet:
apiVersion: chainsaw.kyverno.io/v1alpha1
kind: Test
metadata:
name: hostname-routing
spec:
description: A Service declaring a hostname MUST receive HTTP requests for it.
timeouts:
exec: 30s
steps:
- name: route is accepted and serving
try:
- apply:
file: ../../fixtures/basic-route/
- assert: # polls until true or the timeout
resource:
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: echo
status:
~.(parents):
(conditions[?type == 'Accepted']):
- status: 'True'
- script:
content: |
kubectl exec -n test prober -- \
curl -sS -o /dev/null -w '%{http_code}' \
--retry 5 --retry-connrefused --retry-delay 2 \
-H 'Host: echo.test' \
http://gateway.gateway-system/
check:
($stdout): '200'
catch:
- describe:
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
- events: {}
- podLogs:
namespace: gateway-system
selector: app=gateway-proxy
tail: 100The loop has to respect GitOps. Nothing on our platform gets helm installed, because Argo syncs it, so the test renders the worktree with the same Helm values and overlays Argo uses and then applies the result. What the test sees is exactly what Git produces.
Component tests also skip delivery on purpose. Sync waves, health checks, drift correction and promotion order belong to the Delivery contract, and they get tested once, with Argo and Kargo running. Every other component assumes delivery works. Routing every component test through Git, Kargo and Argo is how a three-week spike turns into a six-month framework.
What I rejected#
I won't use Helm charts or CRDs to define components, because both are implementations. A component boundary is a YAML file in Git that points at whatever the implementation already is, with no controller behind it. The boundaries will be wrong at first, and fixing one means deleting some YAML, which is the cheapest way to be wrong.
Crossplane composes upward, turning many resources into one object, and my first job runs the other way: drawing lines through a platform that already exists. crossplane render proves the YAML is right, which isn't the same as proving a route returns 200. It also adds a control plane, providers and a reconciliation loop to every test setup, which pushes up the one number I track (more on that below). It does fit Secrets and Cloud landing later, because their pieces get created together and drift together, so the composition boundary and the component boundary match. With a contract in place, that move keeps the contract and the tests and changes only the implementation.
I also put off an internal developer portal. A portal copies the patterns you already have, and ours aren't written down yet, so Backstage would be a catalogue with nothing to catalogue.
For the first two components the test runner is a Makefile, because the worst outcome here is six months spent on a platform testing tool instead of the platform. make test starts Kind, installs the dependencies and the candidate, applies the fixtures and runs Chainsaw.
And I won't reorganise the repo first. Restructure before you write down what things do and you can't tell whether you broke anything. Writing the contract often shows the boundary was fine, and it only felt like a mess because nobody had written it down.
How I'll know it failed#
Tests written and guarantees covered both measure effort, and they go up whether or not the work helps. I count how many things must be running to test one guarantee instead. This is the plan for the Gateway:
| To test that... | You need running |
|---|---|
| An HTTPRoute works | Envoy Gateway alone |
| Rate limiting works | Envoy Gateway and Redis |
| A test app is reachable on a domain | Runtime and Gateway |
| Access logs reach Grafana | Gateway and Telemetry, then Grafana Cloud |
If Envoy Gateway needs six other controllers before it routes one request, the boundary isn't real, however many tests pass. That's still a useful finding, because it names a specific coupling that no architecture meeting would surface.
There's a time budget too. A component suite that takes more than five minutes, or flakes more than once in a hundred runs, gets ignored by everyone, me included. So the Gateway setup comes in two sizes: a minimal one without Redis or Loki, budgeted at about 45 seconds, and a full one at about three minutes.
Where it's at#
The capability map is drafted and the build order is set, but no contract is written yet and the loop isn't built. The loop comes first because it's what makes real contracts testable. The Gateway goes first among the capabilities because traffic in and traffic out is the cleanest boundary we have. Secrets follows, because if the seam pattern works there, it works everywhere, and if it doesn't, I want to know early.
The acceptance test for the whole idea is one real incident from the last year. The Gateway loop has to catch it, failing on the old version and passing on the fix. If it can't catch a bug we actually had, nothing else here counts.
What I'll do differently next time#
I'd write the Delivery contract before the chart family. I built the golden path first and described it later, which is why the migration's diff ended up doing the job of a spec.
I'd go broad and shallow first. The contracts are the asset and the tests only keep them honest, so eight capabilities with rough contracts beat one with exhaustive tests and seven that nobody can see. A ten-guarantee contract takes a couple of hours with the right person, so all eight could be drafted in a week.
And I'd treat every incident as a spec bug first, asking which promise was missing, wrong or unwritten before asking why it broke. Nobody writes a wrong guarantee. They just never write the twelfth one, and the twelfth one takes production down.