An agent that can write Terraform can also convince itself the change worked. That's the failure I kept seeing when I measured our coding agents on infrastructure work. They didn't break production. They wrote confident reports that a change had landed, citing evidence that couldn't tell a landed change from a stale one.
In Merge means deploy I said a structured validation contract was a direction I was working towards. This is what it became. Our toolkit's guidance on validating an infrastructure change was a 71-line stub full of TODOs, so agents improvised. Against nine real, merged infrastructure PRs, one unaided run in nine produced a usable check plan, and none traced whether the thing running was the thing that merged. With the workflow below, nine in nine did.
Then I measured whether those plans could catch anything. 92% of the pass criteria the agents wrote still read as satisfied when the change was deliberately broken. This post is about that number, and about why the agent that writes a plan can't be the one that judges it.
The false pass the agents kept writing#
The classic GitOps false pass is "Argo CD says Synced and Healthy". I found applications reporting exactly that whose last sync had finished 30 days earlier. They were healthy. They just weren't running the change anyone was asking about.
The obvious fix is to compare the revision Argo CD reports with the merge commit. It's the check I used to recommend, and on one of our two delivery chains it holds. On the platform repo, a merge reaches the cluster in three hops, and the merge SHA survives to .status.sync.revision. I checked 88 applications on that chain and all 88 matched.
The app repo is seven hops, and the SHA doesn't survive them. Kargo renders manifests into one branch per cluster (the GitOps post has the whole path), and every application on that cluster tracks the same branch. So the revision field shows the branch tip, whoever moved it last. In one case I watched it point at an unrelated application's promotion from a day later.
The field you want is .status.operationState.syncResult.revision, the revision that application's own last sync used. But it's null on multi-source applications, which report a revisions list instead. So the obvious one-liner that checked convergence across the fleet quietly skipped 22 of 250 applications and reported a clean board. Nothing errored. The 22 just weren't on it.
The repaired read treats a missing value as a finding instead of as nothing to report:
# Simplified. Each app and the revision(s) its last sync used.
argocd app list -o json | jq -r '.[] | [ .metadata.name,
( .status.operationState.syncResult
| (.revisions // [.revision])
| map(. // "MISSING") | join(",") ) ] | @tsv'Every unaided agent run assumed Argo's revision equals the merge commit. Empty results fooled them as well. In our setup, an expired SSO session fails in a way that looks like a healthy, empty cluster instead of an error, so "nothing is out of sync" and "I can't see anything" print the same thing. Absent is not empty, and the workflow scores an absent result as a failure.
What a validation plan looks like#
The workflow keeps three words apart, because an agent that uses them interchangeably reports passes it didn't earn:
- Convergence means the automation finished and the system settled.
- Validation means the change works, seen from a caller.
- Verification means live state matches what we declared, and nothing else moved.
A check reports pass, fail or unclear, never "looks fine". Unclear means the check couldn't be scored (a blip, an empty metric, a command that wouldn't run), and an unclear result that survives a retry counts as a failure.
Every plan is a table with one row per claim and seven columns: Claim, Phase, Dur., Reproducible check or human action, Pass when, Who and Result. Dur. is durability. A durable result stays true for the recorded revision, such as what a render produced. A perishable one, such as a live HTTP response, has to be rerun when another session picks the work up. A check the agent can't run still gets a row, with the name of whoever runs it.
Here's a trimmed, illustrative plan for adding a /v2 route to an API gateway, with Result left out because nothing has run yet:
| Claim | Phase | Dur. | Check | Pass when | Who |
|---|---|---|---|---|---|
/v2 reaches the new backend | post-merge | perishable | status of GET /v2/healthz | 200. Today it's 404 | agent |
| A call with no token is refused | post-merge | perishable | status of GET /v2/orders, no token | 401 | agent |
/v1 is untouched | post-merge | perishable | status of GET /v1/healthz | 200, as in the baseline | agent |
| The render adds one route, nothing else | pre-merge | durable | diff of helm template on main and on the branch | one added /v2 rule | executor |
| Merge and promote | gate | human action | human |
In a real plan the check cell holds the exact command, such as curl -s -o /dev/null -w '%{http_code}' "$API/v2/healthz", so anyone can rerun it.
The "Pass when" column gets filled in before any check runs. A number you interpret after the fact is not a check. Criteria are mechanical: a status code, an exact string, a number, an empty diff. If scoring a row needs judgement, the row gets rewritten.
The main functional check goes first, and it has to be sensitive to the change. It must fail against the old state, which is why the first row says what /v2 returns today. A check that already passes before the change can't prove the change worked, and several delivery signals don't add up to one functional check.
Access checks run in pairs: the right caller gets in, and the wrong one is still refused. Only a refusal proves the fence is live, so any new component gets a negative-path row.
Renders get the same suspicion. A successful helm template proves Helm parsed your inputs, not that your value did anything. A key no template reads renders cleanly, changes nothing, and passes review because the command exited zero. So the render row diffs a baseline against a candidate. For a feature, the diff must contain the change. For a refactor meant to preserve output, the empty diff is the pass.
Every claim also names how strong its evidence is:
- The tool said so. CI is green, the plan rendered, Argo says Healthy. This is the weakest.
- The system said so. You read back the intended revision and the settled state from the platform.
- The outside said so. A real caller used the changed thing. This is the strongest.
"CI passed" is never the headline. A run that exists is not an apply that happened. Facts also carry a label for where they came from, SOURCE (declared in Git), LIVE (read from the control plane, with a date) or HISTORICAL, and only LIVE facts can support an operational conclusion.
Before the plan: challenge it, then size it#
Planning starts with an argument against the change: what the agent is assuming, who else depends on the thing, what would show this is the wrong fix, and how it could fail quietly. An unanswered challenge blocks planning.
Two rules outrank everything else, because in both cases the diff doesn't show the damage. For traffic, never add a path and remove one in the same change, because the diff shows a route added, not the route that stopped resolving. For data, say what survives, what's destroyed and what a revert restores, before the merge. "I think it persists" is a stop, not a caveat.
Then the change gets a tier, rated on its blast radius instead of its resource type:
| Tier | Meaning | Gets |
|---|---|---|
| None | No runtime effect | One line saying why. Stop. |
| Cheap-reversible | Undo in minutes, one system, no traffic or data | Read the whole diff, one line on undo cost |
| Routine | Reversible, one system, obvious failure | The loop, kept small |
| High | Traffic, data, identity, shared dependencies, quiet failure | Full loop, the griller when complex, and human approval of the plan |
The cheap-reversible tier has one test. If this is wrong, what does undoing it cost, and does the undo restore behaviour or only the file? If the answer doesn't fit in one sentence, it's a higher tier.
That tier caps rigour on purpose, which is the opposite of how most validation processes work. An hour spent proving a five-minute change is theatre, and theatre teaches people to skip the process on the day it matters.
The catch: 92%#
Plans in that format looked rigorous, so I measured whether they were. I built a scoring harness around a broken world, a version of the change that is silently wrong, and scored each pass criterion the agent wrote against it. A criterion that still reads as satisfied there can't catch the defect.
92% of them still read as satisfied. The plans were the green-check theatre I was trying to remove, with extra steps.
My first fix was the obvious one. I asked the planning agent to review its own plan, and measured that too. It moved the result from 0 of 41 criteria catching the defect to 0 of 46. It added criteria, and none of them caught anything. The agent that wrote a plan is the worst judge of whether the plan catches anything.
So I added an adversary, which the workflow calls the validation-plan griller. It's a separate agent, and it can't edit anything. It gets the plan and builds a concrete broken world first: this route silently resolves nowhere, this secret is the old version, this alert has no receiver. Then it attacks each row with one question. Does this check fail in that world? If not, the row goes back for repair, and the griller attacks the repair too. It stays the same agent across rounds, so it remembers what it already broke, and it closes with a verdict on the plan as written, not the plan the author meant to write.
A typical finding looks like this (illustrative):
Row: Argo app is Synced and Healthy.
Broken world: the rollout finished, but pods still run the previous image.
Attack: the row still passes in that world.
Verdict: repair it. Prove the running digest is the promoted digest.
The repaired row checks artifact identity, which no unaided run ever did:
# Simplified. Fails unless every pod runs the promoted digest.
[ -n "$DIG" ] || { echo "no promoted digest: unclear" >&2; exit 1; }
kubectl get pods -n "$NS" -l "app=$APP" -o json \
| jq -e --arg d "$DIG" '
[.items[].status.containerStatuses[].imageID]
| length > 0 and all(endswith("@" + $d))'The length > 0 matters as much as the comparison. all() over an empty list is true, so without it the check passes on a service with no pods at all.
The adversary was the only thing in the workflow that found the defects.
Roles that tools enforce, not prompts#
The workflow runs as seven subagents with one job each: validation-plan-griller, verification-executor, pr-validation-reviewer, convergence-checker, validation-verification-checker, fix-implementer and pre-deploy-fixer. A human does every merge, apply, sync, promotion, approval and cutover. Controllers can react after a human acts, and the agents watch that reaction without triggering it. Two design calls in the split matter more than any prompt.
Checkers can't edit. Four of the seven carry disallowedTools, so the file-editing tools aren't in their tool set at all. Only the two fixers can edit. A checker that can fix ends up scoring a system it just changed, and the cheapest fix is always to move the expectation, not the infrastructure. The boundary has a regression test: eval case L6-checker-proposes-fix fails if a checker even offers a fix.
---
name: convergence-checker
description: Reports whether a merged change converged on every target. Read-only.
disallowedTools: Write, Edit
---That's trimmed to the line that matters. The real file also carries the brief and the evidence rules.
The executor can't grade. It runs the rows and returns the command, the exit status, the output and a timestamp, and nothing else. The plan's author scores them. The executor runs on the cheapest model because there's no judgement in its job. An agent that both runs and grades can talk a nonzero exit into a pass, and splitting the roles makes that impossible instead of just discouraged.
The post-merge checkers run isolated on purpose. The convergence checker never receives the baseline, and the validation-verification checker returns two separate verdicts, the functional result first. The griller is the exception, since remembering its last attack is its job.
Turning a sentence into an exit code#
The rule "agents never mutate" started as prose, so I tested whether anything enforced it, and the test taught me something general. An agent inherits its operator's access, so whatever I'm allowed to do, it's allowed to do. The guardrail has to live in the workflow, not in the agent's good behaviour.
So I wrote a hook, 207 lines of Python with a 198-line test suite. It runs before every shell command the agent issues. It allows a list of read-only verbs and exits 2 on everything else. Most of the work is in the evasions:
- It strips quoted heredoc bodies as data.
- It treats backticks and
$(as command boundaries. - It resolves simple assignments, so
A=argocd; $A app syncstill gets caught. - It unwraps
sudo,env,timeoutandxargs, and recurses intosh -c. - It fails closed on anything it can't parse, including a missing
python3.
Failing closed takes more care than it sounds, because of how Claude Code reads hook exits. Only exit code 2 blocks a tool call. Exit 1, the usual Unix failure, is a non-blocking error, and so is a hook that can't start at all: the shell exits 127 and the command runs anyway (hook exit codes). An uncaught Python exception or a missing interpreter would both let the command through, so the hook exits 2 when python3 is absent and on anything it can't parse. A missing interpreter and an unparseable command mean the same thing: I can't prove this is read-only.
The restriction costs nothing. Every Argo application already self-heals, so a manual sync is redundant anyway.
A prompt that can't see its blast radius#
Production promotion used to sit behind an interactive permission prompt on kargo promote. The prompt had no way to classify what it guarded, because the command text doesn't reliably reveal the target, and Stage names aren't authoritative anyway. So the prompt asked a human to approve a command whose blast radius it didn't know.
I replaced it with a classifier that runs before the shell and reads a tier label from the live Stage. Exactly dev means a dev promotion. Everything else, including an empty, missing, unknown or unreadable label, is protected. A missing label isn't permission to guess from the Stage's name.
There's an audit reason too. An agent that acts with my access acts as me. So agents don't promote.
When a row fails#
The repair and stop rules exist so an agent can't quietly redefine success when a check fails. Checkers never fix, push, merge or change an expectation, and a fixer only starts after a human approves it. What happens next depends on what failed:
| What failed | What happens |
|---|---|
| Convergence, or an implementation defect | I approve a fixer. It starts cold, prepares a fix PR, and I merge it |
| A planned rollback trigger | The planned revert PR, which I merge |
| The outcome itself is contradicted | Stop and report that the approach can't work |
| A must-not-change row | Stop and report the collateral movement. Never auto-fixed |
| The check is invalid or blocked | Repair the check, the query or the access path |
| Unclear or still pending | Retry within the bounded window |
| The result can't be tied to this change | Report the limit, without claiming a pass |
A revert and every fix go back through the same human gate and the same live checks as the original change. Some findings stop the work before anyone fixes anything: unexpected destruction, a wrong environment, unexplained drift, unclear ownership, or a High-tier claim that can't be checked at all.
A record that outlives the session#
Apply and rollout run on their own clock, so the session that plans a change often isn't the session that confirms it. The plan and its results live in one maintained PR comment, the validation ledger, and the PR carries one label taken from the ledger's Status: line: validation:pending, validation:passed or validation:failed.
A new session reads the ledger first and picks up from the PR's state:
| State | Continue from |
|---|---|
| Approved, not merged | Recheck the diff and the baseline |
| Merged, checks remain | Trust durable results for the recorded revision, rerun perishable ones |
| Merged, no plan recorded | Ask a human for full validation or convergence only. Write the criteria before reading live results |
| Merged long ago, never validated | Verify current identity and declared state. Report missed comparison windows as unattributable |
The rule under that table is the one I'd keep if I could keep only one. Don't reconstruct missing pre-merge evidence, and don't turn current health into proof of an old change. That's why "unattributable" is its own result. It gives an agent an honest answer that isn't a pass.
What broke: I simplified away the part that worked#
Version three of the workflow had a commit called "Simplify". It cut the main skill from 807 lines to 419, and it read much better. It also deleted the adversary loop, which was the one part the evidence supported.
I reversed it in version four, with the data in the commit history. Version four moved detail into the domain skills instead of deleting it. Negative results, like the 0 of 46, went into the repo too. A null result is still a result.
Where it's at#
| Measure | Unaided | With the workflow |
|---|---|---|
| Runs that produced a check plan, on nine real merged PRs | 1 of 9 | 9 of 9 |
| Runs that traced artifact identity | 0 of 3 | 2 of 3 |
| Criteria a script can score | 75% | 90% |
| Convergence quality, blind-graded out of 5 | 2.7 | 3.7 |
The cost was roughly 15,000 to 20,000 extra input tokens a run and no extra wall-clock time. The samples are small, and an org spend limit capped part of the benchmark, so I read these numbers as a direction, not a precise effect size.
The validation skill is the hub of the nine Claude Code skills I built for the team's toolkit. The domain skills answer "what is true right now", and each of the Argo CD, Terraform and alerting skills hands off to the validation skill as soon as what you're looking at belongs to a change someone intends to merge. Around them sit the seven subagents, the hook, and 35 eval cases in four suites. The workflow runs my own platform work every day.
What I'll do differently next time#
Build the broken world before the plan template. I designed the plan format first and measured it later. The harness that breaks things is what told me the truth, so it belongs at the start.
Enforce boundaries in tools from day one. Every rule I kept as prose, I later had to prove wasn't enforced, then enforce. I'd start with the tool grant and the hook and write the prose after.
Measure before simplifying. The "Simplify" commit felt like progress, and the eval said otherwise. Now nothing gets cut from the workflow without a run against the suites.