Joel Freeman
All writing
04Writing
Article
Published
Reading time
15 min read

Merge Means Deploy: Running AI Agents Against Production Infrastructure

The agent wrote the change and opened a clean PR. In infrastructure that's where the risk starts, because the merge is the deploy. This is what a merge sets off, how an agent proves a change is live, and why I still do the merging.

  • ai
  • platform engineering
  • devops
  • gitops
  • agents

An AI agent will write you a chunk of Terraform or a Helm values change, open a tidy PR with a sensible description, and tell you it's ready. That part is good now, good enough that it's tempting to treat it as the whole job.

In application work, merge means done, and the deploy is a later step that's tested, staged and reversible. In infrastructure, the merge is the deploy. Where I work, about 36 reviewed Terraform changes a week go through pull requests, and the apply starts when the PR merges. Every service's config reaches its cluster through Argo CD. So the risk starts at the agent's PR.

I use Claude Code and Codex every day against real Terraform, Argo CD, Kubernetes, DNS and IAM on Google Cloud. As the first infrastructure hire, I had to decide what an agent may do to production, because nobody had written it down. This post is the general loop I settled on, and most of it is about what happens after the PR. Running Terraform safely with agents has the Terraform specifics, and the post on validation plans has the planning half. Everything here is an interactive agent that I drive in a terminal, and I'm present at every merge and apply.

A merge is a trigger#

A merge doesn't deploy anything itself. It starts two different machines, and they behave differently.

Terraform is push. The merge fires an apply workflow, which applies an ungated stack straight away and holds a production stack for a named approver. A push-triggered job fires once, so if it was queued, cancelled or paused, nothing happened and nothing will. Even Apply complete! only means the cloud API accepted the calls. A Google-managed certificate can then take up to 60 minutes to provision, and an IAM change typically takes two minutes, sometimes seven or longer.

Argo CD is pull. It keeps comparing Git with the cluster and syncing until they match, but it isn't instant either. Without a webhook it polls every 120 seconds, plus up to 60 seconds of jitter, and only then does it sync and start the rollout. Where Kargo sits in front, the merge lands the first stage only. Each promotion after that is a separate change with its own convergence and validation, and promoting is a human's call, like merging.

A merge to main starts two chains. Terraform pushes once: the apply job runs for the commit, a gated stack waits for a named approver, Apply complete means the API accepted the calls, certificates, IAM and DNS settle later, and a follow-up plan shows no changes. Argo CD pulls and keeps reconciling: it notices the commit by polling every two to three minutes, syncs, rolls out until pods are Ready, the resources settle, and a Kargo promotion to the next stage is a new change. An agent that checks right after the merge reads the old system as green.

None of those steps waits for the agent's session, and each one can leave the old system looking healthy while it runs.

Why the app-work loop breaks#

Agents do well on application code because ground truth is cheap and local, rollback is a revert, and the whole system usually sits in one repo the agent can read. Infrastructure takes all of that away.

Ground truth needs the real world. "Did it work?" means terraform plan against real remote state, an Argo CD diff against a live cluster, or dig against real DNS. There's no fast green on a laptop.

A revert isn't a rollback either. Reverting the commit doesn't bring back a deleted PersistentVolumeClaim, a changed CRD, a recreated load balancer or a rotated secret, and on the Terraform side a revert can be a second outage.

And the system spans repos the agent can't see. I hit this one myself. An agent working in the cluster repo needed a subnet and a shared service account. Both are created in the landing-zone repo and handed over as remote-state outputs. That repo wasn't in the agent's worktree, so the agent did the reasonable thing with what it could see, and it declared both in the cluster repo.

The landing-zone repo, outside the agent's worktree, creates the subnet and the shared service account and passes them to the cluster repo as remote-state outputs. The agent, working only in the cluster repo, could not see them and declared a new subnet and a new shared service account there. The plan was clean, and merging it would have created a second, divergent copy of each.

The plan was clean, because in that state the resources really didn't exist, and the PR read as tidy and self-contained. Merged, it would have stood up a second service account and subnet next to the ones the landing zone already owned, and the next apply over there would have fought it. The agent wasn't careless. It reasoned correctly about half a system as if it were the whole thing.

Give the agent the map#

The direct fix is a CLAUDE.md that the agent reads before it changes anything. It lists the cross-repo dependencies that the agent can't discover by itself. This is a simplified version:

Cross-repo dependencies: read before you change anything

  • The landing-zone repo creates the VPC and subnets, the shared GKE service accounts and the DNS zones. It exports them as remote-state outputs. Do not declare them here.
  • This repo consumes those outputs. If you need a subnet or a service account, read it from the landing-zone remote state.
  • IAM bindings for workloads live in the landing-zone repo, not next to the app. Changing a workload's identity is a two-repo change.

It doubles as the architecture note that new engineers needed anyway, and it sets the merge order. If the subnet doesn't exist yet, the landing-zone change merges, applies and gets checked first, and the cluster change waits. The workflow's rule is to validate foundations before you merge dependent work. The map holds structure only. Lists that change often stay off it, for reasons the Terraform post covers.

The map is still a patch. An agent can't trace Terraform dependencies across a repo boundary unless someone pastes the context in, so I proposed merging our two Terraform repos, with a done-when line that an agent can trace landing-zone DNS to the GKE cluster in one worktree. The proposal kept the GitOps repo separate, because every commit to it invalidates Argo CD's manifest cache.

Before merge: shape the deploy#

The loop is the same for a Terraform stack, an Argo CD app, a DNS record or an IAM binding, and I run it myself. I write the intent as a paragraph first: the outcome, the target environment, and what must not change. For anything bigger than one resource, I decide the PR boundaries before the agent writes a line. In GitOps every PR is a deploy, so a PR boundary is a deploy boundary. An agent told to "sort out the networking" will either pile everything into one diff that nobody can review or spray a dozen small PRs, which is a dozen deploys. Choosing the boundaries is the orchestration. If the agent chooses them, I end up reviewing a deploy I didn't design.

Then the agent makes one scoped change and shows me what it will do: the Terraform plan, the rendered manifests or the Argo CD diff. I read that before the code, because it's what will happen. The review has three possible answers:

  • It did what I meant. We carry on.
  • It errored, or the result doesn't match the intent, and the output goes straight back. This is where an agent earns its place. It reads its own plan, finds its own mistake and fixes it faster than I would.
  • It changed something the intent never mentioned, and we stop. That resource is either drift the agent just surfaced or a second problem riding along, and either way it comes out of this PR and becomes its own ticket. The most dangerous PR does what you asked plus one thing you didn't notice.

Two stop rules outrank that review, one for traffic and one for data, and the validation post has both.

Deletes the diff doesn't show#

Some of the most destructive changes look small in review. Our services are placed on clusters by a placement config that an Argo CD ApplicationSet fans out into Applications. Removing a cluster from that list doesn't pause the app there. By default the ApplicationSet controller deletes the Applications it no longer generates, and their resources go with them unless the ApplicationSet sets preserveResourcesOnDeletion. So commenting out one placement entry is a cascade delete, and the diff is one line.

Terraform has its own versions of this, such as an authoritative IAM binding that removes every member it doesn't list. There the agent doesn't summarise the plan. It lists every delete and replace from the plan's JSON, and the Terraform post shows how.

Hand-checked isn't checked#

I inherited a GitOps template that was meant to generate a baseline set of apps for every cluster from one ApplicationSet. It had only ever been checked by hand, and on the page it looked right. Run against a real Argo CD, it generated zero apps, because a generator that matches nothing reports success, not an error. I fixed three bugs in it and built a CI check that starts a real Argo CD and checks that the apps generate and that their permissions and secrets resolve. It caught all three, and it gives the agent the same ground truth it gives me.

After merge: is my commit live?#

After the merge, the agent runs read-only commands while I watch, and it reports evidence. The order matters, because each check means nothing until the one before it passes.

Did the trigger fire for this commit?#

For Terraform, the pass is that the apply job for the merge commit succeeded, its log ends Apply complete!, and the resource counts match the plan's. A green run somewhere doesn't pass, because a run's conclusion covers every stack in the push, not just yours. For a gated stack, "it's waiting on approval" is a real answer, and any metric window starts when the apply finished, which can be hours after the merge.

For Argo CD, the agent reads the app's .spec.syncPolicy before it waits on anything. An automated block means the app rolls out on its own. Without one, a human has to sync, and that step belongs in the plan instead of turning into a mystery timeout.

Is my revision the live one?#

This is the false pass that worries me most. An app that was Synced and Healthy before the merge is still Synced and Healthy a second after it, on the old revision. An agent that checks straight away reads the old system and certifies it green, which is worse than no check. argocd app wait --sync --health doesn't help, because it has no flag for a revision and returns at once on an app that was already healthy.

So the revision comes first. This is a simplified version of the check for a single-source app that tracks the branch I merged into:

bash
#!/usr/bin/env bash
# Simplified example: one single-source Argo CD app that tracks the branch I merged into.
app="$1" merged="$2"

for _ in $(seq 60); do  # about 15 minutes
  status="$(argocd app get "$app" -o json)"
  live="$(jq -r '.status.sync.revision' <<<"$status")"
  git fetch --quiet origin

  # Revision first: an app that was already Synced and Healthy stays that way on the old commit.
  if git merge-base --is-ancestor "$merged" "$live" 2>/dev/null &&
    jq -e '.status.sync.status == "Synced" and .status.health.status == "Healthy"' \
      <<<"$status" >/dev/null; then
    echo "$merged is live on $app (synced at $live)"
    exit 0
  fi
  sleep 15
done

echo "$merged is not live on $app after 15 minutes" >&2
exit 1

It asks whether my commit is an ancestor of the live revision, not whether the two are equal. In a busy repo someone else merges after me, and a strict match would turn my live change into a permanent false negative. A multi-source app has no single value in that field, so the script times out instead of passing. This check only works where the merge commit reaches Argo CD intact. On our Kargo chain it doesn't, and the validation post explains what to read there instead.

Did the platform settle?#

Healthy in Argo CD describes the manifests, not the behaviour. Argo CD can report Synced and Healthy while the Gateway reports CertificateNotReady, because the manifests applied perfectly and the change still isn't live. When the two disagree, the resource wins. So "finished" has a definition:

  • metadata.generation equals status.observedGeneration.
  • The rollout is complete and the pods are Ready.
  • The Certificate is Ready and the Gateway listeners are Programmed.
  • Jobs are Complete.
  • No sync wave is still pending. An app can report Healthy while a later wave hasn't started, because resources that don't exist yet can't be unhealthy.

On the cloud side the agent reads the finished state. The certificate is ACTIVE, the record answers from a public resolver with dig, and the policy troubleshooter says the IAM binding is in effect. It never swaps a timer for the finished state. A certificate still PROVISIONING isn't a failure if its DNS record came in the same change, because the hour of provisioning starts only once DNS and the load balancer have propagated.

For Terraform, the last check is a follow-up plan. With -detailed-exitcode, terraform plan exits 0 when there are no changes, 2 when there are and 1 on an error, so an agent can't read a non-empty plan as a pass. A non-empty one means drift or a provider mismatch. It goes into the report, and the agent doesn't chase it in the same session.

The merge that doesn't deploy#

Sometimes a merge doesn't deploy at all, and everything still looks fine.

Our rollback is to promote the previous Freight again in Kargo. Since Kargo 1.11, a manual promotion of Freight that isn't the current auto-promotion candidate records a hold. Auto-promotion from that origin to that Stage then stops until someone promotes the current candidate, or turns auto-promotion off and on again. So after a rollback, the next merge builds and its Freight appears, but the Stage doesn't move. Argo CD stays Synced and Healthy on the old Freight, and an agent that reads only health reports success for a change that never reached the Stage. That's the strongest reason the revision check comes before the health check.

Close on evidence#

Apply, sync and rollout run on their own clocks, so the session that writes a change and the session that checks it converged can be days apart. The validation record lives on the PR for that reason, as one comment the workflow keeps updating, and the validation post shows how a later session resumes from it. Each PR also carries a validation:pending, validation:passed or validation:failed label. When the workflow loads, it checks once for merged PRs that are still pending, which finds the merges nobody came back to.

Each meaningful change closes on evidence tied to the merge commit. A closing note reads something like this (simplified):

Terraform: the apply for <sha> succeeded and the counts match the plan. The follow-up plan shows no changes. Argo CD: <sha> is live, Synced and Healthy, and the pods are Ready. Gap: no smoke test exists for this path, so runtime behaviour is unproven.

The honest "unproven, and here's the gap" is worth more than a confident tick. The effort matches the risk, though, so a docs change gets no report at all.

The human keeps the merge#

People keep every merge, apply, promotion and cutover, even where the apply runs automatically after an approved merge. The merge is the deploy, and an agent that acts with my access acts as me.

So the agent never recommends "merge this". It reports that the PR is ready for review, that it's blocked because the plan doesn't match the intent, that it's blocked because the change is risky and needs its owner, or that it needs validation. A good handoff is a PR that carries its intent, its change surface, the machine evidence, a risk rating and the known gaps. I review a case, not a diff.

I'd like more of this orchestrated, with an agent that sequences the changes and checks each plan, and that's where I'm taking it. But on a system like this, a human at the merge is part of the design.

What I'd do differently#

Check the check on every chain. The first version of this post recommended comparing Argo CD's revision field with the merge commit. That holds on one of our two delivery chains and is wrong on the other, as the validation post explains. I'd checked it on the chain I was looking at, which is the agent's cross-repo mistake, made by me. The script above says which apps it covers, and it times out on the rest.

Put dependent stages in one repo before writing the map. The context map works, but it's a patch over a boundary the agent can't see across. I'd rather remove the boundary first.

Make "did it deploy at all?" the first post-merge check from day one. A queued job, a pending approval and a Kargo hold each produce a merge that didn't deploy, and every one of them looks healthy.

04More writing
7 posts

Keep reading