A coding agent can write correct Terraform faster than I can read it. That stopped being the interesting part a while ago. The interesting part is what a Terraform merge starts: an apply that can drop a database, rewrite an IAM binding or repoint a DNS record. The undo isn't free either. Revert an API enablement, an immutable field or a bucket, and the "rollback" can be a second outage.
There are two easy positions on this, and I think both are wrong. "Let the agent write it and apply it" bets production on a model's judgement, and "keep agents away from Terraform" throws away a real speed-up. I stopped trying to trust the agent and built a system that doesn't need to. The agent does the typing, the plan shows what the change will do, and a person owns the merge and the approval.
When I started as the first infrastructure hire, engineers ran Terraform from their laptops with administrator credentials. Today about 36 reviewed changes a week go through a pipeline I built (Building a better execution system for Terraform covers how), and I drive coding agents against that Terraform every day. For the first stretch I was the whole team, so nobody had time to watch every plan. Every safety property had to live in something durable.
Two constraints hold for everything below. The agent is an interactive tool with me in the loop, so nothing runs unattended, and it never merges, approves or applies. And CI is deterministic: it checks diffs, and it never invokes the agent.
The path from agent to apply#
The agent shows up three times on that path. It writes the change, it reads the plan, and after the apply it checks what happened. Every step that changes infrastructure starts with a person or with CI, and none of the CI gates know or care whether a model wrote the diff. That property is the whole design.
Stacks are the blast radius#
A plan or an apply can only touch what's in the state file it runs against. So the way you split state is your blast-radius design, whether you meant it to be or not. A stack called misc is a blast radius nobody chose.
My landing zone is a fork of Google's Cloud Foundation Fabric FAST, which layers stages from the organisation down to workloads. A simplified view of the stages:
0-bootstrap, which can break everything.1-resman, for folders, projects and organisation resource management.2-networkingand2-security(IAM and entitlements).3-gke, for clusters and platform workloads.
Shared code lives in modules/, outside the stages.
Small stacks matter more with an agent than without one. An agent works from what's in front of it, and it will confidently make a change whose blast radius it can't see. If one stack owns the network, a database and DNS, a single agent-written plan can put all three in the same apply. Small stacks cap the worst case, and they keep each plan short enough that a person reads it properly instead of skimming it.
The Terraform guidelines I wrote are for humans and agents alike. They treat 150 to 200 managed resources as a warning sign, not a limit, because the real signal is cohesion: 180 tightly related resources can be a healthy stack, and 40 unrelated ones are already wrong. A new stack is right when the resources have a different lifecycle or owner, or when destroying one set shouldn't put another at risk. Modules get the matching rule, which is to extract one only when a second caller exists.
Dependencies flow one way, from networking to DNS to databases to applications. Stacks read each other's outputs through terraform_remote_state, never each other's resources, and I treat those outputs as the stack's public contract. A stack exports service_account_email or dns_zone_name, not every resource ID it holds.
The plan is the contract#
terraform plan shows what will be added, changed, destroyed and replaced without touching anything. It's the one artifact that the agent and I both reason about.
The agent-specific trap is terraform validate. It checks that the configuration parses and is internally consistent. It says nothing about real infrastructure. An agent that reports "validate passed" as reassurance has answered a question nobody asked.
The plan has limits too, and agents rarely mention them. It reads real state, so it can show drift you didn't cause, and that drift needs an explanation before the change goes anywhere. A clean plan also means only that the change was calculable. Quotas, organisation policies and API-side validation only fail at apply.
I read the plan against the stated intent, never on its own. "3 to add, 1 to change, 0 to destroy" is only reassuring if that's what I meant. And when an agent says a change is "just a refactor", the plan has to agree: a pure refactor with moved blocks plans 0 to add, 0 to change, 0 to destroy, and any replacement means an address is wrong.
Some parts of a diff earn a closer read, whoever wrote them:
- Anything destroyed or replaced.
- IAM. The
_bindingand_policyresources are authoritative and remove every member they don't list, while_memberonly adds one. - State-holding resources like databases, buckets and disks, where I check
deletion_protectionandforce_destroyby hand. - Immutable fields that force a replacement.
- API enablement with
disable_on_destroyunset. - Any change that reaches into a second project.
Gates that don't care who wrote the diff#
The pipeline has its own post, so here's only what matters from the agent's side. None of these properties depend on the agent behaving well.
Every changed stack gets a plan on the pull request, as one comment per stack that's patched in place on each push. So the effect of an agent's change is on the PR before anyone reviews it. The plan job also runs a label policy check.
Each stack applies as its own least-privilege identity. Workflows authenticate through Workload Identity Federation with no stored key, and each stack has a separate service account that holds only the roles for its own project. A mistake in one stack's apply stays inside that stack's blast radius.
The registry fails closed. A small registry file maps each stack to its identity and its approval gate. A gated stack with no approvers fails to load, a registry that doesn't parse aborts the run, and a stack missing from the registry can't authenticate at all. That last one catches agents. A new stack without its registry entry dies at google-github-actions/auth, and the error reads like a cloud outage, so the code standards tell the agent it's a missing entry and the fix is in the code.
Gated stacks wait for a named approver, and an approval covers one commit. Before the apply, a guard checks that the approved commit is still the current commit for that stack. The guard compares full 40-character SHAs and runs before cloud auth, so a refused run never gets credentials. If the commit is current, the workflow plans again and applies.
Nothing on this list is about AI, and I like that. The agent is bound by the same gates I am, and it reads the same plan comment.
Count the deletes with a tool, not a prompt#
"Don't destroy production resources", written into an agent's prompt, is a suggestion. Models don't reliably follow it, and I won't bet a database on the ones that usually do. A count of deletes has to come from something deterministic.
Terraform's JSON plan format makes that cheap. Each entry in resource_changes has a change.actions list, and the valid values are ["no-op"], ["create"], ["read"], ["update"], ["delete"], ["delete", "create"] and ["create", "delete"]. A plain destroy is ["delete"]. A replace is one of the last two, depending on create_before_destroy. HashiCorp says they shaped it this way on purpose, so a caller can scan the list for delete and catch all three cases where an object gets deleted.
This is the scan, trimmed. I tested it with jq 1.8 against a sample plan with a destroy, both replace orders, an update and a no-op:
terraform plan -out=plan.tfplan
terraform show -json plan.tfplan > plan.json
# "// []" covers a plan with no resource changes at all.
deletes=$(jq -r '(.resource_changes // [])[]
| select(.change.actions | index("delete"))
| "\(.address): \(.change.actions | join(","))"' plan.json)
if [ -n "$deletes" ]; then
echo "::error::plan destroys or replaces resources"
echo "$deletes"
exit 1
fiThe Terraform skill my agents load carries a jq recipe like this one. So when an agent reports "two replaces, no destroys", the count came from jq and names every address, instead of coming from a model summarising a long plan. People get the same rule: count the destroys and the replaces, and write the number down. As a failing CI step, the scan stops the job and prints the addresses, whatever the agent concluded about its own change. The prompt makes the agent useful, and the check makes it safe.
How I drive an agent through a change#
The general loop is in Merge means deploy, and the validation plans the agent writes are in Don't let them grade their own homework. These are the parts that are specific to Terraform.
The intent comes first, in one paragraph, with a list of what must not change. For example: "must not replace the DNS zone, touch public DNS, or affect non-prod." That list is the part people skip, and it's the part that catches the agent. If the intent doesn't fit in a paragraph, it's more than one change.
One intent per pull request. I don't mix a refactor with a behaviour change, a provider upgrade with a resource change, or a state move with a configuration change. A mixed plan is one nobody can reason about, so it gets rubber-stamped, and the real failure mode is the plan that hid a destroy in the noise.
The agent reads the source before it writes. Left alone, it writes from what it remembers of the provider, and the plan will happily calculate a change built on a wrong assumption. So it starts with the provider documentation for the exact version pinned in the stack's versions.tf, and with the variables of any module it calls. Then it runs the same pre-commit hooks as CI (fmt, validate, tflint and trivy) instead of a hand-written command list that drifts from CI.
The agent reports on the plan, in a fixed vocabulary. It reads the plan comment from the PR, diffs it against the intent and reports the delta. Its answer is always one of "ready for review", "blocked: plan doesn't match intent" or "needs owner", and never "apply this". The most useful thing it tells me is usually the change the intent didn't mention, which is either drift it just surfaced or scope creep. Either way, it comes out of this PR.
Merged is not applied. Terraform convergence works differently from GitOps, and the differences are easy to miss. A gated stack changes nothing after the merge until a named person approves, so "it's waiting on approval" is a real answer. The pipeline is a push system that fires once, so unlike Argo CD, nothing retries later. If the job was queued, cancelled or paused, nothing happened and nothing will. Even Apply complete! only means the cloud API accepted the calls, while certificates, load balancers, IAM and DNS settle on their own timelines. The Terraform skill puts it in one line: "A run that exists is not an apply that happened, and the run's own conclusion covers every stack in the push, not yours." The proof that a change landed is a follow-up plan of the same stack that shows no changes.
Give the agent the map, and keep the lists off it#
An agent can't trace a dependency it can't see, and Terraform dependencies cross stage and repo boundaries all the time. A networking output feeds a cluster configuration, which feeds an application, and an agent left to guess will guess.
So I encode the map. A root context file (CLAUDE.md, or whatever your agent reads) carries the cross-stage dependency graph: where the DNS zones and shared networking live, which stage owns which service accounts, and how the remote-state wiring hangs together. Per-stage files carry the local detail. It's also the architecture documentation a small team needed anyway and never wrote.
Writing the Terraform skill taught me that some facts make a map worse. Structure that changes slowly belongs on it. Lists that change every week don't, because a stale list looks exactly like a current one. So the skill refuses to list which stacks are gated, and tells the agent where the truth lives and how to ask. Against a simplified registry file like the one in the pipeline post, that looks like this:
# Which stacks wait for a named approver? Ask the registry, not the docs.
yq '.stacks | to_entries | map(select(.value.approval.required)) | map(.key)' stacks.yamlThe skill also has a short section that lists the parts of itself most likely to go stale. That tells the agent which of its own instructions to check before it trusts them.
Footguns an agent will walk into#
prevent_destroy only protects a block that still exists. Terraform doesn't record lifecycle rules in state, except for create_before_destroy (lifecycle docs). Delete the block, and Terraform destroys the resource anyway, because no configuration is left for the rule to live in. Removing blocks is exactly what an agent does when it cleans up "unused" resources, so the backstop is the delete count. To stop managing a resource without destroying it, use a removed block (Terraform 1.7 or later), with destroy = false in a nested lifecycle block:
removed {
from = google_storage_bucket.legacy
lifecycle {
destroy = false # forget it from state, leave the bucket standing
}
}for_each, not count, for anything with a name. I learned this properly from a one-line edit to a count-based list of google_service_account resources. Changing account_id forces a new service account, and removing an item from a count list moves every item after it to a new index. This is the shape of that plan, with the names made up:
It looked like a trivial diff until I read the plan. And Google treats a recreated service account as a separate identity, even with the same name, so it inherits none of the old one's roles. for_each keys each instance by name, so removing one key touches only that instance, as long as the key is stable. My post on the Terraform execution system has the rename that keyed projects on filenames and planned 34 destroys.
Import, never recreate. If a resource already exists, an import block brings it under management, where a reviewer sees it in the plan. An agent that finds no DNS zone in its local state will happily declare a new one, so the guidelines forbid creating a duplicate zone, identity, database, bucket or network just because the local stack doesn't manage one. When ownership is unclear, the agent says so instead of guessing.
Make destructive switches loud. I never hide a destroy behind a friendly variable like delete_database = true. The name should make the reviewer flinch:
# This intentionally replaces the temporary development database.
# This must not be used in prod.
allow_database_replacement = trueThe flinch is the safety mechanism, and it matters more when an agent might be the one setting the flag.
A policy check reads the code literally. Our label policy reads default_labels only as an inline literal. Change it to default_labels = local.labels, the kind of tidy-up any reviewer or agent might suggest, and the check denies every labelled resource in the stack while the code looks correct. It fails closed, so it's loud rather than dangerous, and the code standards document the trap so the agent knows before it tidies.
The prevent_destroy and count traps teach the same thing. Terraform's safety behaviour lives in your current configuration, not in state. Anything you delete from the configuration takes its protections with it.
What still bites#
A scan for delete can't tell an intended replace from an accident. Wired as a gate, it trips on every legitimate replacement, so owner sign-off gets noisy. Noise builds the habit of waving replaces through, which is the rubber stamp the scan exists to stop. I'd still rather clear a false positive than miss a destroy, but the noise is real.
The map in CLAUDE.md also goes stale. Nothing checks that what it says about the remote-state graph matches the code, so a cross-stage rewiring that doesn't update the map leaves the agent reasoning from a description that's wrong. A check that diffs the documented graph against the real terraform_remote_state references is on my list and not built. Until it is, the map is only as good as the last person who touched it.
What I'd do differently#
I'd write every list as a query from the start, the way the Terraform skill does now. A list of gated stacks or approvers in a context file is wrong the first time someone edits the registry, and nothing tells the agent.
I'd also put one question into every Terraform review from the first day. It's on the checklist in my guidelines now: "Could an agent safely modify this later without guessing too much?" A stack that passes it is usually easier for people to change too.