Joel Freeman
All writing
04Writing
Article
Published
Reading time
17 min read

Designing an execution system for Terraform on GitHub Actions

How to take a Terraform change from pull request to apply with no long-lived keys: a plan on every pull request, approved production applies, a guard against stale commits and a least-privilege identity for each stack.

  • terraform
  • gcp
  • github actions
  • security
  • platform

Running terraform apply is the easy part. The hard part is the execution system around it: the machinery that decides which stacks run, which identity each one runs as, which commit it applies, and who has to say yes first.

GitHub Actions gives you the parts for that, and a few of them behave differently from how they look. This post is the design I'd build, and have built and run, for Terraform on GitHub Actions with Google Cloud as the target. The Google parts swap cleanly for AWS or Azure. The GitHub parts don't change.

The design has six properties:

  • Every pull request gets a plan, for every stack it touches, without anyone asking for one.
  • A merge to the default branch is the only way to apply.
  • Production applies wait for a named person, and no runner waits with them.
  • An approval is for one commit. A stale commit is refused before it gets credentials.
  • Each stack applies as its own least-privilege identity, through a short-lived token. No key exists to leak.
  • Every failure is loud. A silent skip is treated as a bug.

The path of one change#

A pull request's changed stacks are detected from the diff, and each stack is planned into its own sticky comment. After merge, workflow A runs one job per changed stack. An ungated stack passes the stale-commit guard, authenticates through Workload Identity Federation as its own service account, and applies. A gated stack plans, opens an approval issue, pings its approvers in Slack and exits, so no runner waits. An approver's comment on the issue starts workflow B, which checks the approver, runs the stale-commit guard, then authenticates, plans again, applies and writes a commit status. A stale commit is refused before any credentials exist.

The unit of work is the stack. A stack is one Terraform root directory with its own state, its own identity and its own gate, so it can be planned and applied without touching any other stack. One pull request can touch one stack or many.

The left column is the short path for an ungated stack, which applies minutes after merge. The right column is the path for a gated stack, such as a production one. It's split across two workflows so that no runner sits waiting while a person decides. The stale-commit guard appears on both paths because both of them apply.

Finding the stacks#

Change detection takes the pull request's diff and turns the changed paths into a list of stack directories. That list becomes a job matrix, so each stack plans in its own job, in parallel, with its own log and its own failure.

The detector has to fail loudly. If the diff command fails, the job fails. A detector that swallows that error reports "no stacks changed", and the pull request goes green with nothing planned, which looks exactly like a change that didn't need a plan.

A directory can't tell you two things about its stack: which identity it runs as, and whether it needs an approver. Put those in a small registry file. This is an illustrative example with made-up names:

yaml
# stacks.yaml (illustrative)
stacks:
  networking/prod:
    service_account: tf-networking-prod@<project>.iam.gserviceaccount.com
    approval:
      required: true
      approvers: [alice, bob]
  cache/dev:
    service_account: tf-cache-dev@<project>.iam.gserviceaccount.com

Two rules keep the registry honest. A gated stack with no approvers fails to load. And a registry that doesn't parse aborts the whole run, instead of skipping the entry it couldn't read. A loud refusal is better than a gate that quietly lets a change through.

The registry is also where a new stack usually breaks first. A stack with no registry entry, or with an identity that has no Workload Identity binding, dies at the google-github-actions/auth step. The error reads like a Google Cloud outage, but it's an authoring mistake, and the fix is in the registry entry or the binding. Say so in the docs for adding a stack, because everyone hits it once.

The plan is the review#

The plan job uses the pull_request trigger, never pull_request_target. That matters because a Terraform plan runs code. Providers run, and an external data source runs whatever program the configuration names.

For a pull_request run from a fork, GitHub downgrades every write permission to read, as long as nobody has turned on the setting that sends write tokens to fork pull requests. That includes id-token, so a fork's pull request can't request the OIDC token it would need to reach Google Cloud. pull_request_target does the opposite. It runs with the base repository's permissions, write token included, even for a fork. Planning a fork's Terraform under it would hand that code your cloud access.

Each stack's plan becomes a sticky comment through tfcmt. A new push patches the existing comment instead of adding another:

bash
# one comment per stack, patched in place on every push
tfcmt -var "target:${STACK}" plan -patch -- terraform plan -no-color

tfcmt finds the comment to patch by the target, so the target has to name the stack and nothing else. Keep a timestamp out of the key. If it's in the key, every push posts a new comment, and the pull request turns into a scroll of stale plans.

The review rules are short:

  • Read the comment, not the green check. Green only means the job ran.
  • Make sure the comment belongs to the head commit. A plan from an earlier push is not evidence.
  • Count the destroys and the replaces, and write the number down.
  • A pure refactor with moved blocks must plan 0 to add, 0 to change, 0 to destroy. Any replacement means an address is wrong.

A plan is good evidence of a narrow thing. It reads the real state, so it can show drift you didn't cause, and then you either explain the drift or stop. A clean plan only means Terraform could calculate the change. Quotas, organisation policies and API-side validation fail at apply, not at plan.

Merged isn't applied, either. A gated stack changes nothing until a person approves it. And nothing retries a push-triggered job on its own. If the apply job was cancelled or never started, nothing happened, and nothing will happen until someone looks. Put the apply's result on the commit, so "did this apply?" has an answer you can find.

Keys are identity#

I've watched a plan comment save a project. A project factory turned one YAML file per project into a Google Cloud project, keyed on the filename. Someone renamed a file, and the plan read that as "create a new project, destroy the old one", including a live state bucket. The plan comment caught it before apply, and the fix was to rename the file back.

In Terraform, a for_each key is identity. If the key comes from something people treat as cosmetic, like a filename or a display name, you've built a trap. Key on a stable ID that nobody has a reason to edit, so a rename is a no-op. Review should be the second line of defence here, not the first.

Who the apply runs as#

The execution system should hold no long-lived credential. On Google Cloud, enforce that with the organisation policy that blocks service-account key creation. It forces the question early, and it means a key can't creep back in later as a quick fix.

I covered how GitHub's OIDC tokens work, with AWS as the example, in Eliminating static AWS credentials in GitHub Actions. The Google side is the same idea with different nouns. GitHub signs a short-lived token for the job. Workload Identity Federation checks that token against a provider in a workload identity pool and swaps it for a federated token. The auth action then uses that to impersonate one service account, and the Google access tokens it gets last an hour by default. Nothing is stored, so there's nothing to rotate and nothing to leak.

The trust has to be narrow at two points. The provider's attribute condition should accept tokens from your repository only, not from anything GitHub signs. And each service account should let only the principals that need it impersonate it. A pool that trusts all of GitHub is a door with a sign on it.

The obvious shortcut is one CI service account with broad rights over every stack. Don't take it. With one account, a bug in a dev stack's apply has the blast radius of production. Instead, each stack applies as its own service account, and that account holds only the roles for its own project. Define those accounts in Terraform too, so a new stack's identity is a reviewed change like anything else.

The apply job reads its row from the registry and authenticates as that account. This is trimmed to the parts that matter:

yaml
permissions:
  id-token: write   # lets the job request GitHub's OIDC token
  contents: read

steps:
  - uses: actions/checkout@v7
  # registry lookup, then the stale-commit guard
  - uses: google-github-actions/auth@v3
    with:
      workload_identity_provider: ${{ vars.WIF_PROVIDER }}
      service_account: ${{ steps.registry.outputs.service_account }}

Per-stack identities have a cost. Every new stack needs an identity, a binding and a registry entry before its first plan works, and when one is missing you get the misleading auth error from earlier. It's a fair price for being able to say exactly what each apply can touch.

Production waits for a named person#

The obvious gate is a job that pauses while it waits for approval. I built that first. It held a runner the whole time, and approvals sometimes took hours, so runners sat idle for hours.

The better design splits the apply into two workflows.

Workflow A runs on merge. For a gated stack, it plans, opens a GitHub issue with the plan in it, pings the named approvers in Slack and exits. Nothing waits.

Workflow B runs when someone comments on that issue. issue_comment fires for every comment on every issue and pull request in the repository, so workflow B first has to decide whether the comment is an approval at all. Then it checks that the commenter is a named approver for that stack, and that the approved commit is still the current commit for the stack. Only after both checks does it authenticate, plan again and apply. When the apply finishes, it writes a commit status, so the commit itself shows whether its apply ran.

The issue is the audit trail. It holds the plan, who approved it, when, and the run that applied it. That's what an auditor asks for, and it costs almost nothing to produce.

A denial closes the issue first and comments second. The close is the control, and the comment is a courtesy. The report steps should fail late. Each one runs on its own, and the job fails at the end and names what didn't land, instead of stopping at the first thing that broke.

GitHub's own environment protection rules are the other option. They don't hold a runner while they wait, and they're less to build. They also give you less: the approval lives in the run, not in an issue with the plan beside it, and which rules you get depends on your GitHub plan. If they fit, use them. The stale-commit guard still applies either way.

These gates are also what make it reasonable to let an agent write Terraform. An agent can open the pull request, but a person still merges it, and a named person still approves a gated apply. I wrote about that side in Running Terraform safely with agents.

You can't test workflow B on a branch#

GitHub runs an issue_comment workflow only if the workflow file exists on the default branch. So workflow B can't run from a pull request, and its first real run happens after it merges.

That changes how you test it. Move the logic out of shell and into a small package with real unit tests. Then run mutation tests against it. A mutation test changes the code on purpose, say by flipping a comparison, and checks that some test fails. If a flipped comparison in the approver check leaves the suite green, the suite isn't testing the approver check. The last step before real use is a rehearsal on a fixture stack that manages nothing important.

Shell in YAML is fine for three lines. Approval logic isn't three lines.

An approval is for one commit#

GitHub lets you re-run a workflow for 30 days, and a re-run uses the same commit SHA as the original run. On an apply workflow, the Re-run button applies an old commit over whatever is newer. Terraform does exactly what you asked. It reverts the newer change and reports success.

Commit A is applied, then commit B is applied and adds a resource. Later, someone clicks Re-run on A's apply, which runs with A's commit SHA. Without a guard, the re-run applies A over B, B's resource is gone, and the run reports success. With the stale-commit guard, the run sees that A is no longer the stack's current commit and refuses before cloud authentication.

So the apply runs behind a stale-commit guard with no bypass:

  • The commit being applied must be the current commit for that stack. If it isn't, the run refuses.
  • The guard accepts only a full 40-character SHA and fails closed on anything else. A 7-character prefix is short enough that someone can compute a commit that matches it.
  • The guard runs before cloud auth, so a refused run never gets credentials.

This is a simplified example of the idea. It compares against the tip of main, which is stricter than a per-stack check, because an unrelated merge also makes the approval stale:

yaml
- name: Refuse a stale commit
  env:
    APPROVED_SHA: ${{ steps.approval.outputs.sha }}
  run: |
    if [[ ! "$APPROVED_SHA" =~ ^[0-9a-f]{40}$ ]]; then
      echo "::error::not a full 40-character SHA: '$APPROVED_SHA'"
      exit 1
    fi
    tip=$(git ls-remote origin refs/heads/main | cut -f1)
    if [[ "$APPROVED_SHA" != "$tip" ]]; then
      echo "::error::approved $APPROVED_SHA, but main is at $tip"
      exit 1
    fi

The SHA goes in through env and not through ${{ }} inside the script, because it comes from an issue, and anything from an issue is untrusted text. If git ls-remote fails, tip is empty, the comparison fails and the run refuses. That's the direction it should fail in.

The tip-of-main version has a cost. On a busy repository, any merge between approval and apply makes the approval stale, and the approver has to say yes again. A per-stack check compares against the last commit that touched the stack's directory instead. It's kinder to approvers and needs more care to get right.

Old runs don't get the new guard#

A re-run uses the workflow file from its original commit, not the current one. So a run from before a guard merges never gets that guard. Until its 30 days pass, deleting the run is the only way to take its Re-run button away.

If you add a safety check to a workflow, clean up the old runs in the same change. They still carry the old rules.

Plan again, or apply the saved plan#

Between the pull request and the apply, two things can change: the code on main, and the real infrastructure. That leaves two choices for what the apply does.

You can save the plan file and apply exactly that. The apply then does only what the reviewer read. But Terraform refuses a saved plan once the state has moved on, and a plan file can hold sensitive values, so it has to be stored like a secret.

Or you can plan again at apply time and apply that. The stale-commit guard makes sure the code is the approved code, and the new plan reflects the infrastructure as it is now. The cost is that the apply may differ from the plan the reviewer read, if something drifted in between.

This design plans again. With the guard pinning the code, a difference between the two plans can only come from drift, and drift is worth knowing about anyway. If you need the stronger promise, save the plan, and make the apply refuse when the new plan's changes don't match the reviewed ones.

Locking and concurrency#

Two runs that apply the same stack at once will corrupt each other's work. There are two layers of defence, and you want both.

The first is the state lock. The gcs backend locks the state while a run writes to it, so a second writer fails instead of racing. Turn on object versioning on the state bucket too, so a bad write can be rolled back. A lock left behind by a killed runner needs terraform force-unlock, and only after you've checked that no run still holds it.

The second is a GitHub concurrency group per stack, so two applies for one stack don't even start together:

yaml
jobs:
  apply:
    strategy:
      matrix:
        stack: ${{ fromJSON(needs.detect.outputs.stacks) }}
    concurrency:
      group: terraform-apply-${{ matrix.stack }}
      cancel-in-progress: false

A concurrency group isn't a queue. It holds at most one running job and one pending job. When a third run arrives, GitHub cancels the pending one and puts the new one in its place.

The cancelled job never starts, so it never reaches its if: always() step, and it never reports. "Queued" doesn't mean "will run eventually". The pending job can just vanish. With the stale-commit guard in place that's usually harmless, because the newer run carries the newer commit. But it's one more reason "merged" and "applied" are different states, and one more reason the commit status matters.

Drift#

The plan on a pull request shows drift only when someone happens to touch that stack. A stack nobody edits can drift for months.

Run a scheduled plan for every stack with -detailed-exitcode. It exits 0 for no changes, 1 for an error and 2 when the plan has changes. On main, a 2 means the real infrastructure no longer matches the code. Open an issue or alert on it. Don't apply it automatically, because drift is sometimes a hand-made fix during an incident, and reverting that fix on a timer is how you get a second incident.

A drift job only reads, so give it a read-only identity per stack. A scheduled job that can write is one bug away from an apply nobody approved.

Failure modes to design for#

Most of this design exists because of a failure that looks like success:

  • A failed diff that reports "no stacks changed". The pull request goes green with nothing planned.
  • A registry that skips the entry it can't parse. The stack runs ungated.
  • A plan comment from an earlier push. The reviewer approves a change they didn't read.
  • A Re-run on an old apply. Terraform reverts the newer change and reports success.
  • An old run from before the guard. It doesn't carry the guard.
  • A pending job replaced in a concurrency group. It vanishes without a report.
  • A clean plan and a failed apply. Quotas and policies only check at apply.

Each fix has the same shape: refuse loudly, before credentials, and leave a record.

What I'd do differently next time#

I'd build the two-workflow gate first. The waiting-job design looked simpler, it wasn't, and I rebuilt it anyway.

I'd key every for_each on a stable ID from day one. I'd rather the design rule out a cosmetic destroy than count on review to catch it.

And I'd start with the tested helper package. I began with shell in YAML and moved it out later. With a test suite from the first day, approval logic is much faster to trust.

04More writing
7 posts

Keep reading