Skip to main content

Landing a PR on main — Runbook

Status: active — 2026-09-04. Owner: devops-engineer. Related: core#2965 (queue enabled), core#2967 (the merge_group wiring), core#2842 (the treadmill this replaces), core#2710 (attribution), core#2397 (closing keywords).

STATUS IS NOT RECORDED HERE — read it with the query in §6. repository.mergeQueue(branch:"main") is the only thing that can answer, and it is two seconds to run.

This block used to carry the answer anyway — "as of 2026-09-04 it returns null, the queue is absent" — and by 2026-09-08 that was false: the queue was live, ALLGREEN, with three entries in it (core#3242). Someone following the recorded status would have taken §4 and hit a 405. The number is deleted rather than updated, because updating it just resets the clock on the same failure. Run the query.

When the query returns null the queue is absent, not merely un-required — unchecking the box discards the configuration (merge method, concurrency, timeout: all gone with it). §4 is the live landing path in that case.

§3 (the queue flow) is conditional on that query, not on this sentence. Everything in §2, §5 and §6 holds either way.

History, so the sequence is not relearned: the queue was enabled and required (#2965) while no workflow triggered on merge_group (#2967), so a direct merge returned 405 Changes must be made through the merge queue and every queued PR would have been ejected after the 3600s timeout. Nothing could land at all. The founder unchecked the box; #2971 wires the triggers.


1. What the merge queue changes, and what it does not​

The queue changes landing. It changes nothing about authoring, pushing, reviewing, or resolving conflicts.

BeforeWith the queue
Getting a behind PR up to dateyou run update-branch, all ~40 checks re-run, ~11.5 minthe queue does it, on a throwaway ref; your branch never moves
Racing other mergesyou lose — main merged 7× in 100 min on 2026-09-02 while one CI cycle costs 11.5 minserialised; there is no race to lose
Verify-then-merge gapthe gap is where the 405 comes from; must be collapsed into one command blockgone; enqueue, then bounded observation per AGENTS.md
Two individually-green PRs that break when combinedundetected — see the #2950/#2953 lint incompatibility belowdetected; the queue builds the composed tree and runs CI on it
Textual conflictsyou resolve themyou still resolve them — the queue refuses a CONFLICTING PR outright
Pushing .github/workflows/**needs an App with workflows: writeunchanged
Attributionsquash-merge preserves the PR authorunchanged — the queue is configured SQUASH

The single most useful sentence: the queue removes the update-branch treadmill (#2842) and the composed-tree blind spot; it removes nothing else.

2. Things that are true regardless of how a PR lands​

These are the parts of the old landing recipes that survive. They are properties of the repo, not of the merge mechanism.

2.1 Squash is load-bearing for attribution​

main is squash-merged. Of the wrong-agent commits ever recorded on main, 100% arrived via the 3.5% of PRs merged with a merge commit — a squash discards the branch commits' identities and keeps the PR author's. The queue is set to SQUASH for this reason, and it is not a stylistic preference.

If you ever merge by hand, supply an explicit commit_title and commit_message. Letting GitHub auto-generate the message is what produces a squash whose author resolves to nobody (core#2710).

2.2 First decide union-or-replace; only then resolve​

Never git checkout --ours / --theirs on a file where only one hunk conflicted. That takes the whole file and silently drops the other side's non-conflicting additions. On #2953 exactly one -run line conflicted, so the guard steps raised no marker; taking ours wholesale went 39 → 38 guard steps and dropped #2950's #2861 guard with no visible sign.

But union is not the universal answer, and assuming it is has its own failure. db.yml conflicted twice in one afternoon with opposite correct resolutions: once the shared -run allow-list, where the right answer was to take the side that deleted it (that list was itself the stranding mechanism); once two guard blocks added at the same location, where taking either side wholesale would have un-registered a guard. A union-only procedure cannot express the first, and a set-difference check that is union-preserving in both directions can never say "this should be gone".

So the first question is which kind of conflict it is, and it is decidable:

Compare the two sides' unit names — for a workflow, the - name: values.

  • Disjoint ⇒ both sides added ⇒ union, keep both whole.
  • Overlapping ⇒ the sides are replacing one another ⇒ pick a side, and say in the commit message which and why.
  • A rename is in play ⇒ the test does not hold. X → Y on one side plus modify X on the other reads as disjoint and will happily union a renamed unit with its own original — invisible to lost/invented, because both names are present. Compare bodies, not names.

Assert the answer mechanically rather than eyeballing it; on #2960 that assertion is what distinguished the second conflict from the first.

The procedure, once you know which:

  1. Enumerate on every dimension the file carries, not just the one that conflicted. For db.yml that is the -run entries and the guard step names.
  2. Compute lost and invented against both parents.
    • For a union: both must be empty.

    • For a deletion, the two parents are asymmetric and must be checked separately — this is the clause the first draft was missing, and read literally it licensed exactly the casualty the arm exists to catch:

      • against the parent that retained the thing, lost is expected to be non-empty and must contain exactly what you meant to drop, by name;
      • against the parent that deleted it, lost must be empty. Anything here is a silent casualty — you are dropping something that side never asked you to.

      A single union-preserving check cannot express either half.

      This arm caught a live one on #2975, hours after it was written. Guard blocks in db.yml across that merge:

      tree- name: "Guard blockscarries the #2918 T16 guard
      main 407199da40yes
      PR tree 61b1d9c840no
      head 93e2036141yes

      A guard block went missing while the count stayed identical, because the PR added one of its own. Cardinality read that as parity; only the by-name lost check sees it. Never assert preservation with a count — a same-sized set is not the same set.

  3. The resolution unit is the block, not the expression. Reconciling two edits inside one if: or one shell line by merging the expressions produces something that parses, passes CI, and means neither thing. Keep both blocks.

Point 3 is not stylistic. On #2960 both guard blocks define a shell variable NEEDLE, and a tidy union folding them into one run: would have clobbered the first — while passing step-name checks, passing lost/invented, and parsing as YAML, leaving one guard green whilst asserting the other's test names. The guarantee is one shell per run:, not that the steps look separate. Measured afterwards: 38 steps in that one job define NEEDLE — 37 by NEEDLE= assignment and one by a for NEEDLE in loop binding — so this is a family-wide latent defect, not a two-step coincidence.

Two drafts of that sentence were wrong before this one, in opposite directions, and the near-miss is worth keeping visible. The first said "38 … define" while having counted mentions — right number, wrong predicate. The correction narrowed to NEEDLE= and got 37, which is the count of one syntactic form, not of the defect: a for NEEDLE in binding defines the variable and is exposed to a block-merge clobber identically. Over this file the two predicates do not even differ — every step that mentions NEEDLE also defines it — so the looser measurement was right by accident and the tighter one was wrong on purpose.

Grepping for the assignment operator is not the same as grepping for the definition. Name the predicate your claim needs, then measure that one; tightening a pattern until it is unambiguous is not the same as aiming it.

Which is the same error as §2.2's missing deletion arm, one clause over: an honestly-named predicate that is not the one the argument rests on.

Cheap cross-check that a base merge changed nothing that matters: compare the blob SHA of the contested file across the update — git rev-parse <head>:.github/workflows/db.yml. Identical blob, nothing to re-review.

2.3 Sweep for closing keywords, decoded, against the API​

GitHub parses closing keywords mid-sentence and decodes HTML character entities before matching. On #2953 the body read clos&#101;d #2872; an unanchored raw-text grep found nothing while closingIssuesReferences was [2872]. The escape removed the evidence and kept the effect.

# sweep the body RAW and html.unescape'd, then confirm against the API —
# never against a grep alone
gh api graphql -f query='
{ repository(owner:"upsquad-ai",name:"upsquad-core"){
pullRequest(number:<N>){ closingIssuesReferences(first:20){ nodes{ number } } } } }'

Assert length 0 unless the closure is intended. A body PATCH does not dismiss reviews — it is not a push — so fixing this is free.

2.4 devops-engineer and backend-sme are the only Apps that can push .github/workflows/**​

Measured 2026-09-04T06:46Z — and this has changed twice, so reproduce it:

sg upsquad-devs -c 'python3 scripts/show-app-permissions.py workflows'
Appworkflows
upsquad-backend-sme[bot]write
upsquad-devops-engineer[bot]write
the other five(absent)

The 2026-08-13 revocation decision has now been executed against frontend-sme and project-manager. docs/runbooks/workflow-yaml-handoff.md still shows the four-holder table; treat the command above as the answer.

The queue does not touch this. GitHub refuses an App-token push whose tree touches .github/workflows/** without the permission, and that refusal happens long before any merge mechanism is involved. The hand-off procedure in workflow-yaml-handoff.md is unchanged.

Corollary that still binds: pushing for someone else makes you a co-author of the branch, which disqualifies you from being the independent approver on it. Sequence: disclose → push → someone else approves → merge.

2.5 [build failed] on a path under .cache/go-build is the disk guard, not your change​

Re-run before investigating. scripts/dev-disk-guard.sh runs every 15 minutes on the devbox and clears the Go build cache when the disk is >= 85%. If it fires while you are running go build ./... or go test, roughly twenty packages report [build failed] at once, on an import path under ~/.cache/go-build.

It looks like you have broken the module. You have not — the compiler lost files underneath itself mid-run. The tell is the import path: a real breakage names a package in internal|cmd|pkg|test, this one names a cache directory. A plain re-run is clean, and costs a minute against the hour it costs to go looking for a dependency cycle that is not there.

Two things follow. Never report a local build result you have not re-run once if the disk was anywhere near the threshold — the artefact is indistinguishable from a real failure in a screenshot. And when the disk is tight, go clean -cache deliberately before you start rather than letting the guard fire at a random moment; the cache is the only lever that moves the number (it regrew 12K to 22G in a single session), and doing it yourself makes the timing yours.

2.6 Verify migration ordinals immediately before the merge, not at authoring​

A migration whose ordinal is at or below a long-lived DB's current version is skipped silently, rc=0. CI is always fresh so CI can never catch it. Re-derive the ceiling against main in the same breath as the landing.

With the queue this gets more important, not less: the queue may hold your PR for an unbounded time while others land ahead of it, so the gap between "I checked the ordinal" and "it applied" is wider than it used to be.

3. The queue flow — use this iff §6's query returns a queue​

open PR ──▶ reviewer bot APPROVEs ──▶ enqueue ──▶ queue builds composed tree
and runs the 10 checks
│
green ────┴──── red
│ │
squashed ejected;
onto main you fix and re-enqueue

What a landing agent does:

  1. Confirm the PR is MERGEABLE, not CONFLICTING. The queue will not resolve a textual conflict for you. If it conflicts, resolve per §2.2 first.
  2. Confirm an independent APPROVE exists, latest-per-reviewer with COMMENTED excluded:
    gh api "repos/upsquad-ai/upsquad-core/pulls/<N>/reviews?per_page=100" \
    | jq -r '[.[]|select(.state!="COMMENTED")] | group_by(.user.login)
    | map(last) | map(.state)'
    Blocked iff any element is CHANGES_REQUESTED. [] means nobody was ever dispatched.
  3. Sweep closing keywords (§2.3) and re-check the migration ordinal (§2.6). Both are cheaper to fix before enqueueing than after.
  4. Enqueue. From the UI it is the "Merge when ready" button. From the API:
    gh api graphql -f query='
    mutation($id:ID!){ enqueuePullRequest(input:{pullRequestId:$id}){
    mergeQueueEntry{ position state enqueuedAt } } }' -f id=<PR node id>
    Dequeue with dequeuePullRequest if you need it out.
  5. Observe through the shared bounded waiter. Follow AGENTS.md → “Bounded in-session waiting and recovery” (core#3439), preserving the successful enqueue receipt/timestamp and exact head. That founder contract supersedes the old enqueue-and-stop rule; the waiter observes only and grants no recovery or landing permission.

Do not run update-branch on a queued PR. It is pure waste — it burns ~11.5 min of CI to produce a head the queue is going to discard and rebuild anyway.

What the queue does not excuse​

  • A skipped required check still satisfies the gate. The queue evaluates the same ten contexts by the same rule. Verify execution, not colour — this is the repo's most-repeated defect class and the queue gives it a straight path to main.
  • An approved-and-green PR is a necessary condition for landing, never a sufficient one. If it falls on CLAUDE.md's human-decision list (PRD/ADR/HLD, prod promotion, security-critical, new dependency, data-destructive), it waits for the founder no matter what the queue says.

4. Direct-merge procedure — use this iff §6's query returns null​

This is the live path whenever there is no queue, which as of 2026-09-04 is the case. It is not history: a queue can be switched off in one click, and has been.

Under strict: true with a busy main, a human-paced update → look → merge round trip loses the race essentially every time. On 2026-09-02 main merged 7 times in 100 minutes against an 11.5-minute CI cycle; two fully-green heads were computed and thrown away purely to base staleness.

The only thing that works is one detached script doing update-branch, poll-to-settle, full verification and the merge with no human round trip in between. Pin sha=<head> on the merge call throughout, so losing the race is a harmless 405/422 and never a wrong merge.

Two traps inside that loop, both of which cost a full cycle when hit:

  • gh-token.py hands back a cached token. REFRESH_BUFFER_SECS = 300, so a "fresh" call can return a token with 9 minutes left on it. Correct for a single gh api call, wrong for any loop longer than ~5 minutes. Re-mint every iteration; the cache makes that a no-op until the buffer.
  • An auth failure can render as "zero check-runs." A swallowed 401 plus .get("check_runs", []) logs as total=0 pending=0 — indistinguishable from a settled head. Write the settle condition as if runs and not pend, assert the required contexts by name, and abort after three consecutive empty enumerations. Never let an error path and a legitimate empty result share a representation.

Two rules from this era that the queue makes obsolete​

Recorded so nobody reconstructs them and thinks they are new:

  • "update-branch does not dismiss the approval" (5 for 5) — true, but the rule that was generalised from it is wrong. The five observations were all of GitHub's update-branch API, and the explanation attached to them — "dismiss_stale_reviews fires on pushes that change the PR's own diff; a base-into-head merge contributes none" — is not the mechanism.

    The discriminator is who performed the merge, not what kind of merge it was. Measured on #2963 (2026-09-04): a conflict-resolution merge commit created locally and git pushed dismissed the architect's approval (5110058052, still pinned to the pre-push head), while five update-branch operations on comparable PRs did not.

    dismisses?
    GitHub's update-branch APIno — GitHub exempts its own operation
    a merge commit you create and pushyes — it is an ordinary push

    Both produce a merge commit and both move the head SHA, which is why the wrong rule survived five confirmations. Budget a re-approval round trip for every hand-resolved conflict — and note you cannot supply it yourself, because pushing the resolution made you a co-author (§2.4).

    Under the queue this stops arising for base staleness — the queue merges on a ref your branch never sees — but it still arises for every textual conflict, which the queue refuses to resolve.

  • "An approval carries iff the merge is clean." The discriminator was mergeable_state == clean, never the review object. Under the queue the approval is evaluated once at enqueue time and the queue owns everything after it. Do not carry this rule forward as a licence to assume an approval exists — on #2954 a coordinator asserted a prior review carried and there was no review at all. "An approval carries" and "no approval exists" are very different states to merge on; measure with the group_by query in §3.2.

5. strict: true under a merge queue​

required_status_checks.strict was true when last measured (2026-09-04); re-derive it rather than trusting that. It is redundant with the queue and actively harmful:

  • The queue already guarantees the stronger property — CI green on the PR's changes applied to the tip of main plus everything ahead of it in the queue. "Branch is up to date" is a weaker approximation of that.
  • strict blocks a behind PR from being enqueued at all, forcing the update-branch dance before entry — i.e. the whole treadmill of #2842 is reimposed as a toll booth in front of the machine built to remove it.

Recommendation: drop strict, keep the ten required contexts. Not changed unilaterally — branch-protection changes are a founder/architect decision (#1901). Tracked on #2965.

5b. Re-enabling the queue — the acceptance test is binding​

Wiring the triggers is not evidence the wiring works. Before the queue is trusted with main again, all of this must hold:

  1. Measure on the gh-readonly-queue/main/pr-N-<sha> head, not the PR head. The PR head's checks came from the pull_request event and say nothing about the merge_group path. This is the whole point and it is easy to get wrong, because the PR head is what every habit reaches for.
  2. All ten required contexts, by name, with conclusion == "success" — and != "skipped" and != null asserted separately, so the three ways a context can satisfy protection without doing its job are each named. A skipped required check satisfies branch protection; null is a third state that a two-way partition mis-sorts.
  3. The first enqueued PR must be docs-only. Every paths-filter answers false on a docs/-only diff, so watching the four filtered lanes (db, proto, security-scan, wave2-smoke) run anyway is the only runtime test of the forced-true half — the half carrying the entire safety argument. A code PR exercises the filters' honest true and proves nothing about the override.

Explicitly rejected as evidence:

offeredwhy it proves nothing
"it merged through the queue"indistinguishable from the rubber-stamp outcome, where a lane skipped-green and the queue merged on one real check
a hand-pushed gh-readonly-queue brancha push event, not merge_group; different trigger, different answer
check-required-check-gating.py passingstatic; it asserts the YAML says the right thing, never that a runner did it

Reassuringly, the queue as wired would have caught its own justifying incident: repoRootFromLint surfaces on Go Unit Tests and the guard-step drop on Bash script tests — both required, both now wired to merge_group.

6. There is a THIRD source of merge gating, and neither existing check sees it​

CLAUDE.md tells you to re-derive branch protection from (1) the classic protection endpoint and (2) the rulesets endpoint. Both are blind to the merge queue. Classic protection has no merge-queue field in its REST schema at all, BranchProtectionRule has none in GraphQL, and rules/branches/main returns []. A queue that stops every merge in the repo is invisible to both calls.

The only way to read it:

gh api graphql -f query='
{ repository(owner:"upsquad-ai",name:"upsquad-core"){
mergeQueue(branch:"main"){
configuration{ mergeMethod mergingStrategy
maximumEntriesToBuild maximumEntriesToMerge
minimumEntriesToMerge minimumEntriesToMergeWaitTime
checkResponseTimeout }
entries(first:20){ totalCount nodes{ position state
pullRequest{ number } } } } } }'

mergeQueue is null when no queue is configured — verified against upsquad-client/main (null) and a non-existent branch (null) on the same call, so a non-null result is a real reading and not an artefact.

That call tells you a queue EXISTS. It does not tell you the queue is REQUIRED, and no read-only endpoint does. The only thing that answers it is attempting the merge:

PUT /repos/upsquad-ai/upsquad-core/pulls/2953/merge
405 {"message":"Changes must be made through the merge queue"}

Measured 2026-09-04 on a PR that was approved, clean, and 39/39 green. Note what the 405 is not: 405 Pull Request has merge conflicts and 405 Pull Request is not mergeable are the same status code with different messages, so read the message, never the code. A conflicted PR cannot distinguish these cases at all — resolve first, then probe.

Its configuration is UI-only. There is no REST endpoint and no GraphQL mutation for it: updateBranchProtectionRule has no merge-queue input field, and the only queue mutations in the schema are enqueuePullRequest / dequeuePullRequest. Changing concurrency, batch size, merge method or timeout requires a human at https://github.com/upsquad-ai/upsquad-core/settings/branches → edit the main rule → Require merge queue. Record any change on #2965, because nothing else will show it moved.