Landing a PR on main — Runbook
Status: active — 2026-09-04. Owner: devops-engineer.
Related: core#2965 (queue enabled), core#2967 (the merge_group wiring),
core#2842 (the treadmill this replaces), core#2710 (attribution), core#2397
(closing keywords).
STATUS IS NOT RECORDED HERE — read it with the query in §6.
repository.mergeQueue(branch:"main")is the only thing that can answer, and it is two seconds to run.This block used to carry the answer anyway — "as of 2026-09-04 it returns
null, the queue is absent" — and by 2026-09-08 that was false: the queue was live,ALLGREEN, with three entries in it (core#3242). Someone following the recorded status would have taken §4 and hit a405. The number is deleted rather than updated, because updating it just resets the clock on the same failure. Run the query.When the query returns
nullthe queue is absent, not merely un-required — unchecking the box discards the configuration (merge method, concurrency, timeout: all gone with it). §4 is the live landing path in that case.§3 (the queue flow) is conditional on that query, not on this sentence. Everything in §2, §5 and §6 holds either way.
History, so the sequence is not relearned: the queue was enabled and required (#2965) while no workflow triggered on
merge_group(#2967), so a direct merge returned405 Changes must be made through the merge queueand every queued PR would have been ejected after the 3600s timeout. Nothing could land at all. The founder unchecked the box; #2971 wires the triggers.
1. What the merge queue changes, and what it does not
The queue changes landing. It changes nothing about authoring, pushing, reviewing, or resolving conflicts.
| Before | With the queue | |
|---|---|---|
Getting a behind PR up to date | you run update-branch, all ~40 checks re-run, ~11.5 min | the queue does it, on a throwaway ref; your branch never moves |
| Racing other merges | you lose — main merged 7× in 100 min on 2026-09-02 while one CI cycle costs 11.5 min | serialised; there is no race to lose |
| Verify-then-merge gap | the gap is where the 405 comes from; must be collapsed into one command block | gone; enqueue, then bounded observation per AGENTS.md |
| Two individually-green PRs that break when combined | undetected — see the #2950/#2953 lint incompatibility below | detected; the queue builds the composed tree and runs CI on it |
| Textual conflicts | you resolve them | you still resolve them — the queue refuses a CONFLICTING PR outright |
Pushing .github/workflows/** | needs an App with workflows: write | unchanged |
| Attribution | squash-merge preserves the PR author | unchanged — the queue is configured SQUASH |
The single most useful sentence: the queue removes the update-branch
treadmill (#2842) and the composed-tree blind spot; it removes nothing else.
2. Things that are true regardless of how a PR lands
These are the parts of the old landing recipes that survive. They are properties of the repo, not of the merge mechanism.
2.1 Squash is load-bearing for attribution
main is squash-merged. Of the wrong-agent commits ever recorded on main,
100% arrived via the 3.5% of PRs merged with a merge commit — a squash
discards the branch commits' identities and keeps the PR author's. The queue is
set to SQUASH for this reason, and it is not a stylistic preference.
If you ever merge by hand, supply an explicit commit_title and
commit_message. Letting GitHub auto-generate the message is what produces a
squash whose author resolves to nobody (core#2710).
2.2 First decide union-or-replace; only then resolve
Never git checkout --ours / --theirs on a file where only one hunk
conflicted. That takes the whole file and silently drops the other side's
non-conflicting additions. On #2953 exactly one -run line conflicted, so the
guard steps raised no marker; taking ours wholesale went 39 → 38 guard steps and
dropped #2950's #2861 guard with no visible sign.
But union is not the universal answer, and assuming it is has its own
failure. db.yml conflicted twice in one afternoon with opposite correct
resolutions: once the shared -run allow-list, where the right answer was to
take the side that deleted it (that list was itself the stranding mechanism);
once two guard blocks added at the same location, where taking either side
wholesale would have un-registered a guard. A union-only procedure cannot express
the first, and a set-difference check that is union-preserving in both directions
can never say "this should be gone".
So the first question is which kind of conflict it is, and it is decidable:
Compare the two sides' unit names — for a workflow, the
- name:values.
- Disjoint ⇒ both sides added ⇒ union, keep both whole.
- Overlapping ⇒ the sides are replacing one another ⇒ pick a side, and say in the commit message which and why.
- A rename is in play ⇒ the test does not hold.
X → Yon one side plusmodify Xon the other reads as disjoint and will happily union a renamed unit with its own original — invisible tolost/invented, because both names are present. Compare bodies, not names.
Assert the answer mechanically rather than eyeballing it; on #2960 that assertion is what distinguished the second conflict from the first.
The procedure, once you know which:
- Enumerate on every dimension the file carries, not just the one that
conflicted. For
db.ymlthat is the-runentries and the guard step names. - Compute
lostandinventedagainst both parents.-
For a union: both must be empty.
-
For a deletion, the two parents are asymmetric and must be checked separately — this is the clause the first draft was missing, and read literally it licensed exactly the casualty the arm exists to catch:
- against the parent that retained the thing,
lostis expected to be non-empty and must contain exactly what you meant to drop, by name; - against the parent that deleted it,
lostmust be empty. Anything here is a silent casualty — you are dropping something that side never asked you to.
A single union-preserving check cannot express either half.
This arm caught a live one on #2975, hours after it was written. Guard blocks in
db.ymlacross that merge:tree - name: "Guardblockscarries the #2918 T16 guard main407199da40 yes PR tree 61b1d9c840 no head 93e2036141 yes A guard block went missing while the count stayed identical, because the PR added one of its own. Cardinality read that as parity; only the by-name
lostcheck sees it. Never assert preservation with a count — a same-sized set is not the same set. - against the parent that retained the thing,
-
- The resolution unit is the block, not the expression. Reconciling two
edits inside one
if:or one shell line by merging the expressions produces something that parses, passes CI, and means neither thing. Keep both blocks.
Point 3 is not stylistic. On #2960 both guard blocks define a shell variable
NEEDLE, and a tidy union folding them into one run: would have clobbered the
first — while passing step-name checks, passing lost/invented, and parsing as
YAML, leaving one guard green whilst asserting the other's test names. The
guarantee is one shell per run:, not that the steps look separate. Measured
afterwards: 38 steps in that one job define NEEDLE — 37 by NEEDLE=
assignment and one by a for NEEDLE in loop binding — so this is a family-wide
latent defect, not a two-step coincidence.
Two drafts of that sentence were wrong before this one, in opposite directions,
and the near-miss is worth keeping visible. The first said "38 … define" while
having counted mentions — right number, wrong predicate. The correction
narrowed to NEEDLE= and got 37, which is the count of one syntactic form,
not of the defect: a for NEEDLE in binding defines the variable and is exposed
to a block-merge clobber identically. Over this file the two predicates do not
even differ — every step that mentions NEEDLE also defines it — so the looser
measurement was right by accident and the tighter one was wrong on purpose.
Grepping for the assignment operator is not the same as grepping for the definition. Name the predicate your claim needs, then measure that one; tightening a pattern until it is unambiguous is not the same as aiming it.
Which is the same error as §2.2's missing deletion arm, one clause over: an honestly-named predicate that is not the one the argument rests on.
Cheap cross-check that a base merge changed nothing that matters: compare the
blob SHA of the contested file across the update —
git rev-parse <head>:.github/workflows/db.yml. Identical blob, nothing to
re-review.
2.3 Sweep for closing keywords, decoded, against the API
GitHub parses closing keywords mid-sentence and decodes HTML character
entities before matching. On #2953 the body read closed #2872; an
unanchored raw-text grep found nothing while closingIssuesReferences was
[2872]. The escape removed the evidence and kept the effect.
# sweep the body RAW and html.unescape'd, then confirm against the API —
# never against a grep alone
gh api graphql -f query='
{ repository(owner:"upsquad-ai",name:"upsquad-core"){
pullRequest(number:<N>){ closingIssuesReferences(first:20){ nodes{ number } } } } }'
Assert length 0 unless the closure is intended. A body PATCH does not
dismiss reviews — it is not a push — so fixing this is free.
2.4 devops-engineer and backend-sme are the only Apps that can push .github/workflows/**
Measured 2026-09-04T06:46Z — and this has changed twice, so reproduce it:
sg upsquad-devs -c 'python3 scripts/show-app-permissions.py workflows'
| App | workflows |
|---|---|
upsquad-backend-sme[bot] | write |
upsquad-devops-engineer[bot] | write |
| the other five | (absent) |
The 2026-08-13 revocation decision has now been executed against frontend-sme
and project-manager. docs/runbooks/workflow-yaml-handoff.md still shows the
four-holder table; treat the command above as the answer.
The queue does not touch this. GitHub refuses an App-token push whose tree
touches .github/workflows/** without the permission, and that refusal happens
long before any merge mechanism is involved. The hand-off procedure in
workflow-yaml-handoff.md is unchanged.
Corollary that still binds: pushing for someone else makes you a co-author of the branch, which disqualifies you from being the independent approver on it. Sequence: disclose → push → someone else approves → merge.
2.5 [build failed] on a path under .cache/go-build is the disk guard, not your change
Re-run before investigating. scripts/dev-disk-guard.sh runs every 15 minutes
on the devbox and clears the Go build cache when the disk is >= 85%. If it fires
while you are running go build ./... or go test, roughly twenty packages
report [build failed] at once, on an import path under ~/.cache/go-build.
It looks like you have broken the module. You have not — the compiler lost files
underneath itself mid-run. The tell is the import path: a real breakage names a
package in internal|cmd|pkg|test, this one names a cache directory. A plain
re-run is clean, and costs a minute against the hour it costs to go looking for a
dependency cycle that is not there.
Two things follow. Never report a local build result you have not re-run once
if the disk was anywhere near the threshold — the artefact is indistinguishable
from a real failure in a screenshot. And when the disk is tight, go clean -cache deliberately before you start rather than letting the guard fire at a
random moment; the cache is the only lever that moves the number (it regrew 12K
to 22G in a single session), and doing it yourself makes the timing yours.
2.6 Verify migration ordinals immediately before the merge, not at authoring
A migration whose ordinal is at or below a long-lived DB's current version is
skipped silently, rc=0. CI is always fresh so CI can never catch it. Re-derive
the ceiling against main in the same breath as the landing.
With the queue this gets more important, not less: the queue may hold your PR for an unbounded time while others land ahead of it, so the gap between "I checked the ordinal" and "it applied" is wider than it used to be.
3. The queue flow — use this iff §6's query returns a queue
open PR ──▶ reviewer bot APPROVEs ──▶ enqueue ──▶ queue builds composed tree
and runs the 10 checks
│
green ────┴──── red
│ │
squashed ejected;
onto main you fix and re-enqueue
What a landing agent does:
- Confirm the PR is
MERGEABLE, notCONFLICTING. The queue will not resolve a textual conflict for you. If it conflicts, resolve per §2.2 first. - Confirm an independent APPROVE exists, latest-per-reviewer with
COMMENTEDexcluded:Blocked iff any element isgh api "repos/upsquad-ai/upsquad-core/pulls/<N>/reviews?per_page=100" \| jq -r '[.[]|select(.state!="COMMENTED")] | group_by(.user.login)| map(last) | map(.state)'CHANGES_REQUESTED.[]means nobody was ever dispatched. - Sweep closing keywords (§2.3) and re-check the migration ordinal (§2.6). Both are cheaper to fix before enqueueing than after.
- Enqueue. From the UI it is the "Merge when ready" button. From the API:
Dequeue withgh api graphql -f query='mutation($id:ID!){ enqueuePullRequest(input:{pullRequestId:$id}){mergeQueueEntry{ position state enqueuedAt } } }' -f id=<PR node id>
dequeuePullRequestif you need it out. - Observe through the shared bounded waiter. Follow
AGENTS.md→ “Bounded in-session waiting and recovery” (core#3439), preserving the successful enqueue receipt/timestamp and exact head. That founder contract supersedes the old enqueue-and-stop rule; the waiter observes only and grants no recovery or landing permission.
Do not run update-branch on a queued PR. It is pure waste — it burns ~11.5
min of CI to produce a head the queue is going to discard and rebuild anyway.
What the queue does not excuse
- A skipped required check still satisfies the gate. The queue evaluates the
same ten contexts by the same rule. Verify execution, not colour — this is
the repo's most-repeated defect class and the queue gives it a straight path to
main. - An approved-and-green PR is a necessary condition for landing, never a sufficient one. If it falls on CLAUDE.md's human-decision list (PRD/ADR/HLD, prod promotion, security-critical, new dependency, data-destructive), it waits for the founder no matter what the queue says.
4. Direct-merge procedure — use this iff §6's query returns null
This is the live path whenever there is no queue, which as of 2026-09-04 is the case. It is not history: a queue can be switched off in one click, and has been.
Under strict: true with a busy main, a human-paced update → look → merge
round trip loses the race essentially every time. On 2026-09-02 main merged 7
times in 100 minutes against an 11.5-minute CI cycle; two fully-green heads were
computed and thrown away purely to base staleness.
The only thing that works is one detached script doing update-branch,
poll-to-settle, full verification and the merge with no human round trip in
between. Pin sha=<head> on the merge call throughout, so losing the race is a
harmless 405/422 and never a wrong merge.
Two traps inside that loop, both of which cost a full cycle when hit:
gh-token.pyhands back a cached token.REFRESH_BUFFER_SECS = 300, so a "fresh" call can return a token with 9 minutes left on it. Correct for a singlegh apicall, wrong for any loop longer than ~5 minutes. Re-mint every iteration; the cache makes that a no-op until the buffer.- An auth failure can render as "zero check-runs." A swallowed 401 plus
.get("check_runs", [])logs astotal=0 pending=0— indistinguishable from a settled head. Write the settle condition asif runs and not pend, assert the required contexts by name, and abort after three consecutive empty enumerations. Never let an error path and a legitimate empty result share a representation.
Two rules from this era that the queue makes obsolete
Recorded so nobody reconstructs them and thinks they are new:
-
"
update-branchdoes not dismiss the approval" (5 for 5) — true, but the rule that was generalised from it is wrong. The five observations were all of GitHub'supdate-branchAPI, and the explanation attached to them — "dismiss_stale_reviewsfires on pushes that change the PR's own diff; a base-into-head merge contributes none" — is not the mechanism.The discriminator is who performed the merge, not what kind of merge it was. Measured on #2963 (2026-09-04): a conflict-resolution merge commit created locally and
git pushed dismissed the architect's approval (5110058052, still pinned to the pre-push head), while fiveupdate-branchoperations on comparable PRs did not.dismisses? GitHub's update-branchAPIno — GitHub exempts its own operation a merge commit you create and push yes — it is an ordinary push Both produce a merge commit and both move the head SHA, which is why the wrong rule survived five confirmations. Budget a re-approval round trip for every hand-resolved conflict — and note you cannot supply it yourself, because pushing the resolution made you a co-author (§2.4).
Under the queue this stops arising for base staleness — the queue merges on a ref your branch never sees — but it still arises for every textual conflict, which the queue refuses to resolve.
-
"An approval carries iff the merge is clean." The discriminator was
mergeable_state == clean, never the review object. Under the queue the approval is evaluated once at enqueue time and the queue owns everything after it. Do not carry this rule forward as a licence to assume an approval exists — on #2954 a coordinator asserted a prior review carried and there was no review at all. "An approval carries" and "no approval exists" are very different states to merge on; measure with thegroup_byquery in §3.2.
5. strict: true under a merge queue
required_status_checks.strict was true when last measured (2026-09-04);
re-derive it rather than trusting that. It is
redundant with the queue and actively harmful:
- The queue already guarantees the stronger property — CI green on the PR's
changes applied to the tip of
mainplus everything ahead of it in the queue. "Branch is up to date" is a weaker approximation of that. strictblocks abehindPR from being enqueued at all, forcing theupdate-branchdance before entry — i.e. the whole treadmill of #2842 is reimposed as a toll booth in front of the machine built to remove it.
Recommendation: drop strict, keep the ten required contexts.
Not changed unilaterally — branch-protection changes are a founder/architect
decision (#1901). Tracked on #2965.
5b. Re-enabling the queue — the acceptance test is binding
Wiring the triggers is not evidence the wiring works. Before the queue is trusted
with main again, all of this must hold:
- Measure on the
gh-readonly-queue/main/pr-N-<sha>head, not the PR head. The PR head's checks came from thepull_requestevent and say nothing about themerge_grouppath. This is the whole point and it is easy to get wrong, because the PR head is what every habit reaches for. - All ten required contexts, by name, with
conclusion == "success"— and!= "skipped"and!= nullasserted separately, so the three ways a context can satisfy protection without doing its job are each named. Askippedrequired check satisfies branch protection;nullis a third state that a two-way partition mis-sorts. - The first enqueued PR must be docs-only. Every paths-filter answers
false on a
docs/-only diff, so watching the four filtered lanes (db,proto,security-scan,wave2-smoke) run anyway is the only runtime test of the forced-truehalf — the half carrying the entire safety argument. A code PR exercises the filters' honesttrueand proves nothing about the override.
Explicitly rejected as evidence:
| offered | why it proves nothing |
|---|---|
| "it merged through the queue" | indistinguishable from the rubber-stamp outcome, where a lane skipped-green and the queue merged on one real check |
a hand-pushed gh-readonly-queue branch | a push event, not merge_group; different trigger, different answer |
check-required-check-gating.py passing | static; it asserts the YAML says the right thing, never that a runner did it |
Reassuringly, the queue as wired would have caught its own justifying
incident: repoRootFromLint surfaces on Go Unit Tests and the guard-step
drop on Bash script tests — both required, both now wired to merge_group.
6. There is a THIRD source of merge gating, and neither existing check sees it
CLAUDE.md tells you to re-derive branch protection from (1) the classic
protection endpoint and (2) the rulesets endpoint. Both are blind to the merge
queue. Classic protection has no merge-queue field in its REST schema at all,
BranchProtectionRule has none in GraphQL, and rules/branches/main returns
[]. A queue that stops every merge in the repo is invisible to both calls.
The only way to read it:
gh api graphql -f query='
{ repository(owner:"upsquad-ai",name:"upsquad-core"){
mergeQueue(branch:"main"){
configuration{ mergeMethod mergingStrategy
maximumEntriesToBuild maximumEntriesToMerge
minimumEntriesToMerge minimumEntriesToMergeWaitTime
checkResponseTimeout }
entries(first:20){ totalCount nodes{ position state
pullRequest{ number } } } } } }'
mergeQueue is null when no queue is configured — verified against
upsquad-client/main (null) and a non-existent branch (null) on the same call,
so a non-null result is a real reading and not an artefact.
That call tells you a queue EXISTS. It does not tell you the queue is REQUIRED, and no read-only endpoint does. The only thing that answers it is attempting the merge:
PUT /repos/upsquad-ai/upsquad-core/pulls/2953/merge
405 {"message":"Changes must be made through the merge queue"}
Measured 2026-09-04 on a PR that was approved, clean, and 39/39 green. Note
what the 405 is not: 405 Pull Request has merge conflicts and
405 Pull Request is not mergeable are the same status code with different
messages, so read the message, never the code. A conflicted PR cannot
distinguish these cases at all — resolve first, then probe.
Its configuration is UI-only. There is no REST endpoint and no GraphQL
mutation for it: updateBranchProtectionRule has no merge-queue input field, and
the only queue mutations in the schema are enqueuePullRequest /
dequeuePullRequest. Changing concurrency, batch size, merge method or timeout
requires a human at
https://github.com/upsquad-ai/upsquad-core/settings/branches → edit the main
rule → Require merge queue. Record any change on #2965, because nothing else
will show it moved.