Spindle fleet · model-seat trials · protocol GLM-BENCH

Can glm-5.3 hold a codex seat?

run 2026-08-16 → 17 · 9 scored cells
plan z.ai LITE · thinking default-max
ground truth: shipped fix commits + CONVERGENCE.ndjson

Verdict: assist everywhere, confirm nowhere. GLM-5.3 reads real code with zero hallucinated citations and keeps finding true defects — but in every cell it either misses one must or under-weighs it. The seats it can take are the ones a stronger confirm still backstops. glm-4.7 is disqualified outright: it rationalized both planted musts as correct code. Late addition — the standout: on the Phase-1 discovery breadth seat, glm-5.3 reproduced ~88% of the real opus fleet's repo-derivable synthesis — including both items the build later proved load-bearing — for about ten cents.

Seat board the decision

Each review seat in the pipeline, stamped with where GLM lands after seven cells. This is what the policy flag enables — nothing more.

loop-iterate

R2…N−1 middle rounds of plan & Phase-4 loops (today: terra)
TAKE when codex hot
Finds real musts with clean severities on non-legal surfaces; the 5.5 confirm catches what it under-weighs.
evidence: Q5 · Q1-53

checkpoint-confirm

Standard-step confirm + heavy R2 at build fences (today: terra)
TAKE — non-consent surfaces
Q1: 5/6 substance. Q2 (consent surface): missed all 3 integrity musts — route consent/money fences back to codex.
evidence: Q1-53 · Q2

R1 breadth

Fresh-eyes first sweep beside the Sonnet swarm
ADDITIVE LENS ONLY
1/8 primary recall solo — but 2 novel valid finds. Worth a free extra pair of eyes, never the swarm's replacement.
evidence: A · B

confirming 5.5

The round that declares convergence
KEEP CODEX
Q5: closed with a data-integrity must still open. A solo GLM confirm = false-clean risk, the one unaffordable failure.
evidence: Q5 · Q2

build.rote

Mechanical step implementation (today: Sonnet)
YES with guardrails
~90% fidelity on a shipped step; 2 fence-invisible defects say keep the cargo-check fence + checkpoint review.
evidence: C

discovery breadth

Phase-1 mechanical codebase enumeration (today: Sonnet/opus @ xhigh)
STRONGEST SEAT
88% of the opus synthesis incl. the any_runner_wanted trap the build's R3 later confirmed. Web + deep-seam + ENHANCE-SCOUT seats stay claude.
evidence: D1

any seat · glm-4.7

The 1× -quota breadth model
BANNED
False-clean: argued both planted musts were correct code, in 155 s. Cheap is not a defense for a reviewer that certifies bugs.
evidence: Q1-47

Recall against ground truth 7 cells

Bar = share of the ground-truth findings GLM independently re-derived (substance match). Dots = the musts: filled caught, hollow missed. Every GLM citation was re-opened against the sandbox tree — hallucination rate 0.00 in all fully-audited cells.

Q1 · glm-5.3checkpoint confirm · CONSENT-ARTIFACT S9
75% 5/6
Q5 · glm-5.35.5-cell confirm · INT7-RT-SHEETS
67% 2/3
Q2 · glm-5.3terra-cell confirm · CONSENT-ARTIFACT S8
50% 3.5/7
D1 · glm-5.3discovery breadth · INT7-RT-SHEETS
88% 11.5/13
D1 · glm-4.7discovery breadth · INT7-RT-SHEETS
38% + 3 anti-hits
A · glm-5.3R1 breadth · MAP1a-2
13% 1/8
B · glm-5.3confirm/middle · MAP1a-2
0% 0/20
Q1 · glm-4.7checkpoint confirm · CONSENT-ARTIFACT S9
0% + false-clean
C · glm-5.3build.rote · MAP1a-2 S4
~90% impl fidelity
glm-5.3 vs ground truth false-clean (certified bugs as correct)  must caught  must missed

The severity tell the recurring defect

When GLM matches a finding, does it weigh it right? On legal/evidentiary surfaces the answer is no — it sees the bug and calls it smaller. This is why the confirming seat stays codex.

Q1 — consent erasure (3 of 5 matches downgraded)

must→shouldenvelope-MAC broken on prior-phone rows
should→couldpreview under-reports lead rows
should→notereceipt over-claims retention

Q2 — consent backfill (both matches downgraded)

must→shouldmissing_legacy never stamped
should→couldbackfill can mint evidence w/o audit row

Q5 — sheets integration (0 downgrades)

must→mustholds transient Err revokes healthy channel
should→shouldholds runbook canary unexecutable
pattern: severity judgment degrades exactly where legal weight enters

Expenditure measured, not modeled

Quota burn measured off the z.ai meter during the runs; capacity assumes 7–8 GLM-relief passes per unit (3–4 fence confirms + ~4 loop iterates). Plan prices approximate — confirm on billing.

2–11% weekly / pass
LITE quota per review pass
Scoped confirm (Q5) = 2%; big-surface confirm (Q2) = 11%. Blended ≈ 5%.
$0.07–0.15 / pass
marginal cost, LITE ≈ $6/mo
≈ $0.50–1 per unit of GLM-relief seats. Codex yardstick: ≈ $5–6 of ChatGPT-Pro spend per unit at the current mix.
~10× cheaper
per absorbed seat vs codex
Every terra seat GLM absorbs frees codex weekly for the deep seats that are irreplaceable.
LITE plan
~3–5 u/wk
PRO plan (6×)
~20–30 u/wk
fleet demand
~7–10 u/wk
units/week of GLM-relief capacity what the fleet actually clears when hot ⇒ LITE can't cover demand · PRO is effectively unconstrained

All cells table view

CellSeat emulatedSurfaceModelGTSubstanceMust recallSeverity callsNovel validHalluc.Verdict
AR1 fresh-eyesMAP1a-2 whole-unitglm-5.38 (2m/3s)1/81/1 explicit—20.00additive lens; not replacement
Bconfirm/middleMAP1a-2 whole-unitglm-5.3~20 (6m)0/200/6—20.00no standalone
Cbuild.roteMAP1a-2 S4 implglm-5.3n/a~90% fidelity2 fence-invisible defects—n/a0.00yes, with guardrails
Q1-47checkpoint confirmCONSENT-ARTIFACT S9glm-4.76 (2m/4s)0/60/2 + rationalized both—02 false claimsBANNED — false-clean
Q1-53checkpoint confirmCONSENT-ARTIFACT S9glm-5.36 (2m/4s)5/6 (75% exact)2/2 content · 1/2 severity3 downgrades20.00yes-with-caveats
Q55.5-cell confirmINT7-RT-SHEETS post-R1glm-5.34 (2m/2s·1 routed)2/3 gating1/2both correct20.00iterate yes · confirm no
Q2terra-cell confirmCONSENT-ARTIFACT S8glm-5.37 at-sandbox (4m/3s)3.5/71/4 · 0/4 at severity2 downgrades1 strongspot-onlyadditive; not consent confirm
D1-53discovery breadthINT7-RT-SHEETS pre-discoveryglm-5.313 repo-derivable11.5/13 (88%)n/a—TTLs via repo cross-ref0.00 (anchors cross-check GT)STANDOUT — breadth parity
D1-47discovery breadthINT7-RT-SHEETS pre-discoveryglm-4.713 repo-derivable5/13 (38%)n/a—03 anti-hitsinvents schema vs migration:none

Method. Each cell re-runs a real, already-shipped review round in a frozen git sandbox at the exact commit the original reviewer saw; ground truth = the findings that round actually produced, grounded in the fix commits. Reviewers get read-only tools; every tool call is audited (cheat-flags 0 across all cells). Scores: Q1/Q5 fully citation-audited against the sandbox tree; Q2 spot-audited (flagged). Q2's answer was recovered from its stream log after a runner collision. Policy state: burn_balance.glm.enabled=false — this page is the evidence for that flip, which also awaits the z.ai PRO upgrade. Protocol + raw ledgers: ~/glm-bench/results/ · scored 2026-08-17.