Verdict: assist everywhere, confirm nowhere. GLM-5.3 reads real code with zero hallucinated citations and keeps finding true defects — but in every cell it either misses one must or under-weighs it. The seats it can take are the ones a stronger confirm still backstops. glm-4.7 is disqualified outright: it rationalized both planted musts as correct code. Late addition — the standout: on the Phase-1 discovery breadth seat, glm-5.3 reproduced ~88% of the real opus fleet's repo-derivable synthesis — including both items the build later proved load-bearing — for about ten cents.
Each review seat in the pipeline, stamped with where GLM lands after seven cells. This is what the policy flag enables — nothing more.
Bar = share of the ground-truth findings GLM independently re-derived (substance match). Dots = the musts: filled caught, hollow missed. Every GLM citation was re-opened against the sandbox tree — hallucination rate 0.00 in all fully-audited cells.
When GLM matches a finding, does it weigh it right? On legal/evidentiary surfaces the answer is no — it sees the bug and calls it smaller. This is why the confirming seat stays codex.
Quota burn measured off the z.ai meter during the runs; capacity assumes 7–8 GLM-relief passes per unit (3–4 fence confirms + ~4 loop iterates). Plan prices approximate — confirm on billing.
| Cell | Seat emulated | Surface | Model | GT | Substance | Must recall | Severity calls | Novel valid | Halluc. | Verdict |
|---|---|---|---|---|---|---|---|---|---|---|
| A | R1 fresh-eyes | MAP1a-2 whole-unit | glm-5.3 | 8 (2m/3s) | 1/8 | 1/1 explicit | — | 2 | 0.00 | additive lens; not replacement |
| B | confirm/middle | MAP1a-2 whole-unit | glm-5.3 | ~20 (6m) | 0/20 | 0/6 | — | 2 | 0.00 | no standalone |
| C | build.rote | MAP1a-2 S4 impl | glm-5.3 | n/a | ~90% fidelity | 2 fence-invisible defects | — | n/a | 0.00 | yes, with guardrails |
| Q1-47 | checkpoint confirm | CONSENT-ARTIFACT S9 | glm-4.7 | 6 (2m/4s) | 0/6 | 0/2 + rationalized both | — | 0 | 2 false claims | BANNED — false-clean |
| Q1-53 | checkpoint confirm | CONSENT-ARTIFACT S9 | glm-5.3 | 6 (2m/4s) | 5/6 (75% exact) | 2/2 content · 1/2 severity | 3 downgrades | 2 | 0.00 | yes-with-caveats |
| Q5 | 5.5-cell confirm | INT7-RT-SHEETS post-R1 | glm-5.3 | 4 (2m/2s·1 routed) | 2/3 gating | 1/2 | both correct | 2 | 0.00 | iterate yes · confirm no |
| Q2 | terra-cell confirm | CONSENT-ARTIFACT S8 | glm-5.3 | 7 at-sandbox (4m/3s) | 3.5/7 | 1/4 · 0/4 at severity | 2 downgrades | 1 strong | spot-only | additive; not consent confirm |
| D1-53 | discovery breadth | INT7-RT-SHEETS pre-discovery | glm-5.3 | 13 repo-derivable | 11.5/13 (88%) | n/a | — | TTLs via repo cross-ref | 0.00 (anchors cross-check GT) | STANDOUT — breadth parity |
| D1-47 | discovery breadth | INT7-RT-SHEETS pre-discovery | glm-4.7 | 13 repo-derivable | 5/13 (38%) | n/a | — | 0 | 3 anti-hits | invents schema vs migration:none |
Method. Each cell re-runs a real, already-shipped review round in a frozen git sandbox at the exact commit the original reviewer saw; ground truth = the findings that round actually produced, grounded in the fix commits. Reviewers get read-only tools; every tool call is audited (cheat-flags 0 across all cells). Scores: Q1/Q5 fully citation-audited against the sandbox tree; Q2 spot-audited (flagged). Q2's answer was recovered from its stream log after a runner collision. Policy state: burn_balance.glm.enabled=false — this page is the evidence for that flip, which also awaits the z.ai PRO upgrade. Protocol + raw ledgers: ~/glm-bench/results/ · scored 2026-08-17.