{"path":"research/steadying-an-outside-judge.md","content":"# Steadying an Outside Judge\n\n**Date**: 2026-08-27 · **Occasion**: the founder, after listening to a podcast on the judge-stability paper (*Jagged Judges*, arXiv:2608.12645): *\"I definitely was thinking that Deliberus should be able to help any judge whether AI or human stay much more levelheaded.\"* · **Status**: mechanisms named, one of them new to the corpus and testable; nothing built.\n\n## What the corpus already had, and the half it was missing\n\n[fragile-checkers-and-the-verification-bottleneck.md](fragile-checkers-and-the-verification-bottleneck.md) answers the **defensive** question: Deliberus's *own* judges do not wiggle because verdicts live as artifacts in a record rather than as states in a channel, so pressure has to write itself down. And the **human** side is in the corpus too, from the debate-alignment literature: structured argument makes a human judge more reliable, with measured support ([deliberus-as-alignment-infrastructure.md](deliberus-as-alignment-infrastructure.md) §; [ai-safety-and-tao-augmentation-research.md](ai-safety-and-tao-augmentation-research.md) — *\"the structured courtroom that makes human judgment more reliable\"*).\n\n**What is absent: the outside AI judge.** Nothing in the corpus asks what a *third party's* model — a grader, a moderator, an evaluator in someone else's system — gains from being able to query this graph while under pressure. That is the founder's question, and it has a sharper answer than the general \"structure helps\" claim.\n\n## Four mechanisms, tactic by tactic\n\nThe paper decomposes pressure into named tactics. Taking them one at a time is what makes the answer concrete rather than a hope.\n\n**1. Fabricated consensus — the strongest measured tactic, and the one that dissolves.** *\"Three experts disagree with you\"* broke a model that was rock-stable against doubt, counterargument and single-expert authority: **60.3% of its verdicts flipped.** It works because the claim is **unfalsifiable inside the conversation** — the judge can comply or resist, and has nothing to check against. Against a persistent attributed record it stops being pressure and becomes a **lookup**: do those three experts hold that view, in the record, with attribution? The corpus already names this for its own judges; the new part is that **the same lookup is available to a judge that is not ours**, and it needs no write access.\n\n**2. The counterargument — new information, or a known move?** A judge folds partly because it cannot tell whether an objection is fresh or already considered and answered. In a graph, an objection that exists as a claim with attacks on it and a badge is **visibly a known move**. Pressure that repeats a settled objection stops resembling evidence. This is the dispute structure the corpus insists is the non-commodity part, doing work for a reader who is not a person.\n\n**3. The correction channel, kept without the corruption.** The paper's uncomfortable finding is that **30–44% of flips are corrective** — sometimes pressure fixes a wrong verdict — and Deliberus's record-mediated design gives that up along with the 56–70% that corrupt. An outside judge querying the graph gets a third option: **corrected by the record without being persuadable by the interlocutor.** It can discover it was wrong from an attributed claim rather than from someone insisting.\n\n**4. Fragility prediction, which is the genuinely new proposal.** The paper's best cheap predictor of which verdicts will wiggle was **jury majority strength across model families** (mean |ρ| = 0.59, beating repeat-consistency at 0.42): *items a jury splits on are the items any single judge folds on*. Deliberus computes structurally similar signals from **accumulated human disagreement** rather than from a model panel — balanced attack-and-support energy, a high hinge, an unclassified terminus, a `does_not_fit` scheme. So the graph could tell a judge **in advance which of its verdicts sit on contested ground**, and it is cheaper than running nine models.\n\n**One respect in which it is better grounded than a model jury**, and the paper supplies the reason itself: its independence condition. Two samples of one model are one juror sampled twice, and the corpus's own anchor is 18 of 30 agents choosing an identical branch name. A split among *people* is not correlated the way a split among models is.\n\n**And the honest caveat, which must travel with it**: the corpus's split reflects **what was contributed**, not what is genuinely contested — the coherent-absence problem ([the-residual-error-taxonomy.md](the-residual-error-taxonomy.md) class 2). A claim nobody attacked may be uncontroversial or may be unexamined, and the graph cannot currently tell those apart. So this is a proxy with a named bias, and the gray band is what keeps it honest.\n\n## Why this composes with the read-only ruling rather than straining it\n\nEvery mechanism above is a **read**. None needs an agent write surface, which the corpus closed for reasons unrelated to this (sybil resistance, source independence, machine-speed selective legibility, and the finding that agents reading a shared public board become correlated by construction). **A judge that reads to check a consensus claim is not writing to the board**, so the argument for read-only is untouched — and this is, so far, the sharpest concrete answer to *what an agent reader would actually use the graph for*.\n\nIt also lands on the fact layer rather than the values layer, which is where the corpus already ruled the first agent-facing tools belong: whether three experts hold a view is checkable; whether they are right is not.\n\n## What would test it\n\nThe cheapest version needs no humans, which makes it a sibling of the substrate run in [agents-as-a-consumer-class.md](agents-as-a-consumer-class.md) §3a: run the paper's own pressure ladder against a judge on claims **drawn from the graph**, twice — once with the model given a graph query tool, once without — and measure the flip rate under fabricated consensus. A drop is the mechanism working. **The prediction to register before running it**: mechanism 1 shows the largest effect, because fabricated consensus is the tactic whose whole power is unfalsifiability, and it is the one a record removes entirely.\n\n## Cross-references\n\n[fragile-checkers-and-the-verification-bottleneck.md](fragile-checkers-and-the-verification-bottleneck.md) (the defensive half, and the paper's numbers) · [agents-as-a-consumer-class.md](agents-as-a-consumer-class.md) (read-only; the dispute-structure argument; the substrate run) · [deliberus-as-alignment-infrastructure.md](deliberus-as-alignment-infrastructure.md) + [ai-safety-and-tao-augmentation-research.md](ai-safety-and-tao-augmentation-research.md) (the human-judge half, already in the corpus) · [the-residual-error-taxonomy.md](the-residual-error-taxonomy.md) (why an unattacked claim is ambiguous) · [the-scrutiny-gap.md](the-scrutiny-gap.md) (agent correlation via a public board)\n"}