{"path":"research/forecasting-room-at-the-top-and-deliberus.md","content":"# Forecasting's Ceiling Question and Deliberus\n\n*Scott Alexander, \"Does Forecasting Have Room At The Top?\" (Astral Codex Ten, read 2026-08-03). Source: [astralcodexten.com/p/does-forecasting-have-room-at-the](https://www.astralcodexten.com/p/does-forecasting-have-room-at-the)*\n\n## What the post argues\n\nThe question: will AI forecasters plateau at human superforecaster level (Scenario A), or vastly exceed it — \"ultraforecasting\" (Scenario B)? Alexander contests Daniel Reeves' claim that prediction markets already sit near the optimum (only 3–6% better than simple statistical models), reframing those small margins as impressive relative to what the baselines already capture — and noting that sports, Reeves' main evidence base, are *deliberately engineered against predictability* (salary caps, drafts), so they understate the room available in geopolitical questions, where superforecasters beat baselines by much more.\n\nHis method is the interesting part: rather than argue abstractly, he triangulates the ceiling with three independent **anchors** — Reeves' market data, the human-vs-Samotsvety gap, and chess engines' growing pawn-handicap over the best humans — landing at 70–30 for Scenario B, but a modest B: roughly 4–12 percentage points of improvement over today's markets. Not miracles; real, measurable, meaningful margin. Two closing caveats matter most here: modern prediction markets show no measurable accuracy progress despite vastly greater volume, and *\"practical policy improvements might exceed headline accuracy gains through better unknown-unknown identification\"* — paired with the observation that a forecaster that much better at markets *\"would also be better at crafting good policy.\"*\n\n## Four connections to Deliberus\n\n### 1. It is the convergence wager rotated ninety degrees\n\nAlexander is measuring forecasting's irreducible remainder: how much uncertainty is **aleatoric** (no amount of cognition dissolves it) versus **epistemic** (better reasoning extracts it). The convergence wager asks the same question about *disagreement*: how much dissolves under decomposition, and what is left is the residue map. His three anchors do for prediction exactly what the residue map does for normative conflict — bound the ceiling empirically instead of asserting it, with the explicit possibility that the answer disappoints.\n\nTyped residues are a richer answer than forecasting has available. A Brier score can tell you *that* you hit a wall, never *what kind* of wall; fittingness vs structural vs axiom-choice vs permissive-zone can. And the accounting discipline is shared: just as the wager only counts if hinge/parity/permissivism ceilings score *against* convergence ([convergence.md](../convergence.md)), Alexander's estimate only counts because he takes seriously the scenario where the room is a fraction of a percent.\n\n### 2. His closing caveat gestures at the missing layer\n\nForecasting is **conclusions-layer** infrastructure in the missing-layer taxonomy ([the-missing-layer.md](../the-missing-layer.md)): markets price the WHAT and leave the WHY illegible. Alexander's own ending points past the layer he is analyzing — the practical value of a better forecaster runs through *unknown-unknown identification* and *policy-crafting*, and both transfers run through reasoning, not scores. A bare \"54%\" is nearly inert for action; the decomposed considerations behind it — the implicit premises, the cruxes, what evidence would move the number — are what a decision-maker can actually use, and that is precisely the structure Deliberus extracts and makes challengeable. When the leading essayist of the forecasting-adjacent world concludes that the accuracy number undersells the value and the reasoning-quality overflow is where the payoff lives, that is the conclusions-layer testifying, from inside, that the reasoning layer is the bottleneck.\n\nThe forward-looking version: if Scenario B arrives, AI ultraforecasters (FutureSearch-style systems already produce rationale trees) shift the bottleneck from *accuracy* to *auditable reasoning substrate* — you will not trust or act on a machine's 62% without inspectable, attackable structure behind it. That is the same opening the LawZero read documented on the safety side ([bengio-safety-from-honesty-and-deliberus.md](bengio-safety-from-honesty-and-deliberus.md)): machine-side epistemics converging on the need for exactly this layer.\n\n### 3. The demand-skeptic tell\n\nAlexander is the named source of the hardest demand-side critique the SFF application confesses. In this post, he says a measurable move from 50% to 54% would leave him *\"very excited.\"* What converts this skeptic-shape is not vision — it is a **small gain made measurable**. That quietly validates the instrument strategy: disagreement-preservation scores, the completeness oracle, hinge, and the residue fraction are Deliberus's Brier score, the apparatus that makes its value legible to precisely this audience.\n\n*Amended 2026-08-17.* This paragraph used to continue: *\"the pre-registered group evaluations ('do people reason demonstrably better on this substrate?') are the equivalent of the calibration record that made forecasting respectable.\"* The evaluations are still planned, but the founder rejected that framing on 2026-08-16 as the wrong **unit** — a commons is not evaluated by what one reader gains, and [lowering-the-cost.md](lowering-the-cost.md) §5 had already argued it formally. The analogy survives in better shape without the individualist measure, because the instruments listed above *are themselves* the calibration record: they are corpus-level, they are published, and they can read badly. Which is exactly what made forecasting respectable — not that any one forecaster improved, but that the scoreboard was public and could embarrass you.\n\n### 4. A methodological rhyme: bounding the ceiling by triangulation\n\nHis chess-handicap anchor — estimate distance-from-optimal by measuring the handicap a stronger player can give — is structurally the move run 3F made ([frontier-extraction-experiment.md](frontier-extraction-experiment.md)): a frontier model executing the pipeline by hand as the quality *ceiling*, the production Flash pipeline as *current*, the gap as the handicap. Same epistemic maneuver, same honesty payoff: you learn how much room is above you without needing to reach it first.\n\n## The open fork (NOT ratified)\n\n**Candidate public framing: \"the residue map measures the aleatoric fraction of disagreement.\"** The aleatoric/epistemic split is an already-legible frame for the EA/rationalist/forecasting audience, and it states the wager's falsifiability in one line — the map exists to find out how much of moral conflict is \"chance-like\" (irreducible under the current move-set) versus extractable by better reasoning.\n\nThe counter-argument, at full strength: aleatoric uncertainty is about *chance*, and a fittingness residue is not chance — it is a determinate, arguable-in-a-different-register normative question that decomposition has made maximally legible. Importing a probabilistic frame could miscast typed residues as noise, exactly the flattening the ontology exists to resist, and could smuggle in the implication that residues are *permanently* fixed when a terminus is only ever a fallible fixed point under the current move-set. Founder decision pending; until ratified, the framing appears nowhere public-facing.\n\n## Honest limits of the analogy\n\nForecasting questions have ground truth arriving on a date; normative disagreements do not, which is why Deliberus needs verdict-claims and challengeable classifications where forecasting needs only resolution criteria. Forecast accuracy is a scalar; residue typing is categorical, so \"how much room\" translates only loosely. And Alexander's post is about a *machine-vs-human* ceiling, while the wager is about a *decomposition-vs-disagreement* floor — the rotation is illuminating, not an identity.\n\n## What this changes and does not change\n\nIt changes nothing in the architecture; every instrument this read validates already ships. It adds: (a) a third independent arrival at the reasoning-layer gap, from forecasting's own leading voice, alongside LawZero (machine safety) and the science-infrastructure lineage — strengthening the missing-layer positioning with a source the target audience already trusts; (b) the registered open fork on the aleatoric framing; (c) a named kinship (ceiling-bounding by triangulation) between run 3F's method and an accepted epistemic practice. It does not license outreach moves — the dogfooding gate holds.\n\n---\n\n**See also**: [convergence.md](../convergence.md) · [the-missing-layer.md](../the-missing-layer.md) · [bengio-safety-from-honesty-and-deliberus.md](bengio-safety-from-honesty-and-deliberus.md) · [frontier-extraction-experiment.md](frontier-extraction-experiment.md) · [convergence-wager-red-team.md](convergence-wager-red-team.md)\n"}