{"path":"research/peer-review-and-the-reasoning-layer.md","content":"# Peer Review and the Reasoning Layer\n\n*What thirty-five years of reform did to peer review, the one function it never touched, and the eight places Deliberus could act. Written Aug 2026, prompted by a referee report its author had to publish on social media to make public.*\n\nPeer review is the **quality-control** row of [the missing-layer taxonomy](../the-missing-layer.md): infrastructure that stamps and annotates finished text without holding the structure of what it checks. This doc works out what that costs and what could be done about it. The survey of science's other infrastructure attempts — nanopublications, scite, ORKG, discourse graphs, Registered Reports — is in [science's missing reasoning layer](science-reasoning-infrastructure.md); the philosophical lineage is in [the science-lineage study](science-lineage-philosophy.md).\n\n---\n\n## The specimen\n\nIn August 2026 a mathematician and AI-risk researcher posted his own referee report to Facebook. The review was double-blind, he judged the manuscript publishable, and he objected to one thing: a coined term of the form *\"structural X-alignment\"*, where *X-alignment* already has an established meaning in the field. His argument was grammatical before it was technical. By the ordinary rules of English compounding, readers will take the coinage to name a *kind of* the established thing. It does not: the established term is about an AI system's own values, the coinage is about how a system reasons about human values. He placed the objection in a wider frame — the century-long practice of coining terms that piggyback on prestigious neighbors — and then made it a safety argument: one route to catastrophe runs through developers confusing *having accurate knowledge of human values* with *being aligned to them*, and a term that implies the first is a species of the second promotes exactly that confusion.\n\nFive different kinds of move sit in that one letter:\n\n| Move | Kind |\n|---|---|\n| \"The manuscript is now publishable as it stands\" | verdict |\n| The compound implies a subsumption relation that does not hold | definitional attack |\n| Confusing knowledge-of-values with alignment-of-values is a catastrophe pathway | causal / mechanism claim |\n| The author has a duty to try to avoid contributing to that | normative claim |\n| Researchers routinely coin terms that piggyback on prestigious ideas | field-level sociological observation |\n\nTo any machine, and to most human readers, those five are indistinguishable: one blob of reviewer prose. The author can only respond to the letter, not to the moves. And the reasoning reached one editor, one author, and whoever happened to scroll past a screenshot.\n\n## The evolution: a long unbundling\n\n**It is far younger than its authority implies.** *Philosophical Transactions* (1665) is the usual origin story, but that is partly retrospective myth-making: early journals ran on a single editor's judgment, and the Royal Society's Committee on Papers (1752) is the better-documented ancestor. Einstein's 1905 papers were not peer reviewed in the modern sense. When *Physical Review* sent his 1936 gravitational-waves paper (with Rosen) to a referee, he was affronted enough to withdraw it and never publish there again — the referee, Robertson, was right and Einstein was wrong. **Nature** had no formal external peer review until **1967**; *The Lancet* until 1976. The gold standard is younger than television.\n\n**Post-war institutionalization** came from four pressures at once: submission volume, photocopying making distribution cheap, government funding demanding accountability, and specialization outrunning any single editor's competence.\n\n**Then the empirical critique.** Peters and Ceci (1982) resubmitted twelve already-published psychology papers under fictitious names and institutions; eight of nine were rejected. The BMJ-linked error-insertion studies planted major errors in manuscripts and found reviewers catching only a minority. Cochrane's review states that editorial peer review is used worldwide with \"little empirical evidence\" that it ensures the quality of biomedical reports. The honest reading is not that peer review is useless — it catches things, improves manuscripts, and supplies social accountability — but that it is an **opaque institutional process rather than a recomputable representation of reasoning**.\n\n**Then reform, function by function.** Read as a series, every major change of the last thirty-five years pried one bundled job loose from the others:\n\n| Year | Move | What it separated |\n|---|---|---|\n| 1991 | arXiv | dissemination from certification |\n| 2006 | PLOS ONE | soundness from importance (\"is it correct\" before \"does it matter\") |\n| 2012 | PubPeer | error-catching from pre-publication timing |\n| 2013 | Registered Reports (Chambers, *Cortex*) | protocol quality from result attractiveness |\n| 2013– | OpenReview | the review record from private correspondence |\n| 2023 | eLife's model | assessment from accept/reject |\n| ongoing | PREreview | reviewing from journal gatekeeping |\n\nEach made criticism more **visible**. None made it **structured**. The reviewer's argument is still welded shut inside prose, which is why the best distribution channel available to a referee with an important objection is a screenshot on a social network.\n\n**And the 2020s squeeze:** submissions growing far faster than the reviewer pool, LLM-written reviews as a live policy fight, and authors embedding hidden prompt injections in preprints to farm favorable machine reviews.\n\n## Eight places Deliberus could act\n\nOrdered by fit. Shipped instruments are marked, because several of these need no new machinery.\n\n### 1. Terminology objections become reusable artifacts\n\nThe sharpest fit, and the specimen is the argument for it. \"Two senses under one term, plus a false is-a-kind-of relation manufactured by grammatical form\" is a precise description of the [contested-concept layer](../worldview-lenses.md). Represented once — a definitional claim pinning the established sense, a sibling sense node, an explicit not-a-subtype-of relation — the objection is inherited by the next paper that coins the same shape, instead of requiring a second referee to notice it from scratch. This is the `@[simp]` flywheel applied to vocabulary.\n\nThe safety framing is the reviewer's own, not ours: he argues the confusion carries catastrophic downstream risk. On that framing a terminology-drift instrument is an alignment instrument, not pedantry.\n\n### 2. Reviewer effort stops evaporating\n\nThe largest efficiency claim. A strong report reaches one editor and up to three authors, then is destroyed. Every rejected paper's review is total loss. Reviewers at different venues independently re-derive the same objection, and none of them can see the others.\n\nAttach the objection to the **claim** instead of to the **manuscript** and the next submission of that claim inherits the existing attack surface. This is the general \"wrong moment\" lesson from the cross-case pattern, applied: intervene after any text exists, not at publication time.\n\n### 3. Uptake becomes measurable\n\nLongino's second condition for social objectivity is that criticism must be *able to change the state of inquiry*. Peer review is where that is least visible. The specimen opens by noting the author \"has done a serious job of revising the manuscript in the light of my previous report\" — real uptake, observed by exactly one person, recorded nowhere.\n\nComputable version: challenge → claim-strength change → whether the challenged node was decomposed, conceded, or defended. This attacks reviewing's incentive problem at the root, because it makes reviewing legible as **contribution** rather than as unpaid invisible labor. Already named as a buildable counter-instrument in [the red-team synthesis](red-team-synthesis-2026-07.md).\n\n### 4. Pre-submission self-review *(shipped)*\n\nRun the [completeness oracle](../depth.md) on your own draft: unsupported value premises, unanswered critical questions, the open sorry-frontier. Walton's schemes already *are* field-specific review checklists — an expert-opinion argument owes credibility, domain, and consistency checks; a causal claim owes mechanism, confounder, and alternative-explanation checks.\n\nAuthors would arrive with the cheap objections handled, so scarce human attention goes where only a human expert can help. **This requires no institution to change anything**, which makes it the correct entry point.\n\n### 5. Split reviewers get localized *(shipped)*\n\nReviewer 1 accepts, Reviewer 2 rejects, and an editor resolves it in prose. The [discursive-dilemma flag](red-team-synthesis-2026-07.md) is built for exactly that shape: premise-level agreement coexisting with conclusion-level disagreement. The hinge score names the single contested node the decision actually turns on. An editor currently reconstructs that by hand, badly, under time pressure.\n\n### 6. Registered Reports have a native representation here\n\nThis identification appears to be new. Reviewing a protocol before results exist *is* reviewing a reasoning structure with a declared empty evidence slot — which is precisely a **sorry marker**. Preregistration becomes a graph with open sorry nodes; the results phase fills them; deviation from protocol becomes a visible graph diff rather than a buried paragraph in a discussion section.\n\nRegistered Reports is the most structurally sound reform anyone has shipped, and it turns out to be asking for the object Deliberus already has.\n\n### 7. Retraction that propagates\n\nToday a retracted paper keeps being cited for years and nothing downstream updates. With typed support edges, a falling foundational claim weakens everything leaning on it through the same published gradual semantics used everywhere else in the graph. scite typed citation contexts; this makes them load-bearing.\n\n### 8. Reviewer matching by graph position\n\nWho has already reasoned about *this premise*, rather than who matches these keywords. A claim graph knows who attacked, supported, or decomposed neighboring claims — including, deliberately, who disagrees productively. Listed last because it is the most exposed to the objection below.\n\n## Why this domain is harder than the others\n\n**Peer review sits at the institution rung of [the scale ladder](fractal-scales-and-temporal-frame.md)**, and each rung up adds an adversary. Here the adversary is capture and gaming, and it bites in a specific way: **a legible review record is also a weaponizable one.** Anonymity currently protects junior reviewers criticizing senior authors, and permanently attributed, computable reasoning would chill exactly the reports most worth having. That is constructive ambiguity in its sharpest real form — some criticism survives only because it is deniable. Any design here owes an answer to \"who is protected by the current opacity?\" before it starts removing it.\n\n**The convergence-illusion risk is acute.** Referee reports are often deliberately sharp. The specimen is grumpy, invokes professional duty, and *pleads*. An extraction pass that renders it as \"the reviewer suggests alternative terminology\" has destroyed the object it was supposed to preserve. [Disagreement preservation](red-team-synthesis-2026-07.md) is the gate, not a nicety.\n\n**The field already has an LLM-review problem.** Reviews written by models, and hidden prompt injections planted to farm favorable ones. Deliberus must not become a machine for generating plausible-looking objections at volume, which is the obfuscated-argument threat wearing reviewing's clothes.\n\n**And Longino's limit binds hardest of all:** uptake in the graph does not guarantee uptake in institutions. Deliberus can show that a claim weakened. It cannot make a journal care.\n\n## Run 5: the extraction, executed by hand (Aug 12, 2026)\n\nThe pipeline path needs an authenticated submission, which is a human-only step, so this\nwas executed the way [run 3F](frontier-extraction-experiment.md) was during the Gemini\noutage: a frontier model performing all eight passes by hand on the real text, producing\na quality ceiling the pipeline can later be compared against. Source: the specimen report\nabove, ~450 words of referee prose.\n\n**Source handling.** The report was posted by its author to a friends-audience social feed,\nand he explicitly weighed the anonymity of a live double-blind review before sharing it\n*with friends*. Republishing it verbatim on a public site would change that calculus\nwithout his consent, and it fails the Reader Test. So this section records the extraction\nstructure and findings with the author unnamed, the coinage generalized to\n*structural X-alignment*, and claims stated analytically rather than quoted. The verbatim\nartifact belongs under `.private/` and has not been written anywhere else.\n\n### The passes\n\nPass 1 found five arguments: the verdict, the definitional attack, the field-practice\nobservation, the catastrophe argument, and the duty claim. Pass 2 decomposed them into\n**22 atomic claims** — 8 definitional, 6 empirical, 5 normative, 2 value premises, and one\nthat fits none of the four types (see finding R3). Pass 3 produced 14 support edges, 2\nqualifies edges, and **zero attack edges**. Pass 4 found the contested concept the whole\nreport is about: one term carrying two senses, contested in the *criterion* sense.\n\nThe load-bearing chain, which is where the report is strongest:\n\n> *X-alignment* denotes alignment of a system's own values (definitional) · the coinage\n> denotes how a system reasons about human values (definitional) · those are distinct\n> properties (definitional) · English compounds denote subtypes of the term they modify\n> (definitional, metalinguistic) → readers will infer a subtype relation (empirical\n> prediction) + the inference is false (definitional consequence) → **the term is\n> misleading** (evaluative) → it should be replaced (normative)\n\n### Scoring the registered predictions\n\n**Prediction 1 — the definitional attack extracts well, the sociological observation\nextracts badly. CONFIRMED.** The definitional chain above decomposed cleanly, every\npremise typed unambiguously, every edge determinate. The field-practice observation did\nthe opposite: it yielded three claims about a professional culture with no evidence\noffered, no scheme fitting them, no support edges, and no critical question to attach —\ntechnically extracted, epistemically orphaned. Hedged meta-observation remains the\nhardest register the pipeline faces.\n\n**Prediction 2 — the duty claim surfaces as an unsupported value premise. CONFIRMED, and\nthere are two of them.** \"The author has a duty to try to avoid contributing to such a\ncatastrophe\" and \"the established sense is tremendously important\" are both asserted\nvalues with no support subgraph. Both are also short and interesting to descend: does a\nduty to avoid contributing scale with the size of one's contribution, and is a\nterminological choice a contribution at all? That is exactly the descent the\nno-copout-axioms check exists to open.\n\n### Findings the predictions missed\n\n**R1. The five-move hypothesis holds, and there is a sixth move.** Alongside verdict,\ndefinitional attack, causal-harm mechanism, normative duty and field observation sits a\n**speaker-attitude hedge**: the reviewer declining, in advance, to be read as waging a\ncampaign. It is not a claim about the world, not a verdict, and not a normative claim\nabout the author — it is a claim about the reviewer's own stance, doing real argumentative\nwork by pre-empting the objection that he is being pedantic. The four-type taxonomy\n(empirical / definitional / normative / value premise) has no slot for it, and dropping it\nloses an anticipated-objection move. This is a taxonomy gap distinct from the\n[ought-bias found in the deep-descent experiment](deep-descent-experiment.md): that one\nwas about residue types, this one is about claim types.\n\n**R2. The core argument runs on a metalinguistic premise, which is a register the corpus\nhas never held.** \"Modifier-plus-term compounds denote subtypes\" is a claim about how\nEnglish works, and the entire objection rests on it. Every argument the graph has\nextracted so far has been empirical, causal, or evaluative. A premise about language is\nnew — and unusually tractable, because it is an empirically testable claim about reader\nparsing rather than a value posit. If the reasoning layer is going to serve terminology\ndisputes, this register has to be first-class.\n\n**R3. Verdict and critique are orthogonal — zero attack edges in a critical review.** The\nreport says both \"publishable as it stands\" and \"please change this term\", and nothing in\nthe structure connects them adversarially. A pipeline that assumed critique implies\nrejection would mis-model this badly. What is striking is that this is precisely the\nseparation eLife's 2023 model made institutional, appearing spontaneously in ordinary\nreferee practice: reviewers already think assessment and accept/reject are different\nobjects, and the paperwork was the thing forcing them together.\n\n**R4. The catastrophe argument is Argument from Consequences with the magnitude left\nunstated.** \"If the term catches on, it will promote that confusion\" carries no\nprobability, and the reviewer's own \"I fear that\" marks it as low confidence. The critical\nquestions for that scheme would ask how likely adoption is and how likely the confusion\nfollows — so the hinge of the strongest part of the report is a quantity nobody wrote\ndown. That is [run 3F's headline finding](frontier-extraction-experiment.md) — the crux is\nimplicit — recurring inside a *single-author* text rather than across a debate pair.\n\n**R5. Referee prose is roughly three times denser than debate prose.** 22 atomic claims\nfrom ~450 words is about one claim per 20 words, well above the run-3F debate texts. Review\nwriting is compressed argument by professional habit, which makes it unusually good\nextraction material and is independent support for treating this as a high-fit surface.\n\n### What the pipeline still has to be checked against\n\nThis is the ceiling, not the measurement. When an authenticated run is possible, the\ncomparison is: does the pipeline recover the metalinguistic premise (R2), does it avoid\ninventing an attack edge between verdict and critique (R3), does it flag the missing\nprobability (R4), and what does it do with the speaker-attitude hedge (R1) — drop it,\nor mis-type it as normative?\n\n---\n\n**See also**: [science-reasoning-infrastructure.md](science-reasoning-infrastructure.md) · [science-lineage-philosophy.md](science-lineage-philosophy.md) · [the-missing-layer.md](../the-missing-layer.md) · [fractal-scales-and-temporal-frame.md](fractal-scales-and-temporal-frame.md) · [red-team-synthesis-2026-07.md](red-team-synthesis-2026-07.md) · [depth.md](../depth.md)\n\n\n**Update (2026-08-20)**: the demand side got desperate — submission floods (NeurIPS doubled, arXiv slop-rejections 4%→10–12%, OSF's generalist server closed to submissions), reviews themselves increasingly machine-written, and a major-press publishing head calling for \"radical change.\" Urgency evidence for the pre-submission entry point: [ai-slop-and-the-reasoning-layer.md](ai-slop-and-the-reasoning-layer.md).\n"}