{"path":"research/ranking-reasoning-quality-without-bedrock.md","content":"# Ranking Reasoning Quality Without Bedrock\n\n**Date:** 2026-07-05\n**Status:** AI-conducted research (web survey of primary and academic sources, conducted by a Claude research agent for the Deliberus project; all substantive claims cited)\n\n---\n\n## The claim under test\n\nDeliberus's convergence wager rests partly on a load-bearing claim: that a deliberation platform can help resolve *which views are better-informed and better-reasoned than others* — even in normative and worldview domains where no verifiable endpoint exists. If \"better-reasoned\" is just a compliment we pay to arguments whose conclusions we already accept, the claim is circular and the wager collapses. This document tests whether \"better-reasoned\" can be made rigorous and non-circular, in four moves:\n\n1. **Operationalization** — does argumentation theory yield a workable, non-arbitrary metric?\n2. **Empirical validation** — does reasoning quality predict accuracy where oracles *do* exist?\n3. **Adversarial stress-testing** — what do AI debate and scalable-oversight experiments show about argument-checking under optimization pressure?\n4. **Foundations** — is there any non-circular ground for ranking reasoning norms themselves, and if not, what is the best available response?\n\nThe short answer, defended below: **\"better-reasoned\" survives as a rankable, empirically validated, axiomatically constrained property — but only as a defeasible, procedural verdict, never as a truth-oracle.** The circularity hole is real and bottomless; the correct response is not to fill it but to widen the circle and make the scoring norms themselves inspectable and contestable.\n\n---\n\n## 1. Argumentation theory delivers an operational metric — with known limits\n\n### Schemes and critical questions\n\nDouglas Walton's argumentation schemes formalize everyday defeasible reasoning as stereotyped inference patterns (argument from expert opinion, from consequences, from analogy, and so on), each paired with **critical questions (CQs)** that enumerate the specific ways an instance of that pattern can fail ([Walton, Reed & Macagno, *Argumentation Schemes*, Cambridge](https://www.cambridge.org/core/books/argumentation-schemes/9AE7E4E6ABDE690565442B2BD516A8B6)). This converts \"is this a good argument?\" from a matter of taste into a checklist of typed, answerable challenges: an argument from expert opinion is only as strong as the expert's credibility, field-match, and consistency with other experts — each a CQ, each independently investigable.\n\nTwo honest caveats from the theoretical literature. First, in the classical treatment, CQ-based evaluation is **qualitative and binary**: one unsatisfactorily answered question defeats the argument, which is too coarse for ranking ([Yu & Zenker, \"Schemes, Critical Questions, and Complete Argument Evaluation,\" *Argumentation* 2020](https://link.springer.com/article/10.1007/s10503-020-09512-4)). Second, the scheme–CQ linkage itself lacks a unified theory; recent work argues CQs should be partly disentangled from schemes and reconstructed on independent grounds ([*Argumentation* 2023](https://link.springer.com/article/10.1007/s10503-023-09613-w)), and CQ generation is now an active NLP research problem in its own right ([Calvo Figueras & Agerri, arXiv:2410.14335](https://arxiv.org/pdf/2410.14335)). Schemes and CQs are a strong *skeleton* for quality assessment, not a finished metric. *(Confidence: high.)*\n\n### Bayesian argumentation: fallacies rehabilitated as graded evidence\n\nThe decisive theoretical upgrade comes from Ulrike Hahn and Mike Oaksford's Bayesian reanalysis of the informal fallacies ([*Psychological Review* 2007](https://pubmed.ncbi.nlm.nih.gov/17638503/); [*Synthese* 2006](https://philpapers.org/rec/HAHABA-3)). Classical fallacy theory treated argument from ignorance, circularity, and slippery slopes as *formally* defective. Hahn and Oaksford showed that each form has both strong and weak instances depending on **content** — the same structure that is fallacious in one context is compelling in another (\"this drug has passed extensive safety trials without adverse findings\" is an argument from ignorance, and a good one). Argument strength becomes a matter of Bayesian updating: priors, likelihoods, evidence reliability. Crucially, their experiments found that **ordinary people's strength judgments track the Bayesian factors** — sensitivity to prior belief, evidence strength, and source reliability — meaning the normative model and human intuition are not at war ([Hahn, Harris & Oaksford, \"Rational argument, rational inference,\" *Argument & Computation* 2013](https://doi.org/10.1080/19462166.2012.689327)).\n\nThis is the rehabilitation of fallacy theory: \"fallacy\" stops being a rule-violation label and becomes a region of low evidential strength on a continuous scale. Hahn and Hornikx have argued that the most fruitful normative model of argument quality **combines the scheme approach with Bayesian probability** — schemes supply the typed structure and the CQ checklist; Bayesian strength supplies the graded verdict. *(Confidence: high on the theory; the combination as engineering practice is younger.)*\n\n### Gradual semantics and QBAFs: principled score propagation\n\nQuantitative Bipolar Argumentation Frameworks (QBAFs) — graphs of arguments with attack and support relations plus base scores — provide the machinery for propagating strength through a debate. The pivotal result is that this propagation need not be ad hoc: Baroni, Rago and Toni systematized the space of desirable properties (balance, monotonicity, and their variants) into a principled spectrum that candidate semantics can be checked against ([*International Journal of Approximate Reasoning* 2019](https://www.sciencedirect.com/science/article/pii/S0888613X18304651)). Potyka's Quadratic Energy Model — the gradual semantics Deliberus uses — was designed precisely to satisfy the full principle set, including duality (symmetric treatment of attack and support) and open-mindedness, and handles cyclic graphs via continuous dynamics ([Potyka, KR 2018; tutorial at arXiv:1811.12787](https://arxiv.org/pdf/1811.12787)).\n\nTwo facts matter for the rigor question. First, the axioms genuinely **constrain**: many intuitive-seeming semantics violate them (Df-QuAD saturates; Euler-based semantics treats attack and support asymmetrically), so the choice of semantics is not arbitrary. Second, the axioms **do not uniquely determine** a semantics — they define a family, and picking within it is a convention. A QBAF score is therefore best read as: *given this graph of claims, attacks, supports, and evidence, and given propagation principles that any reasonable person should accept in the abstract, this is the current dialectical strength.* The same research group has already carried this machinery into judgmental forecasting, where argumentation frameworks are used to structure probability estimates and flag irrational updating ([Irwin, Rago & Toni, \"Forecasting Argumentation Frameworks,\" KR 2022](https://arxiv.org/abs/2205.11590)) — a direct bridge between gradual semantics and the accuracy literature in §2. *(Confidence: high.)*\n\n### Natural-language argument quality can be annotated — imperfectly\n\nFor arguments in the wild, Wachsmuth et al.'s taxonomy decomposes quality into 15 dimensions across three layers — **logical** (cogency: premise acceptability, relevance, sufficiency), **rhetorical** (effectiveness), and **dialectical** (reasonableness in the context of the wider debate) — and showed that theory-based expert annotation and untrained practical assessment largely converge ([EACL 2017](https://aclanthology.org/E17-1017/)). Recent work finds LLMs produce consistent annotations with moderately high agreement with human experts across most of these dimensions ([Mirzakhmedova et al., \"Are Large Language Models Reliable Argument Quality Annotators?\", RATIO 2024](https://arxiv.org/abs/2404.09696)).\n\n**Verdict on (a):** Argumentation theory gets surprisingly far. A quality metric built from scheme-typed decomposition + CQ coverage + Bayesian-style evidence sensitivity + principled gradual propagation is operational, non-arbitrary (axiomatically constrained), and empirically anchored to expert human judgment. What it measures, however, is the state of a *represented debate* — how well a position survives the challenges currently on the table — not the state of the world. That gap is what §2 and §4 are about.\n\n---\n\n## 2. The empirical bridge: reasoning quality tracks accuracy where we can check\n\nThe key epistemological move available to Deliberus is this: if independently-specified reasoning-quality properties predict accuracy in domains *with* oracles, that provides a defeasible license to trust the same properties in domains *without* them. The evidence that they do is substantial.\n\n### Forecasting tournaments\n\nIn IARPA's multi-year geopolitical forecasting tournaments, the Good Judgment Project identified stable individual differences in accuracy. Skill was real, not luck: year-to-year accuracy correlated at ~0.65, and a structural model combining dispositional, situational, and behavioral variables achieved a multiple correlation of 0.64 with accuracy ([Mellers et al., \"Identifying and Cultivating Superforecasters,\" *Perspectives on Psychological Science* 2015](https://journals.sagepub.com/doi/abs/10.1177/1745691615577794); [AI Impacts summary of GJP evidence](https://aiimpacts.org/evidence-on-good-forecasting-practices-from-the-good-judgment-project/)). The predictors are recognizably *reasoning-quality* variables: actively open-minded thinking, cognitive reflection, use of comparison classes and base rates, granular probability use, frequent belief updating — and, importantly, **trainable** ones: a one-hour probabilistic-reasoning training module produced accuracy gains persisting at least a year, and teaming (sharing rationales) improved accuracy further. Reasoning style is not just correlated with accuracy; interventions on reasoning style *cause* accuracy improvements. *(Confidence: high.)*\n\n### Rationale text predicts accuracy\n\nThe most direct bridge study: Karvetski, Mellers, Tetlock and colleagues applied NLP to the **written rationales** of forecasters across multiple tournaments and found that textual reasoning properties — higher dialectical complexity (weighing arguments on both sides), more comparison classes, more past-oriented (base-rate) reasoning — distinguish top forecasters, and that accuracy-boosting interventions push rationale profiles in the same direction ([Karvetski et al., \"What do forecasting rationales reveal about thinking patterns of top geopolitical forecasters?\", *International Journal of Forecasting* 2022](https://www.sciencedirect.com/science/article/abs/pii/S0169207021001473); [SSRN preprint](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3779404)). This is the crucial existence proof for Deliberus's premise: **reasoning quality visible in the text of an argument carries signal about whether its author's conclusions turn out true.** *(Confidence: high that the correlation exists; moderate on its size — these are meaningful but not overwhelming correlations, and they are population-level, not verdicts on individual arguments.)*\n\n### Debate: the correct side argues better\n\nIn human-debate experiments on hard reading-comprehension questions, judges who never saw the source text reached 84% accuracy from adversarial debate between one honest and one dishonest expert, versus 74% for single-advocate \"consultancy\" ([Michael et al., \"Debate Helps Supervise Unreliable Experts,\" 2023](https://arxiv.org/abs/2311.08702)). With LLM debaters, optimizing debaters for *persuasiveness alone* — with no truth label in the loop — **increased** judges' ability to identify true answers, and Elo ratings showed debaters assigned the correct answer were systematically more persuasive than those assigned the incorrect one ([Khan et al., \"Debating with More Persuasive LLMs Leads to More Truthful Answers,\" ICML 2024](https://arxiv.org/abs/2402.06782)). This is direct evidence for the truth-asymmetry hypothesis: at least in evidence-anchored domains, it is structurally easier to argue well for true positions. *(Confidence: high within these domains.)*\n\n### Structured deliberation platforms measurably improve reasoning\n\nIARPA's CREATE program (2017–2021) funded four platforms to test whether crowdsourcing plus structured analytic techniques improves reasoning quality ([IARPA CREATE](https://www.iarpa.gov/research-programs/create)). The University of Melbourne's SWARM platform — collaborative drafting of intelligence-style reports with contending analyses, closely analogous in spirit to Deliberus — produced reports scored substantially higher on the US Intelligence Community's analytic tradecraft standards than the same professionals' normal methods: mean 21.4/32 versus 15.9/32, **Cohen's d = 1.37** ([van Gelder et al., *Journal of Cognitive Engineering and Decision Making* 2020](https://journals.sagepub.com/doi/10.1177/1555343420926287)). Deliberative-polling research adds population-level support: across 100+ deliberative polls in ~30 countries, structured deliberation with balanced materials produces measurable knowledge gains and durable depolarization, most famously in the \"America in One Room\" experiment ([Fishkin, *Hastings Center Report* 2021](https://onlinelibrary.wiley.com/doi/full/10.1002/hast.1316); [overview](https://en.wikipedia.org/wiki/Deliberative_opinion_poll)) — though critics dispute how much of the gain is normatively meaningful. *(Confidence: high for the SWARM effect on quality ratings; note the outcome there was rated reasoning quality, not verified accuracy. Moderate for deliberative-polling claims.)*\n\n### The honest caveat: myside bias resists exactly these virtues\n\nThe bridge has a weak plank, and it must be named. Keith Stanovich's research program finds that while actively open-minded thinking predicts avoidance of almost every cognitive bias, **myside bias is the exception**: it is largely uncorrelated with AOT, with need for cognition, and with intelligence itself ([Stanovich & Toplak, \"Actively Open-Minded Thinking and Its Measurement,\" *Journal of Intelligence* 2023](https://pmc.ncbi.nlm.nih.gov/articles/PMC9966223/)). Worse, Kahan and Corbin found that on climate change, partisans *high* in AOT were **more** polarized than those low in it — cognitive sophistication can be recruited to defend identity-defining positions more skillfully. Since worldview-laden normative claims are precisely where myside bias lives, the accuracy-license established in oracle domains transfers *most weakly* to the domains where Deliberus's claim is boldest. This does not break the bridge — the debate results above show adversarial structure catches what individual dispositions miss — but it means individual reasoning virtue cannot carry the load alone; the *structure* must do the work. *(Confidence: high.)*\n\n**Verdict on (b):** The empirical bridge exists and bears real weight. Reasoning quality — measured by independently specified, text-visible properties — predicts accuracy where accuracy can be checked, and structured adversarial deliberation improves both. The license this grants to oracle-free domains is genuine but defeasible, and weakest for identity-charged questions.\n\n---\n\n## 3. AI-era results: argument-checking survives contact with adversaries — mostly\n\n### The theoretical frame\n\nIrving, Christiano and Amodei's \"AI Safety via Debate\" formalized the hope: two adversarial arguers plus a limited judge can, in theory, verify claims far beyond the judge's unaided capacity (a polynomial-time judge with optimal debaters can decide any PSPACE problem, versus NP for direct checking), resting on the conjecture that it is harder to lie convincingly than to refute a lie ([arXiv:1805.00899](https://arxiv.org/abs/1805.00899)).\n\n### The obfuscated arguments problem — the sharpest known threat\n\nBarnes and Christiano's 2020 negative result is the most important stress-test for any decomposition-based platform: a dishonest debater can decompose a claim into a large argument where **the flaw exists but neither debater can locate it** — each subclaim looks probably-true, the error probability is spread thin across the structure, and the honest side can only say \"somewhere in there is a mistake,\" which is indistinguishable from sour grapes. Their conclusion: \"We can't see a way to distinguish a certain class of obfuscated dishonest arguments from honest arguments,\" and obfuscation emerged *naturally* in their human experiments, not just in theory ([Barnes & Christiano, \"Debate update: Obfuscated arguments problem,\" Alignment Forum 2020](https://www.lesswrong.com/posts/PJLABqQ962hZEqhdB/debate-update-obfuscated-arguments-problem)). Decomposition — Deliberus's core mechanism — is thus double-edged: it is both how confusions get exposed and how a sophisticated bad-faith actor could bury an error.\n\nThe response literature is active and partially reassuring. Doubly-efficient debate establishes formal guarantees under explicit assumptions about verification cost ([Brown-Cohen, Irving & Piliouras, arXiv:2311.14125](https://arxiv.org/abs/2311.14125)), and prover-estimator debate directly targets obfuscation by requiring **stability**: arguments only win if small changes in subclaim probabilities don't flip the conclusion, which defuses the thin-spread-error trick ([arXiv:2506.13609](https://arxiv.org/abs/2506.13609)). Anthropic's debate agenda similarly builds in \"graceful failure\" — debates may end *undecided* when neither side is convincing, rather than forcing a verdict on obfuscated material ([Anthropic Fall 2023 Debate Progress Update](https://www.alignmentforum.org/posts/QtqysYdJRenWFeWc4/anthropic-fall-2023-debate-progress-update)). The problem is mitigated, not solved. *(Confidence: high that the problem is real; moderate that the mitigations suffice in practice.)*\n\n### The empirical record 2022–2025\n\n- **Human-assisted oversight works at proof-of-concept scale:** humans working with an unreliable AI assistant substantially outperformed both the model alone and their own unaided performance ([Bowman et al., \"Measuring Progress on Scalable Oversight,\" 2022](https://arxiv.org/abs/2211.03540)).\n- **Debate beats single-advocate formats robustly:** across the largest study to date (9 task types, ~5M model calls), debate outperformed consultancy everywhere; but against *direct* question-answering the advantage was clear only under information asymmetry, and stronger debaters improved judge accuracy more modestly than earlier studies suggested ([Kenton et al., \"On scalable oversight with weak LLMs judging strong LLMs,\" NeurIPS 2024](https://arxiv.org/abs/2407.04622)). Debate's value is real but task-dependent.\n- **Optimizing for correctness alone degrades checkability:** OpenAI's prover-verifier games showed that training only for right answers makes solutions *less* legible, while training against a small verifier (\"checkability game\") preserved human evaluators' ability to check work — at roughly half the raw performance gain. Sneaky provers, meanwhile, got progressively better at fooling time-limited humans ([Kirchner et al., arXiv:2407.13692](https://arxiv.org/abs/2407.13692); [OpenAI summary](https://openai.com/index/prover-verifier-games-improve-legibility/)). Lesson: legibility and quality-checkability are properties you must explicitly optimize for; they do not come free.\n- **Weak judges can elicit strong performance, imperfectly:** GPT-4 fine-tuned against a GPT-2-level supervisor with a confidence-loss recovered close to GPT-3.5-level performance — well above its supervisor, well below its ceiling ([Burns et al., \"Weak-to-Strong Generalization,\" 2023](https://arxiv.org/abs/2312.09390)). Anthropic's 2025 research agenda identifies the residual hard case precisely: oversight signals with *systematic* errors that the stronger system learns to exploit ([Anthropic, Recommended Research Directions 2025](https://alignment.anthropic.com/2025/recommended-directions/)).\n- **LLM judges are usable but bias-laden:** the documented catalog includes position bias, verbosity bias, self-preference, authority and style biases; frontier models fail a majority of advanced bias probes, and formatting changes alone can flip verdicts ([Ye et al., \"Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge,\" arXiv:2410.02736](https://arxiv.org/pdf/2410.02736); [self-preference bias, arXiv:2410.21819](https://arxiv.org/pdf/2410.21819)). Established mitigations — order-swapping, decomposed rubric criteria rather than holistic scores, reference-based judging, multi-judge panels — reduce but do not eliminate these effects. For argument quality specifically, LLM annotators reach moderately high agreement with human experts on the Wachsmuth dimensions ([Mirzakhmedova et al. 2024](https://arxiv.org/abs/2404.09696)). *(Confidence: high.)*\n\n**Verdict on (c):** Argument-checking is not doomed. The truth-asymmetry that debate relies on has repeated empirical support in ground-truth domains, and the known structural attack (obfuscation) has principled countermeasures — stability requirements, undecided verdicts, checkability training. But every positive result comes from domains with an answer key, judge-gaming is demonstrated and recurring, and nothing here validates verdicts on oracle-free questions directly.\n\n---\n\n## 4. The circularity hole: how deep it goes\n\nAny reasoning-quality score presupposes norms of good reasoning. What justifies *those*? Here the philosophical literature is unambiguous: **there is no non-circular ground, for anyone, for anything** — and it is important to state this plainly rather than pretend Deliberus has an exemption.\n\nWilliam Alston showed that any argument for the reliability of a basic epistemic source (perception, induction, reasoning itself) must ultimately use that source — \"epistemic circularity.\" His considered view was that such arguments can still confer justification if the source is in fact reliable, but he remained \"disquieted,\" and the disquiet is warranted ([Alston, \"Epistemic Circularity,\" 1986; overview at the Internet Encyclopedia of Philosophy](https://iep.utm.edu/ep-circ/); [Lynch & Silva, \"Why Worry about Epistemic Circularity?\"](https://philarchive.org/archive/LYNWWA)). Paul Boghossian sharpened the problem for inference rules: justifying modus ponens requires reasoning that uses modus ponens (\"rule-circularity\"), and Carroll's tortoise blocks the escape route of adding the rule as a premise. Boghossian's own conclusion is telling: if we are to have knowledge of basic logical principles *at all*, it must be via rule-circular reasoning — inference can be \"blind but entitled\" ([Boghossian, \"Blind Reasoning,\" Aristotelian Society 2003](https://as.nyu.edu/content/dam/nyu-as/philosophy/documents/faculty-documents/boghossian/Boghossian_Blind-Reasoning.pdf)). Michael Bergmann's refinement distinguishes **malignant** circularity (where the argument is the *only* source of support and could never fail) from **benign** circularity (where the practice generates ongoing, potentially disconfirming feedback) ([Bergmann, \"Epistemic Circularity: Malignant and Benign\"](https://philosophy.rutgers.edu/joomlatools-files/docman-files/Bergmann.pdf)).\n\nTwo constructive traditions convert this from paralysis into method:\n\n- **Reflective equilibrium** (Goodman on induction, Rawls in ethics): norms and particular judgments are justified jointly, by mutual adjustment until coherence — no foundations, every fixed point revisable ([Stanford Encyclopedia of Philosophy](https://plato.stanford.edu/entries/reflective-equilibrium/)). This is the standard justification structure in ethics, and its critics' main complaint (initial intuitions get *some* weight) is a bullet every working epistemology bites somewhere.\n- **Pragmatist vindication** (Peirce): epistemic norms are justified by their long-run performance in a self-correcting, public, communal process of inquiry — truth as the ideal limit of such inquiry, fallibilism all the way down ([SEP: Peirce](https://plato.stanford.edu/entries/peirce/); [SEP: Pragmatism](https://plato.stanford.edu/entries/pragmatism/)).\n\nThe empirical bridge of §2 is exactly a Bergmann-benign, Peircean move: the norms embedded in a reasoning-quality score (consider both sides, use base rates, answer the critical questions, keep evidence chains stable) are not validated by fiat but by their **track record against outcomes in every domain where outcomes can be checked** — and that validation loop stays open, so the norms remain falsifiable and revisable. This is still circular in Alston's sense (we use reasoning to evaluate the track record). But note the crucial symmetry: **the demand for a non-circular justification is itself an epistemic norm that cannot be non-circularly justified.** Science, law, mathematics, and everyday cognition all operate inside the same circle. The meaningful engineering question is never \"is the circle escapable?\" (no) but \"is the circle wide, self-correcting, and discriminating?\" — does it incorporate independent checks, does it revise its own norms under evidence, and does it actually sort better from worse in checkable cases? On all three, the material in §§1–3 answers yes. *(Confidence: high that this is the state of the art in epistemology; this section's conclusion is a philosophical judgment, not a finding.)*\n\n**Verdict on (d):** The hole is bottomless but universal, and the best available response — wide reflective equilibrium plus pragmatist track-record validation, with the norms themselves kept revisable — is available to Deliberus in an unusually literal way, because a claim graph can *contain its own scoring norms as challengeable claims*.\n\n---\n\n## Implications for Deliberus\n\n**What an honest reasoning-quality score can claim:**\n\n1. **A relative, procedural verdict:** \"given the claims, evidence, attacks, supports, and critical questions currently in the graph, this position withstands scrutiny better than that one.\" This is a statement about the state of a deliberation, not about the world — and it is exactly the kind of statement the SWARM/CREATE results show can be measured reliably and improved dramatically (d = 1.37) by platform structure.\n2. **Non-arbitrariness:** the score's propagation obeys published axioms (balance, monotonicity, duality, open-mindedness) that constrain the space of defensible semantics; its quality dimensions (CQ coverage, evidential sufficiency, dialectical complexity) come from independently developed theory, not from the platform's preferences about conclusions.\n3. **A defeasible truth-license:** the same properties the score rewards demonstrably predict accuracy in oracle domains (forecasting rationales, debate experiments). Trusting them in oracle-free domains is an induction — explicitly defeasible, but rationally grounded rather than circular in the malignant sense.\n4. **Diagnosis, not just evaluation:** because the score decomposes (this CQ unanswered, that premise unsupported, this subtree unstable), it tells contributors *what would change the verdict* — which is what makes it an invitation rather than a gavel.\n\n**What it cannot claim:**\n\n1. **Conclusion-truth in oracle-free domains.** \"Best-reasoned position currently on the platform\" and \"true\" must never be conflated in UI copy, API naming, or public framing. Better-reasoned-but-wrong is a live, permanent possibility; the GJP structural model's ceiling (multiple R ≈ 0.64 *with* outcome feedback) is a reminder of how much variance reasoning quality does not explain even under ideal conditions.\n2. **Immunity to motivated sophistication.** Myside bias is uncorrelated with the very dispositions the score rewards, and high-sophistication partisans can polarize *more*. The platform's protection is structural (adversarial CQs, bridging detection, evidence requirements), not dispositional — and should be advertised as such.\n3. **Immunity to obfuscation-by-decomposition.** Deliberus's own core mechanism is the known attack surface. Countermeasures worth building in: stability checks (does the verdict survive small perturbations of subclaim strengths — the prover-estimator insight), explicit \"undecided\" states instead of forced verdicts on murky subtrees, and suspicion of arguments whose strength depends on error probability spread thinly across many barely-checked steps.\n4. **Judge neutrality without engineering.** Any LLM-assisted scoring inherits position, verbosity, and self-preference biases; order-swapping, decomposed rubrics, and multi-model panels are necessary hygiene, and the residual bias should be assumed nonzero.\n\n**Design consequences worth acting on:**\n\n- **Put the norms in the graph.** The scoring rules themselves — which CQs a scheme gets, how strength propagates, what counts as sufficient evidence — should exist as claims subject to the platform's own decomposition and challenge mechanics. This is the reflective-equilibrium response to circularity made literal in software, and it converts the platform's deepest philosophical vulnerability into its most distinctive feature. (It is also the self-similar decomposition principle applied to the metric itself.)\n- **Build the oracle-bridge internally.** Some claims on the platform will resolve (forecasts, empirical predictions, reproducibility outcomes). Tracking whether high-scoring reasoning predicts resolution *within Deliberus* creates a standing, platform-native validation loop — the Peircean track record, continuously updated, and the strongest possible answer to \"why trust this score?\"\n- **Version and label the verdict.** Scores should carry timestamps and graph-state provenance (\"as of this state of the debate\"), making explicit that they are revisable states of a self-correcting process, not pronouncements.\n- **On the convergence wager:** none of this research proves that values converge under decomposition — but the wager does not need to be presupposed for ranking to be legitimate. Ranking reasoning quality is licensed independently (§§1–3); convergence, if real, will show up as an *empirical pattern* in the graph over time. The platform can therefore treat its boldest thesis as a hypothesis it is instrumented to test, which is a far stronger epistemic position than treating it as an axiom.\n\n---\n\n## Strongest objection to my own conclusions\n\nEvery empirical result marshaled above comes from domains with answer keys: geopolitical questions that resolve, reading-comprehension questions with correct answers, math problems, intelligence-tradecraft rubrics scored by trained raters. The inference \"reasoning quality tracks truth where we can check, therefore trust it where we can't\" is an induction across precisely the boundary that matters — and the best-documented finding about that far side is Stanovich's: the biases that dominate worldview-laden reasoning are the ones *uncorrelated* with measured reasoning virtue, and sophistication can amplify them. Meanwhile, Goodhart's law guarantees that once a reasoning-quality score allocates status on a public platform, optimization pressure against the score will produce high-scoring sophistry, and the obfuscated-arguments result is a formal proof that such sophistry can be constructed so that no checker can localize the flaw. It is therefore possible that in exactly the normative domains Deliberus most cares about, its score will function less as a truth-tracker than as a fluency-and-effort tracker — rewarding those with the time and skill to build imposing argument structures.\n\nI do not think this objection wins — the debate experiments show adversarial structure catching what dispositions miss, persuasiveness empirically correlated with being right when both sides get equal arms, and the mitigations (stability checks, undecided verdicts, internal calibration on resolvable claims) are concrete rather than hopeful. But the objection sets the terms Deliberus must live by: the score is safe only while it is presented as defeasible, kept contestable down to its own norms, and continuously audited against every oracle the platform can reach. The moment it is treated as bedrock, it becomes the thing it was built to replace.\n"}