{"path":"research/assumption-ranking.md","content":"# Assumption Ranking: Weakest to Strongest\n\n**Date**: March 28, 2026\n**Method**: Critical examination of every load-bearing assumption in the Deliberus project, ranked by available evidence. Produced after 29 research documents, 8 conceptual threads, steelmanned critiques, and multi-model consensus.\n\n---\n\n## TIER 1: Genuinely Uncertain (thin or no direct evidence)\n\n### 1. Two-axis voting produces reliable signal — WEAKEST\n\nLessWrong/EA Forum implemented agree/disagree + quality voting in 2022 with \"mixed results\" and \"cognitive overhead.\" We have NO empirical data on whether users can reliably separate \"I agree with the conclusion\" from \"this reasoning is sound.\" This is foundational — if it fails, bridging arguments fail, and the novel theoretical contribution collapses. The epistemic gamification research ([epistemic-gamification.md](epistemic-gamification.md)) acknowledges this but doesn't resolve it.\n\n*Evidence: one partial implementation with mixed results. Zero data on argument-specific two-axis voting.*\n\n### 2. Bridging arguments are detectable and valuable in practice\n\nThe most NOVEL claim, therefore the LEAST tested. PAKT (Heidelberg 2024) provides a data model but doesn't implement cross-group detection. Blair et al. (2025) improved bridging *statements* but not bridging *arguments*. The concept is intellectually compelling but relies on assumption #1 working. Could be that cross-group \"well-argued\" ratings are too rare to produce signal, or that the signal drowns in noise.\n\n*Evidence: zero implementations. Theoretical framework only ([bridging-arguments.md](bridging-arguments.md)). Depends on #1.*\n\n### 3. Human-in-the-loop correction is tolerable as a sustained workflow\n\nGPT-5.2 flagged this hardest ([second-opinion-mvp-strategy.md](second-opinion-mvp-strategy.md)): \"Wizard of Oz systematically lies about interaction costs.\" Wikipedia editors correct, but for community status. Metaculus forecasters calibrate, but that's a different cognitive task. Whether correcting LLM-extracted argument maps feels rewarding or feels like grading homework is genuinely unknown. The whole \"correction UX IS the product\" thesis rests on this being tolerable enough that people return.\n\n*Evidence: analogies from different domains (Wikipedia, Metaculus). Zero direct testing.*\n\n### 4. The fact/value boundary is reliably classifiable at extraction time\n\nThe 30°C example (object-model.md) is clean. Real claims blur: \"Climate change will cost $X trillion\" — factual premise of a normative argument? Emanuel's 2013 discourse-layer insight ([voice-memo-emanuel-sofia.md](voice-memo-emanuel-sofia.md)) goes deeper: even correctly classified claims carry discourse-level meaning that classification misses. Two people endorsing the same \"factual\" claim may be making opposite normative moves.\n\n*Evidence: theoretical framework (Thread 3). No extraction-level testing. Emanuel's voice memo actively challenges the assumption.*\n\n### 5. Polis-style clustering + argument structure = emergent insight\n\nThe combinatorial bet. Intellectually beautiful. Zero empirical evidence. People's voting patterns might not align with their logical reasoning in the ways the theory predicts. An opinion cluster isn't necessarily a reasoning cluster — Group A might agree on conclusions for 5 completely different reasons, making \"bridging arguments\" for Group A incoherent.\n\n*Evidence: zero. Pure theory ([bridging-arguments.md](bridging-arguments.md)). The combination has literally never been tried.*\n\n---\n\n## TIER 2: Reasonable but unproven in this context (evidence from adjacent domains)\n\n### 6. LLMs change the friction equation ENOUGH\n\nStrong theoretical argument (the \"Why Now?\" table in [conceptual-threads.md](../conceptual-threads.md)). LLM argument advisors doubled quality scores in one study (IntechOpen). But 50-60% F1 on real political text is sobering ([steelmanned-critiques.md](steelmanned-critiques.md)). The \"eager but flawed intern\" framing is right — the question is whether this intern is good ENOUGH that the correction burden doesn't overwhelm the extraction benefit. The claim extraction experiment will directly test this.\n\n*Evidence: Claimify paper, one study on argument quality. Strong theory, limited empirical validation on messy real-world text.*\n\n### 7. Single-player utility actually solves cold-start\n\nStrong analogical evidence (Roam, Obsidian, Notion all started personal — [single-player-utility.md](single-player-utility.md)). But those tools serve EXISTING cognitive behaviors (note-taking, planning). \"Argument mapping\" is NOT an existing behavior — the tool must teach a new skill while delivering value. \"Help me think about X\" reframing helps, but whether it dissolves the cold-start or just postpones it is unknown.\n\n*Evidence: 5 strong analogies, but all from tools matching existing behaviors. Argumentation is a new behavior.*\n\n### 8. Progressive disclosure works for argument complexity specifically\n\nExcellent evidence from adjacent fields: cognitive load theory, expertise reversal effect, Wikipedia usage patterns, semantic zoom studies ([progressive-disclosure.md](progressive-disclosure.md)). But the research itself notes: \"No direct study on progressive complexity in deliberation yet.\" The theory is solid; the specific application to argument graphs is extrapolated.\n\n*Evidence: strong adjacent-domain evidence. Zero deliberation-specific evidence. Identified as a publishable research opportunity.*\n\n### 9. The rationalist/EA community is the right beachhead\n\nThey demonstrably want these tools (EA Forum discussions). They tolerate friction. They value calibration. But the market is 50K-200K globally ([steelmanned-critiques.md](steelmanned-critiques.md)), many already use Metaculus/LessWrong/Kialo, and this community is notoriously hard to convert from \"I agree this should exist\" to \"I use this daily.\"\n\n*Evidence: demonstrated demand via discourse ([adoption-problem.md](adoption-problem.md)). Unknown conversion rate from interest to usage.*\n\n### 10. Grant/philanthropic funding can sustain epistemic infrastructure\n\nPolis: grant-dependent, fragile. Kialo: single patron, existential dependency. Signal: $50M/year costs. The governance research ([governance-and-financing.md](governance-and-financing.md)) found viable models (Mozilla hybrid, Metaculus PBC) but they all took years to reach sustainability. Getting FROM zero TO sustainable is the unproven step.\n\n*Evidence: models exist but are hard to replicate. No clear path from \"solo dev\" to \"funded infrastructure.\"*\n\n---\n\n## TIER 3: Well-supported (strong evidence, high confidence)\n\n### 11. People want to reason better (the market exists, however small)\n\nMetaculus has active forecasters. EA Forum: 300K+ monthly visitors. Kialo: 1M+ registered. LessWrong sustained for 15+ years. The FB group had 452 members. The market is small but real. The question is size (50K or 500K?), not existence.\n\n*Evidence: multiple sustained platforms serving this need. Kialo's 1M users proves SOME scale.*\n\n### 12. Formal argumentation adds value over informal discussion\n\nDecades of evidence: van Gelder's studies, Monk's research, DeliData (64% of groups found better solutions), Scheuer meta-analysis ([academic-foundations.md](../academic-foundations.md)). Among the most well-established findings in the field.\n\n*Evidence: multiple meta-analyses, decades of replication.*\n\n### 13. A solo developer with AI augmentation can build this\n\nFredrik's demonstrated track record (8-15x multiplier, shipped projects). The 2026 AI dev landscape. This is about execution capability, which is proven.\n\n*Evidence: empirical track record across multiple projects.*\n\n### 14. No platform currently combines formal argumentation + LLMs + usable interface + probabilistic reasoning\n\nVerified across 29 research documents and every competitor analysis ([competitive-landscape.md](../competitive-landscape.md), [kialo-deep-dive.md](kialo-deep-dive.md), [polis-deep-dive.md](polis-deep-dive.md), [deliberation-io.md](deliberation-io.md), [argsbase-argument-web.md](argsbase-argument-web.md)). Kialo has no AI. Polis has no argument structure. Deliberation.io has no formalism. This gap is a factual observation, not an assumption.\n\n*Evidence: comprehensive competitive analysis. This is fact, not assumption.*\n\n### 15. 2026 AI capabilities make this possible in ways 2012 didn't — STRONGEST\n\nThe \"Why Now?\" table in [conceptual-threads.md](../conceptual-threads.md) documents 15 capabilities that didn't exist. LLM extraction, embedding similarity, NLI models, CRDTs, WebGL graph rendering, defeasible argumentation frameworks. These capabilities exist. The assumption that they SUFFICE is weaker (see #6), but their EXISTENCE is not in question.\n\n*Evidence: the capabilities demonstrably exist. This is observation, not assumption.*\n\n---\n\n## Genuinely UNCLEAR (open questions, not assumptions)\n\nThese aren't weak assumptions — they're **unresolved design questions** where the research doesn't point in a single direction:\n\n- **The ontology**: AIF? Custom? How many node types? Explicitly deferred, genuinely open\n- **What the \"aha moment\" actually feels like**: We theorize but have zero experiential data\n- **The discourse layer problem**: Emanuel identified it in 2013 ([voice-memo-emanuel-sofia.md](voice-memo-emanuel-sofia.md)). No one has proposed a practical solution. May be inherent and unresolvable\n- **The right AI autonomy balance**: Too much AI = \"the system decided for me\" mistrust. Too little = \"this is too much work.\" Where's the sweet spot?\n- **Whether the combinatorial vision composes into coherent UX**: Each component is buildable. Whether they create something coherent rather than overwhelming is a design question, not a research question\n\n---\n\n## Strategic Implication\n\nThe **foundation is strong** (#11-15): market gap real, timing right, capability exists, formal argumentation helps.\n\nThe **strategic bets are reasonable** (#6-10): analogical evidence from adjacent domains, untested in this specific context.\n\nThe **novel contributions are the weakest link** (#1-5): genuinely untested because they're genuinely new. This is expected — novel contributions by definition lack prior validation.\n\n**The claim extraction experiment + combinatorial demo directly tests the weakest assumptions first.** This is the correct sequencing: attack the highest-risk assumptions before investing in infrastructure.\n"}