{"path":"research/session7-deep-review-and-strategic-assessment.md","content":"# Session 7: Deep Review and Strategic Assessment\n\n**Date**: March 28-31, 2026\n**Type**: Full corpus review → ontology design → implementation sprint → production deployment\n**Duration**: Multi-day session — the longest and most productive in the project's history\n\n---\n\n## The Research-vs-Decisions Boundary\n\n**Fredrik's statement (verbatim, Session 7):**\n\n> \"I am open-mindedly exploring architecture and ontology and processes and everything in this stage. And so I'm using the research documents as inspiration, but ultimately all architectural and ontology questions are and should be up to me. So I would like the main key documentation to make this clear and to maintain a boundary between the discoveries and insights and recommendations that might surface in all of the research docs. This boundary is permeable and soft yet I am the final arbiter. I simultaneously find myself having deep convictions and intuitions about what's possible and desirable when it comes to everything concerning Deliberus and its potential. Also being deeply curious about limitations, possibilities or warnings or advice or ideas in all of the research that we have amassed and will continue to amass. But I need to go through it in order to make decisions - all flows through me.\"\n\n**Implication**: The 47 research documents describe a DESIGN SPACE — options, trade-offs, analogies, warnings, inspirations. They are not a specification. The extraction pipeline, ontology, object model, UX patterns, and architectural choices are all exploratory until Fredrik makes decisions. AI research agents surface possibilities; the founder decides.\n\n---\n\n## Honest Audit: What Exists vs What Research Describes\n\n**What's built** (Sessions 3-6, ~14 hours):\n- 6-pass extraction pipeline (scout → decompose → decontextualize → relationships ∥ concepts → self-eval → embed+autolink)\n- FalkorDB graph with 130 claims across 2 sources\n- FastAPI REST API with SSE streaming\n- SvelteKit web app (dark design, sorry markers, voice input, previous extractions, claim detail, interactive D3 ego-graph)\n- 64 tests, deployed to Darwin\n\n**The gap between code and research**: The research corpus describes a deeply interactive system. The code is currently a display layer — an extraction viewer. This is exactly right for the stage, but the distance is worth naming:\n\n| Research describes | Code currently has |\n|---|---|\n| Sorry markers as functional contribution entry points (Lean §1) | Sorry markers as inert status badges |\n| Definitions as first-class argument nodes (semantic disambiguation) | Definitions as one of four claim types, no special UX |\n| Correction UX that provokes engagement (P14 interpassivity) | Read-only claim display |\n| One shared graph (Wikipedia model) | Per-extraction display with auto-linking between |\n| Deprecation over deletion (P12) | No edit/versioning mechanism |\n| Composing = retrieval (P13) — typing searches the graph | No compositional search UX |\n\nThis gap is not a problem — it's the work ahead. The research maps the destination; the code is the first step.\n\n---\n\n## Three Horizons Framework\n\n**Horizon 1 (weeks)**: Personal thinking tool. \"Help me think about X.\" Paste text or URL → get structured argument analysis → edit/correct/decompose. Single-player utility that dissolves cold-start. **The correction UX builds toward this.**\n\n**Horizon 2 (months)**: Combinatorial demo. The moment extraction + four-type classification + two-axis evaluation interact and reveal something no existing tool shows — a bridging argument, a semantic fork, a value premise hiding under an apparent empirical disagreement. **This justifies existence vs. contributing to Kialo or Polis.**\n\n**Horizon 3 (years)**: The civilizational deliberation graph. All worldviews mapped onto one structure, navigable through any lens. **This is the destination, not the next step.**\n\nThese are not sequential phases with gates. They're nested — Horizon 1 is an entry point INTO Horizon 3. The Wikipedia model: you come to read, gradually start editing. Personal → social → civilizational.\n\n---\n\n## The Dialectic Is Fractal\n\nThe analysis ↔ attunement dialectic (vision.md §Core Dialectic) recurs at every level of the project. This was identified in the spec session (docs/research/spec-session-reasoning.md) and deepened through this session's full-corpus review:\n\n| Level | Analysis pole | Attunement pole |\n|---|---|---|\n| Philosophy | Decompose, verify, compute (Need for Cognition) | Inhabit other worldviews, understand WHY (perspective-taking) |\n| Strategy | Civilizational graph (ingest everything, decompose everything) | Personal tool (start where human IS, from THEIR text/question) |\n| Architecture | Graph nodes — discrete, categorical, committed (left hemisphere) | Embedding space — continuous, contextual, relational (right hemisphere) |\n| UX | Claim detail page, formal structure, type badges | \"Help me think about...\" free text box, voice input |\n| Data model | Four-type classification (explicit categories) | Worldview emergence from voting patterns (implicit clustering) |\n| Adoption | Formalization paradox (structure repels users) | LLM absorption of structuring burden (users \"just talk\") |\n\nThe dialectic isn't a theme decorating the project — it IS the project's structure. Every design question resolves to \"how much analysis, how much attunement, and how do they feed each other?\"\n\n---\n\n## Pipeline Testing Principle: Don't Go Meta\n\n**Fredrik's statement (verbatim, Session 7):**\n\n> \"My gut tells me that this is going a bit too meta. I think it would be more of a learning, valuable learning experience to feed something different and more concrete about various topics. It can be complex, and that's probably very valuable to see how complex material with high argumentative density and somewhat longer content than the articles we've fed in thus far is handled by Deliberus, and then to iterate based on that. But there's no reason to and actually I think it should be avoided maybe to go too meta too early without having battle tested the pipeline and architecture and ontology and all HITL elements etc etc of Deliberus first in a wave of more targeted testing using less meta-meta-material.\"\n\n**Implication**: Don't test the extraction pipeline on papers ABOUT argumentation theory. Test it on the kind of material it's actually FOR — substantive debates about real topics. AI risk, democracy, cognitive biases, economics, climate policy, ethics. The pipeline should prove its value on content users would actually want to analyze, not on academic meta-commentary about its own domain.\n\nPapers about argument mapping are useful as research references. They should not be the test corpus.\n\n---\n\n## PDF Archive Discovery\n\nFour distinct PDF archives found across drives:\n\n| Archive | Location | Count | Character |\n|---|---|---|---|\n| **PDF Archive (Dropbox)** | FERMI drive, Dropbox PDF Archive | **1,063** | 2009-2016 research library with dedicated subfolders |\n| **Books (MBP)** | FERMI drive, books collection | ~50 | Academic papers + books |\n| **Teleological Evolution** | BOHR drive, Google Drive Backup | ~30 | Consciousness, philosophy, transhumanism |\n| **Deliberus (BOHR)** | BOHR drive, Google Drive Backup | 2 PDFs + .gdoc stubs | Original business plan + canvas |\n\n### Dropbox PDF Archive Subfolders\n\n```\nAGI MEDO Existential Risk/    — AI risk, neural nets, Bostrom, Omohundro, Yudkowsky\nArgument mapping/              — 15 papers: van Gelder, decision mapping, visualization\nBlablabla/                     — Mixed (Coffin argumentation principles, misc)\nBooks/                         — General books/papers\nCritical Theory/               — Flyvbjerg (rationality & power)\nCrowdsourcing/                 — Rob Miller on crowd computing\nFood Ethics/                   — (includes cognition-related)\nGrand:Integral Theorizing/     — Integral theory, alternative economics\nHarvard Business Review/       — (unchecked)\nHealth Psychology/             — Theory of mind, cognitive testing\nkarolinska-sömn-ork-studie/   — Sleep + decision-making research\nManuals/                       — Technical manuals\nnextnature_teaching_kit/       — (unchecked)\nRelevant to Kaus/              — (Kaus project materials)\nStuff/                         — Unsorted\n```\n\n### Best Candidates for Extraction (Non-Meta, Substantive Topics)\n\n**Tier 1 — Concrete topics with rich argumentation, right length:**\n\n1. **`cognitive-biases-global-risk_yudkowsky.pdf`** — Yudkowsky on how cognitive biases threaten existential risk assessment. Rich empirical + normative claims. Tests pipeline on rationality critique.\n2. **`Ethical Issues in Advanced Artificial Intelligence (Bostrom).pdf`** — Bostrom on AI ethics. Dense normative reasoning with empirical premises. Tests fact/value boundary.\n3. **`bostrom-savulescu-enhancement-ethics-stateofdebate.pdf`** — Human enhancement ethics debate. Multiple competing ethical frameworks in one text. Tests contested concepts (what counts as \"enhancement\"?).\n4. **`democracy-for-the-21th-century-research-challenges.pdf`** — Democratic theory. Normative + institutional claims. Tests political theory extraction.\n5. **`PredictiveLiquidDemocracy.pdf`** — Liquid democracy mechanisms. Mix of formal mechanism design + normative claims about governance.\n6. **`EcologicalEconomicsShortDescription2015V2.pdf`** — Alternative economics. Tests pipeline on paradigm-challenging claims with definitional disagreements.\n7. **`BoardDecisionMakingWhitepaperDraft.pdf`** (486KB) — Corporate decision-making. Practical argumentation with empirical evidence. Tests pipeline on business/institutional text.\n\n**Tier 2 — Longer or more specialized, excellent material:**\n\n8. **`Albert Bayesian Rationality and Decision Making A Critical Review.pdf`** — Bayesian rationality critique. Technical but deeply argumentative.\n9. **`cooperation-equilibrium-logic.pdf`** — Game theory + cooperation. Formal reasoning with normative implications.\n10. **`Why-and-how-to-upgrade-human-collective-wisdom-by-Christer-Nylander.pdf`** — Swedish author on collective wisdom. Directly relevant topic, non-meta.\n11. **`empathetic-superintelligence.pdf`** — Empathy + AI. Tests the analysis ↔ attunement tension on a concrete topic.\n12. **`Collective intelligence.pdf`** — Collective intelligence survey. Broad, empirically grounded.\n13. **`the_case_for_developmental_methodologies_in_democratization.pdf`** — Developmental approaches to democracy. Complex normative + empirical interplay.\n\n**Tier 3 — Book-length (RLM needed):**\n\n14. **`Nick Bostrom, Milan M. Cirkovic - Global Catastrophic Risks.pdf`** (3.9MB, BOHR) — X-risk anthology. Multiple chapters, multiple authors, rich disagreement.\n15. **`Andy Clark - Supersizing the Mind.pdf`** (1.1MB, BOHR) — Extended mind thesis. Not about argumentation — about cognition.\n16. **`Daniel Dennett - Consciousness Explained.pdf`** (5.4MB, BOHR) — Philosophy of mind. Dense argumentation throughout.\n17. **`Deliberus Lecture Notes.pdf`** (13MB) — Online Deliberation book. Full RLM test case.\n\n### The Argument Mapping Subfolder (Reference, Not Test Corpus)\n\nThese 15 papers are valuable as **research references** but should NOT be the primary test corpus (per the \"don't go meta\" principle):\n\n- `Argument Mapping Encyc Submission Final.pdf` (127KB) — van Gelder encyclopedia entry\n- `Computer-supported argumentation.pdf` (867KB) — Survey\n- `enhancing-our-grasp-of-complex-arguments.pdf` (464KB) — van Gelder on argument mapping benefits\n- `Interview with Tim van Gelder - TheReasoner-4(2).pdf` (418KB) — Interview\n- `ISVC08VisualizingArgumentStructure.pdf` (532KB) — Visualization research\n- `What is DecisionMapping.pdf` (493KB) — Decision mapping theory\n- `Hi-Trees and Their Layout.pdf` (4.5MB) — Layout algorithms\n- `webmap.pdf` (1.3MB) — Web-based argument mapping\n- `No Computer Program Required- Even Pencil-and-Paper Argument Mapp.pdf` (246KB)\n- `writingandinferencing.pdf` (151KB)\n- `BoardDecisionMakingWhitepaperDraft.pdf` (486KB) — (this one IS substantive, listed in Tier 1)\n- `artclISSAConf.pdf` (258KB) — Conference paper\n- `Sherlynn+Bessick2+Corrected.pdf` (801KB)\n- `Strong Theory Weak Dialogue.pdf` (232KB)\n- `Wise Delinquency.pdf` (903KB)\n\n---\n\n## RLM Angle Assessment\n\n**What RLM is**: Recursive Language Models (MIT 2512.24601) treat the source text as an external object the LLM navigates via tool calls (peek, search, chunk). The model holds \"steering logic + goal\" and accesses text through search, never needing it all in context simultaneously. See RLM optimization plan (local reference doc) for implementation patterns.\n\n**When the current pipeline works fine** (no RLM needed):\n- Academic papers under ~30 pages (~15K words)\n- Wikipedia articles\n- Blog posts and essays\n- Most of the Tier 1 and Tier 2 PDFs above\n\n**When RLM becomes necessary**:\n- Book-length texts (>100 pages): Dennett, Clark, Bostrom/Cirkovic, the Online Deliberation book\n- Multi-article corpora: extracting across a set of related articles\n- Full forum thread archives: entire r/changemyview threads\n- The FB group archive (406 posts, 1,956 comments — 1.3MB of text)\n\n**How it would work for the extraction pipeline**:\n1. Store full text externally (already done — JSON files)\n2. Scout phase uses search/peek tool calls to identify argument-dense sections across the full text\n3. Focused extraction zooms into relevant passages, with system message containing section + global context summary\n4. Cross-structure analysis connects findings from distant sections\n5. Each pass operates on sections but maintains awareness of the whole\n\n**PDF support — IMPLEMENTED (Mar 31, 2026)**: Gemini reads PDFs natively via `inline_data` — no `pymupdf`/`pdfplumber` needed. `POST /extract/pdf` with drag-and-drop frontend. Tier 1/2 (<30 pages). Gemini `inline_data` with `response_schema`. Book-length still needs RLM. See CLAUDE.md §PDF ingestion.\n\n---\n\n## Steelmanned Critiques Summary (From Session 7 Agent Synthesis)\n\nThe full critique analysis is in `docs/research/steelmanned-critiques.md`. Key findings that should inform current decisions:\n\n1. **The rationality market may be structurally tiny** (~100K-200K globally). Not disqualifying — but scope to high-stakes institutional use where external incentives substitute for intrinsic motivation.\n2. **The formalization paradox** (Wittgenstein): meaning arises in USE. The \"stranger test\" filters out the most politically important material. Deliberus must handle irreducible ambiguity, not just clean propositional claims.\n3. **The platform graveyard** (20+ failures): alternative hypothesis is structural incompatibility between rigorous reasoning and online social dynamics. Mitigation: single-player utility FIRST.\n4. **LLM extraction quality**: ~90% on clean benchmarks → 50-60% on real-world political discourse. System must be honest about this — visible confidence scores, human correction as primary workflow.\n5. **The complexity ceiling**: comprehensive map of one major debate requires 10K-100K+ claims. Working memory (7±2 items) limits cognitive engagement. Semantic zoom helps navigation but doesn't solve maintenance.\n6. **The gaming threat**: structured format creates veneer of rationality that's exploitable. Wikipedia-level manipulation is possible. Manipulation resistance must be architectural from day one.\n7. **The neutrality illusion**: every design choice embeds epistemological assumptions. Radical humility: \"one framework\" not \"the framework.\"\n\n---\n\n## Cross-Thread Synthesis Findings (From Session 7 Agent Synthesis)\n\nThe full synthesis is in `docs/research/cross-thread-synthesis.md`. Key emergent tensions from thread interactions:\n\n1. **Adoption × Gaming** (The Gamification Trap): Extrinsic motivation systematically corrupts the behavior it rewards. Calibration optimization may cause users to avoid genuinely uncertain territory.\n2. **Hybrid Intelligence × Consensus** (System Ideology Risk): If the system aggregates opinion → surfaces content → shapes opinion → feedback loop, it develops ideological drift. Early adopter bias seeds baseline norms; LLM extraction biases apply consistent political lens.\n3. **Fact/Value × Probabilistic** (What Does \"70% Likely We Should Ban X\" Mean?): Could mean sociological fact, empirical probability, moral uncertainty, or framework-conditional probability. Conflating these is conceptual confusion.\n4. **Scalability Paradox**: Governance overhead grows faster than content. Evidence suggests systems strain past ~150 concurrent participants on one topic. The civilizational graph vision assumes scaling far beyond what's been demonstrated.\n5. **AI Dependency Risk**: The entire adoption thesis rests on \"LLMs absorb the structuring burden.\" LLM quality plateau risk is real. More capable models make more sophisticated errors, not fewer.\n\n---\n\n## Wise Next Steps (Session 7 Assessment)\n\nIn priority order, with reasoning:\n\n**1. Make sorry markers functional** (days)\nThe single highest-leverage transformation from viewer to product. ◇ → decomposition flow; ○ → evidence submission with P13 composing-as-retrieval. Tests sorry model tolerability, correction loop, and composing-as-retrieval simultaneously.\n\n**2. ~~Add PDF ingestion~~** — DONE (Mar 31). Gemini native PDF, no pymupdf. See CLAUDE.md §PDF ingestion.\n\n**3. Extract 2-3 substantive PDFs** (days)\nRun the pipeline on Bostrom AI ethics, Yudkowsky cognitive biases, or ecological economics. Reveals pipeline weaknesses on real academic argumentation — longer, denser, more contested than Wikipedia.\n\n**4. Surface semantic disambiguation actively** (week)\nWhen two extractions share a concept but with different senses, the UI should highlight this. Uses existing infrastructure (embeddings + contested concepts + SIMILAR_TO links). Tests the unique value proposition: \"oh, we were using the same word differently.\"\n\n**Deprioritized**:\n- Graph browsing page — the card/claim-based UI IS the right primary interface (Thread 6)\n- Two-axis evaluation — needs user accounts\n- Fix kamal deploy — workaround exists\n- Cross-graph attack detection — advanced, deferred correctly\n\n---\n\n## Key Insights From Agent Syntheses (All 47 Docs)\n\nFive agents read the complete corpus in parallel. Here are the highest-value insights surfaced, organized by theme:\n\n### From the Lean Documents (lean-deliberus-analogies, lean4-deep-dive, lean-social-system-research)\n\n- **`sorry` is architectural, not decorative**: In Lean, `sorry` enabled 25 strangers to formalize a 33-page proof in 3 weeks (Tao's PFR project). Each person filled in ONE sorry without understanding the whole proof. The contribution barrier collapsed because: (a) gaps were visible, (b) each gap was self-contained, (c) the compiler tracked transitivity. For Deliberus, sorry markers should function identically — each is a self-contained contribution opportunity.\n- **The `@[simp]` flywheel is compounding knowledge automation**: Each community-vetted claim makes the system better at auto-connecting future arguments. Like Lean's `exact?` tactic that searches the entire Mathlib library for a matching lemma.\n- **Deprecation, never deletion**: When a better formulation emerges, the old one isn't deleted — it's linked with \"superseded by.\" The history of how understanding evolved IS content.\n- **Three-skill model**: Contributing arguments, organizing deliberations, curating the canonical graph — different skills for different people. Maps to Shneiderman's Reader → Contributor → Collaborator → Leader.\n- **The kernel is small and auditable**: Lean's trusted kernel is 5,000 lines of C++. Everything else (tactics, automation, UI) is untrusted but funneled through the kernel. For Deliberus: the \"argument validity kernel\" should be small, explicit, and deliberately limited — only structural rules, never truth claims. Community judges meaning.\n\n### From Competitive Analysis (kialo, polis, habermas-machine)\n\n- **Kialo survived through education-first pivot**: 1M+ users, 18K+ debates. Binary pro/con trees + 500-char limit. What fails: binary reductionism (complex positions don't fit), popularity over logic, no API, no cross-discussion linking. Lesson: richer relation types from day one; API + structured export.\n- **Polis scales but lacks structure**: 10M+ participants. Opinion clustering works. vTaiwan produced 26 laws. But: no argument structure, no reasoning quality assessment, no explanation of WHY groups agree or disagree. The gap Deliberus fills.\n- **Habermas Machine outperformed human mediators** (56% preference) but optimizes for approval, not validity. Aggregation without structure. The counter-model: transparent disagreement over hidden consensus.\n\n### From Philosophy (zizek-schmachtenberger, embeddings-tension, steelmanned-critiques)\n\n- **Žižek's parallax**: Some contradictions are irreducible. No neutral position exists. The system should mark \"parallax claims\" where disagreement is structural, not informational. Two-axis voting is the defense against cynical distance.\n- **Schmachtenberger's Consilience Project failed**: Content-heavy (20 articles over 5 years), not substrate-based. Validates participatory architecture over centralized editorial.\n- **Anti-rivalrous dynamics**: Every feature should become MORE valuable as more people use it. Knowledge graph contribution has anti-rivalrous properties.\n- **The embeddings tension is real**: Embeddings trained on internet text converge toward average meaning, flattening contested distinctions. Same mechanism as \"AI slop.\" Resolution: embeddings SUGGEST, the graph COMMITS. Users must mediate. Disambiguation is productive and should be facilitated, not eliminated.\n\n### From Adoption and UX (adoption-problem, progressive-disclosure, epistemic-gamification, mobile-argument-ux)\n\n- **Every failed platform demanded structure as INPUT**: Deliberus must generate structure as OUTPUT. LLMs absorb the structuring burden. This is the prerequisite that didn't exist pre-2023.\n- **Progressive disclosure is mandatory** (expertise reversal effect): Fixed complexity always fails someone. Three levels: casual (vote, read summaries), curious (premise chains, sorry markers), expert (probabilistic weights, falsification history). Same graph, three renderings.\n- **The \"inspectability premium\"**: Mere AVAILABILITY of deeper inspection increases trust even when rarely used. Knowing you COULD check increases trust at the summary level.\n- **Epistemic gamification must reward virtues, not winning**: Calibration score (Brier-scored), argument quality (cross-adversarial), intellectual honesty (public belief revision), evidence contribution. NEVER aggregate into one number (Stack Overflow's failure mode).\n- **Separate well-argued from agreement**: This is load-bearing for the entire platform. A well-argued position you disagree with is the richest territory in deliberation.\n- **The graph is a map, not the product**: Mobile users are the primary data source. Their evaluations produce the quality scores that desktop analysts rely on. The graph is for orientation (10% of users), not the primary experience.\n\n### From the Pipeline Documents (extraction-pipeline-design, first-pipeline-run-analysis, improvement-research)\n\n- **DnDScore split validated**: Decomposing THEN decontextualizing produces 61% more atomic claims and 5x more definitional claims than combined approach. The tension between isolation and context-insertion is real.\n- **Criterion contestedness is pioneering**: No NLP paper addresses criterion contestedness (competing normative groundings for evaluative terms). The pipeline's detection of this is genuinely novel.\n- **Self-consistency voting improves quality**: Running Pass 2 three times with varied prompts and taking 2/3 consensus produces 7-point macro-F1 gain.\n- **Implicit premise recovery** (from MArgE and multi-agent debate research): Surfacing hidden \"because\" clauses that drive disagreements. Potentially transformative for later pipeline versions.\n- **Walton's 96 argument schemes** could provide structured critical questions: \"What's the evidence?\" → reveal more as argument develops. The \"VALID badge\" concept (progressive scheme validation) is novel.\n\n### From History and Primary Sources (meteor-prototype-analysis, chat-logs-analysis, simplenote-archive-analysis)\n\n- **Meteor prototype validated short-form claims**: Average 44 characters. Users naturally wrote concise, declarative statements. The substitution/rewording mechanism was the most-engaged feature — users iteratively refined formulations toward precision.\n- **Multi-axis voting worked**: 24% used \"Irrelevant\" — genuinely separating quality from agreement. Balanced participation (21 For vs 20 Against). Devil's advocate participation generated substantive counter-arguments.\n- **Simplenote genesis notes** (2011-2012): \"Coarse-grained discourse → fine-grained discourse = advent of language itself\" (from Leverage Research notes). Elevates the mission beyond \"better tool.\" Also: De Bono thinking hats as contribution funnels (never formalized in later docs), \"queues of needy claims\" (evidence gap detection as feed design, 2012).\n- **The \"Focus on trusting the process instead of the people\"** quote (2012 article outline) is the earliest formulation of trustless collaboration — the process (structured argumentation with machine-checked validity) is trustworthy even when the people are unknown.\n\n---\n\n## The \"Neither Pole Should Win\" Exploration (Session 7)\n\nA deep examination of the dialectic tenet, culminating in the naming decision \"attunement\" and a sharper formulation of how the poles relate.\n\n### Why It's Valuable\n\nThe tenet correctly identifies the two known failure modes: all-analysis (20+ platforms in the graveyard) and all-attunement (Habermas Machine — consensus without rigor). The unique value proposition IS the combination. \"Neither should win\" prevents collapsing to either known failure.\n\n### Where It Gets Dangerous\n\n1. **Decision avoidance**: Every concrete design choice leans one way. \"Neither should win\" doesn't help when staring at a specific UI component. The four-type classification IS analysis imposing structure on fluid discourse. The \"no copout axioms\" principle IS the analytical pole insisting on itself even in the domain of felt values.\n\n2. **McGilchrist's actual thesis is stronger**: He argues the right hemisphere (holistic, contextual — the attunement pole) should be the **Master**, and the left hemisphere (analytical, categorical) should be the **Emissary**. Not equal partners — analysis IN SERVICE OF attunement. vision.md already says this: \"use left-hemisphere capabilities in service of right-hemisphere understanding.\"\n\n3. **The current system is already analysis-dominant**: Everything built so far IS the analytical pole. The attunement pole (worldview filters, perspective-taking) is entirely future/aspirational. \"Neither should win\" may obscure this asymmetry.\n\n4. **Žižek's critique applies to the dialectic itself**: Maintaining \"productive tension\" without resolution is itself a liberal perspectivism move — the Beautiful Soul position appearing to transcend the tension while avoiding commitment.\n\n5. **Users will self-select into one pole**: Analytical users (rationalists) want QBAF weights and calibration. Attunement users (mediators, educators) want worldview filters and emotional resonance. Progressive disclosure partially solves this — different levels for different users. But \"neither should win\" doesn't address which user you design DEFAULTS for.\n\n### The Sharper Formulation\n\n**\"Neither should win\" is the right aspiration but the wrong operational principle.** More actionable alternatives emerged:\n\n- **\"Analysis serves attunement.\"** The formal structure exists to enable the human experience of understanding why someone thinks differently.\n- **\"Analysis is the scaffolding. Attunement is the building.\"** You need scaffolding to build, but people live in the building, not the scaffolding.\n- **\"The test of every analytical feature is whether it produces an attunement insight that wasn't possible without it.\"** Semantic disambiguation revealing \"oh, we meant different things\" = analysis producing attunement. QBAF weights that nobody understands = analysis winning over attunement.\n\n### The Naming: \"Empathy\" → \"Attunement\"\n\nThe dialectic was renamed from \"analysis ↔ empathy\" to **\"analysis ↔ attunement\"** based on:\n\n| Candidate | Punch | Warmth | Cognitive depth | Breadth |\n|---|---|---|---|---|\n| Empathy | HIGH | HIGH | MEDIUM — default reading is emotional | MEDIUM — misses holistic/perceptual |\n| **Attunement** | **HIGH** | **HIGH** | **HIGH** — perceptual sensitivity, not just emotion | **GOOD** — covers listening, sensing, inhabiting |\n| Comprehension | MEDIUM | LOW — cerebral, dry | HIGH (etymologically perfect) | LOW |\n| Communion | VERY HIGH | VERY HIGH | MEDIUM | MEDIUM — too spiritual |\n| Resonance | HIGH | HIGH | MEDIUM | GOOD — but passive |\n\n**\"Attunement\" won because**: it encompasses cognitive perspective-taking + emotional resonance + holistic understanding + active inhabitation. Like a musician attuning to an ensemble — a skill and a practice, not merely a feeling. Less likely to be read as merely \"be nice to people who disagree.\"\n\n**The etymological insight**: *ana-lysis* (Greek: \"loosening apart\") vs *com-prehendere* (Latin: \"grasping together\"). The dialectic is built into language itself. Analysis loosens apart; its counterpart grasps together. Added to vision.md §Core Dialectic.\n\n---\n\n## Ontology Is the Primary Concern (Session 7 — Fredrik's Declaration)\n\n**Fredrik (verbatim, Session 7):**\n\n> \"Getting the ontology right is a largely orthogonal and primary concern, it's my main priority. This needs to be documented!\"\n\nAnd on auto-population:\n\n> \"I love the direction of auto-populating Deliberus, but it MUST be done with discernment and wisdom.\"\n\n**Implication**: UI/UX can always be iterated — progressive disclosure means the frontend can always show less or more of the underlying complexity. But the ontology — what the graph's nodes and edges MEAN, how claims relate to premises, how schemes map to CQs, how evidence decomposes — is the structural foundation. Getting this right is more important than any UI decision. The frontend is a view of the ontology; the ontology is not a reflection of the frontend.\n\n---\n\n## Spec Session: Scheme-Bounded Decomposition (Session 7)\n\nSeven rounds of Socratic questioning produced a full implementation spec. Key decisions:\n\n| Question | Decision | Reasoning |\n|---|---|---|\n| Scheme detection unit | **Relationship edge** | Same claim participates in multiple inference patterns via different edges |\n| CQ authorship | **System-attributed** | Honest about provenance; CQs are structural prompts, not author assertions |\n| CQ graph representation | **Both polarities pre-generated** + Question node | Cleanest QBAF semantics; evidence accumulates on both sides |\n| Auto-connect behavior | **Auto-link, flag for review** | Low friction with safety net |\n| Negative CQ answer | **Badge + strength (both)** | Human-readable signal AND computational signal |\n| CQ recursion depth | **Flywheel-gated** | Depth determined by graph density; natural termination |\n| Scheme scope | **All 96 + cluster tags** | Ambitious; clusters provide coarse fallback for rare schemes |\n| Claim IDs | **UUID migration** (claim_{uuid}, cq_{uuid}) | Clean slate with Data Freshness Directive |\n| Multi-scheme detection | **Union of CQs, deduplicated** | Most thorough; catches all relevant questions |\n| CQ text specificity | **Parameterized** (specific + template) | Display: specific instantiation; Search: template type matching |\n| Badge terms | **Separate per scheme type** | VALID for deductive, CREDIBLE for source-based, etc. Maximum precision |\n| Badge computation | **QBAF propagation** | Full probabilistic; v1 uses weighted average |\n| Pipeline integration | **Pass 3a integrated + Pass 6 new** | Scheme detection saves calls; auto-connect decoupled |\n| Time budget | **Under 5 minutes** | Tight but achievable with 8-way LLM parallelism |\n| CQ decontextualization | **Stranger test enforced** | Every CQ claim fully standalone, matching Pass 2b rigor |\n\nFull spec files: `.claude/specs/scheme-bounded-decomposition/` (requirements.md, design.md, tasks.md)\n\n### Implementation Gotchas Identified\n\n1. **Both-polarity CQs create ~500 nodes per extraction** — graph becomes ~85% system-generated CQ scaffolding. This is either exactly right (scaffolding IS the contribution surface) or overwhelming (UI must hide most by default).\n2. **CQ parameterization NEEDS an LLM call** — design.md initially contradicted itself on this.\n3. **FalkorDB edge-to-edge workaround** — Question nodes with parent_edge_from/parent_edge_to properties; orphan cleanup needed if parent edge deleted.\n4. **Rare scheme accuracy ~50%** — cluster-level classification is the real workhorse; specific scheme is a refinement.\n5. **Cross-scheme template matching only works via embeddings, not template IDs alone** — analogous CQs from different schemes have different template IDs.\n\n### The Stranger Test Is a Feature, Not a Claimify Artifact\n\nDecontextualization was initially questioned (\"claims require context for meaning\") but reaffirmed as essential for Deliberus. The \"one graph, not canvases\" principle means claims are navigated from many directions — not just from the original source. Someone arriving at a UBI claim via a minimum wage debate needs the claim to be self-sufficient. The graph provides structural context (edges, parents, source); the claim text provides semantic self-sufficiency. CQ claims are inherently decontextualized (generated from templates with specific parameters). The parameterization prompt MUST enforce the same stranger test as Pass 2b.\n\n### Concrete Simulation: UBI Central Claim With CQs\n\nFor one central claim (\"UBI should be implemented to reduce poverty\") with 4 relationship edges:\n- 14 Questions generated (6 + 3 + 3 + 2 from four different schemes)\n- 28 polarity claims (14 positive + 14 negative)\n- = 42 system-generated nodes for ONE claim's relationships\n\nFull extraction (~74 claims, ~26 edges): 74 extracted + ~117 Questions + ~234 CQ polarity claims = **~425 total nodes (83% CQ scaffolding)**\n\nThe CQs would appear collapsed by default under each relationship, expandable on demand — progressive disclosure at work.\n\n---\n\n## Implementation Sprint (Mar 30-31, 2026)\n\nThe session evolved from reading → design → implementation, becoming the most productive sprint in the project's history. Everything built:\n\n### Pipeline Architecture\n- **Scheme detection** integrated into Pass 3a (44 Walton schemes, CQ templates)\n- **CQ generation** module (parameterized with stranger test enforcement)\n- **Auto-connect** Pass 6 (embedding pre-filter → LLM classification, 8-way parallel)\n- **QEM gradual semantics** replacing weighted average (Potyka, KR 2018)\n- **Temporal workflow** with 11 activities, 30s heartbeat recovery, cleanup on restart\n- **UUID migration** + 7-issue audit fix (all STARTS WITH prefix → EXTRACTED_FROM edge traversal)\n\n### Correction UX (THE product)\n- `POST /claims/{id}/decompose` — value premises get sub-claims extracted from user reasoning\n- `POST /claims/{id}/add-evidence` — supports/attacks with evidence (embedded for discovery)\n- `POST /claims/{id}/answer-cq` — answers CQs, shifts QBAF strength via QEM energy\n- Invitation cards (Aha Generator style) with functional buttons, polarity pills\n\n### Voting (Option B Decision)\n**Fredrik's question**: \"Is 'well-argued' a meaningful signal, or too subjective?\"\n**Answer**: Option B — no separate quality vote. The QBAF badge IS the quality signal (computed from CQ evidence). Human vote is ONLY agree/disagree. LessWrong found >95% of two-axis votes go the same direction; Deliberus's CQ system decomposes \"well-argued\" into concrete structural questions instead.\n\nImplementation: agree/disagree buttons per claim, Bayesian-smoothed percentages (virtual prior 5+5), User nodes in FalkorDB, AGREES_WITH/DISAGREES_WITH edges.\n\n### Feed Algorithm (8 modes)\nThe culmination — every piece feeds into this:\n\n**\"Bridging\" mode** (THE NOVEL SIGNAL): `bridging_score = disagreement_factor × QBAF_strength`. Claims where people disagree on the conclusion BUT the QBAF badge is green (CQs well-answered). \"I can't refute this reasoning, which makes me reconsider.\" Computable ONLY because Deliberus combines voting + structural argument analysis + scheme-based quality.\n\nOther modes: divisive (50/50 split), needs-help (sorry density), needs-votes, contested (semantic), recent, challenged (low QBAF), robust (high QBAF).\n\n### Visual Upgrades (Stitch-Inspired)\n- Bold stat counters (large numbers, gradient strip)\n- Color-coded argument group headers (by conclusion_type)\n- Option B relationships (grouped by target claim, scheme pills)\n- Contested concepts hero (large term, diagnosis + analysis, sense cards)\n- Invitation cards (Aha Generator style for sorry blocks)\n- Foundation doc cards on landing page (6 cards, /about page for full README)\n\n### Research\n- 35 recent academic papers surveyed (Oct 2025 – Mar 2026)\n- QBAF gradual semantics: 6 options compared, QEM selected\n- Unbuilt features: 8 researched with priority roadmap\n- Key finding: Toni's group at Imperial (AAMAS 2026) = Deliberus's thesis in academic form\n\n### Coherence Audit\n5 issues found and fixed: QBAF reads actual CQ evidence, CQs visible from both sides, user claims embedded, CQ answer UI, CQ polarity edges carry scheme info.\n\n### Counts\n- 297 tests passing\n- ~3,500 lines of new code\n- 53 research documents\n\n---\n\n## Cross-References\n\n- [vision.md](../vision.md) — The soul, key tensions, the core dialectic\n- [ux-principles.md](../ux-principles.md) — North Star UX principles\n- [conceptual-threads.md](../conceptual-threads.md) — 8 cross-cutting threads\n- [extraction-pipeline-design.md](extraction-pipeline-design.md) — Pipeline architecture decisions\n- [assumption-ranking.md](assumption-ranking.md) — 15 assumptions, weakest to strongest\n- [consensus-path-forward.md](consensus-path-forward.md) — Multi-model strategy consensus\n- [steelmanned-critiques.md](steelmanned-critiques.md) — 7 serious challenges\n- [cross-thread-synthesis.md](cross-thread-synthesis.md) — Thread interactions and gaps\n- [spec-session-reasoning.md](spec-session-reasoning.md) — Dialectic-is-fractal discovery\n- [scheme-bounded-decomposition-and-evidence-as-subgraph.md](scheme-bounded-decomposition-and-evidence-as-subgraph.md) — CQs as premises, auto-connect, evidence decomposition, cost analysis\n- `.claude/specs/scheme-bounded-decomposition/` — Full implementation spec (requirements, design, tasks). In the repo, not web-served, so referenced as a path rather than a link\n\n- [vision.md](../vision.md) — The soul, key tensions, the core dialectic\n- [ux-principles.md](../ux-principles.md) — North Star UX principles\n- [conceptual-threads.md](../conceptual-threads.md) — 8 cross-cutting threads\n- [extraction-pipeline-design.md](extraction-pipeline-design.md) — Pipeline architecture decisions\n- [assumption-ranking.md](assumption-ranking.md) — 15 assumptions, weakest to strongest\n- [consensus-path-forward.md](consensus-path-forward.md) — Multi-model strategy consensus\n- [steelmanned-critiques.md](steelmanned-critiques.md) — 7 serious challenges\n- [cross-thread-synthesis.md](cross-thread-synthesis.md) — Thread interactions and gaps\n- [spec-session-reasoning.md](spec-session-reasoning.md) — Dialectic-is-fractal discovery\n"}