{"path":"research/prediction-markets-argumentation.md","content":"# Prediction Markets and Structured Argumentation\n\n**Research date**: March 27, 2026\n**Purpose**: Exploring whether argument strength could be treated as a market, and how prediction market mechanisms relate to Deliberus's structured deliberation goals.\n\n---\n\n## Table of Contents\n\n1. [How Prediction Markets Work](#1-how-prediction-markets-work)\n2. [Robin Hanson's Futarchy](#2-robin-hansons-futarchy)\n3. [Epistemic Calibration at Scale](#3-epistemic-calibration-at-scale)\n4. [Metaculus Question Decomposition](#4-metaculus-question-decomposition)\n5. [Argument Markets: Existing Systems](#5-argument-markets-existing-systems)\n6. [Reputation as Currency](#6-reputation-as-currency)\n7. [Limitations: Normative Claims](#7-limitations-normative-claims)\n8. [Decentralized Models: Augur and Gnosis](#8-decentralized-models-augur-and-gnosis)\n9. [Design Sketch: Prediction-Market-Backed Argument Graph](#9-design-sketch-prediction-market-backed-argument-graph)\n10. [Implications for Deliberus](#10-implications-for-deliberus)\n\n---\n\n## 1. How Prediction Markets Work\n\n### Core Mechanism\n\nA prediction market lets participants buy and sell contracts on real-world outcomes. Each contract pays $1.00 if the event happens and $0.00 if it does not. The trading price IS the market's consensus probability. A contract at $0.65 means the crowd estimates a 65% chance of the outcome occurring.\n\nThe industry processed over **$44 billion in trading volume in 2025** (Kalshi alone reported $1B+ on the Super Bowl). This is no longer experimental.\n\n### Three Market Mechanisms\n\n**Continuous Double Auction (CDA)**\nThe mechanism used by Polymarket and most real-money markets. Buyers post bids, sellers post asks; trades execute when prices match. Advantages: familiar to traders, price discovery through natural supply/demand. Disadvantages: can suffer from thin liquidity (no trades happen if no counterparty exists), and probabilities across outcomes may not sum to 100% without arbitrageurs.\n\n**Logarithmic Market Scoring Rule (LMSR)**\nInvented by Robin Hanson. The market maker is an algorithm, not a counterparty. Core math:\n\n- Cost function: `C(q) = b * log(sum(e^(q_i/b)))` where `q` = outstanding shares per outcome, `b` = liquidity parameter\n- Instantaneous price of outcome i: `p_i(q) = e^(q_i/b) / sum(e^(q_j/b))`\n- Prices always sum to 1 (coherent probabilities by construction)\n- Maximum market maker loss is bounded: `WorstCaseLoss <= b * log(n)` where n = number of outcomes\n\nThe liquidity parameter `b` controls market depth. Higher b = deeper market requiring more volume to move prices. Lower b = more responsive to individual trades. LMSR guarantees continuous liquidity (you can always trade) without needing a counterparty. This is the mechanism most relevant to Deliberus because it works in thin markets (few participants) and naturally produces coherent probabilities.\n\n**Parimutuel**\nAll bets are pooled; payouts divided proportionally among winners after the event resolves. Simple but prices only finalize at close, so no continuous price signal during trading. Used historically in horse racing.\n\n### Why LMSR Matters for Deliberus\n\nLMSR solves the \"thin market problem\" that would plague any argumentation platform. You don't need thousands of active traders to get meaningful probability signals. Even a single trader updates the price. The bounded loss property means a platform could subsidize market making with a known maximum cost. Hanson's combinatorial LMSR extension allows markets on combinations and conditionals (e.g., \"probability of A given B\"), which maps directly to argument decomposition.\n\n---\n\n## 2. Robin Hanson's Futarchy\n\n### \"Vote on Values, Bet on Beliefs\"\n\nHanson's futarchy proposal (first published 2000, revised through 2024-2025) separates two functions of governance:\n\n1. **Values**: What outcomes do we want? (decided by democratic vote)\n2. **Beliefs**: Which policies achieve those outcomes? (decided by prediction markets)\n\nElected representatives define measurable welfare metrics. Speculators then bet on which policy options maximize those metrics. The policy with the highest predicted welfare wins.\n\n### Decision Markets (the mechanism inside futarchy)\n\nA decision market creates conditional prediction markets: \"What will GDP be if we adopt Policy A?\" vs. \"What will GDP be if we adopt Policy B?\" The market with the higher conditional estimate wins. Crucially, only the market corresponding to the chosen policy pays out (the other is voided), which incentivizes information aggregation only from people with genuine knowledge.\n\n### The Causal vs. Conditional Problem\n\nA major technical challenge: conditional prices reflect correlations, not necessarily causation. If a CEO's firing correlates with company distress (the board fires CEOs of troubled companies), the market might show \"stock price conditional on firing\" as low — not because firing causes decline, but because both share a cause. Hanson has argued (2025) that if traders apply the same decision theory as the decision-maker, conditional and causal estimates converge. This remains debated.\n\n### Relevance to Deliberus\n\nFutarchy's core insight maps powerfully to argumentation: separate the question \"what do we value?\" from \"what is true?\" In Deliberus terms:\n\n- **Normative claims** (\"We should reduce carbon emissions\") = voted on, not marketed\n- **Factual premises** (\"Policy X would reduce emissions by Y%\") = prediction-marketable\n- **Logical structure** (\"If premise P and premise Q, then conclusion C\") = the argument graph itself\n\nThis suggests a hybrid architecture where the argument graph holds the logical structure, markets price the factual premises, and voting or deliberation handles the value judgments.\n\n---\n\n## 3. Epistemic Calibration at Scale\n\n### What Makes Prediction Markets Accurate?\n\nPrediction markets consistently outperform polls, pundits, and expert panels. Key evidence:\n\n- **Polymarket demonstrates 73% accuracy** across resolved markets since 2023, outperforming traditional polls by 8 percentage points in political markets (Fensory Research, Feb 2026)\n- Across five US presidential elections (1988-2004), prediction markets provided more accurate estimates than **74% of individual polls** (cited in Manhattan West, Jan 2026)\n- Columbia Economic Review (Mar 2026) notes markets aggregate \"diverse, fragmented information\" that no single participant possesses\n\n### The Brier Score\n\nThe standard metric for probabilistic calibration. For a prediction p and outcome o (0 or 1):\n\n```\nBrier Score = (p - o)^2\n```\n\nA perfectly calibrated forecaster achieves a Brier score of 0. Random guessing on binary questions yields 0.25. Polymarket's aggregate Brier score across resolved markets is competitive with institutional forecasters (Marginal Revolution analysis, Oct 2025).\n\n### Why Markets Beat Polls\n\n1. **Skin in the game**: Financial incentives punish overconfidence and reward calibration\n2. **Marginal trader effect**: Markets need only a few informed traders to move prices toward truth, even if most participants are noise\n3. **Continuous updating**: Prices adjust in real-time as new information arrives\n4. **Self-correcting**: Mispricing creates profit opportunities that attract correcting trades\n\n### Calibration Curves\n\nA well-calibrated market should show: events priced at 70% should actually occur ~70% of the time. When plotting predicted probability vs. actual frequency across thousands of resolved markets, the best prediction markets hew close to the diagonal (perfect calibration). Known biases: slight overconfidence in high-probability events, and a \"favorite-longshot bias\" where extreme probabilities are slightly miscalibrated.\n\n### Good Judgment / Superforecasters\n\nPhilip Tetlock's IARPA tournament ($20M, 4 years) discovered that the top 2% of forecasters (\"superforecasters\") dramatically outperform others. Key traits:\n\n- **Treat beliefs as hypotheses**, not convictions\n- **Granular probability updates** (adjusting by small percentages, not rounding to nearest 10%)\n- **Active open-mindedness**: seek disconfirming evidence\n- Foxes (know many things) > Hedgehogs (know one big thing)\n\nGood Judgment's four keys: **talent-spotting, training, teaming, and aggregation**. Teams of superforecasters outperform both individuals and prediction markets. This suggests Deliberus could benefit from identifying and weighting high-calibration contributors, not just counting votes.\n\n---\n\n## 4. Metaculus Question Decomposition\n\n### Structure\n\nMetaculus supports five question types:\n- **Binary**: Will X happen by date Y? (probability 0-100%)\n- **Numeric range**: What will the value of X be? (continuous distribution)\n- **Multiple choice**: Which of {A, B, C, ...}?\n- **Conditional pairs**: P(A|B) and P(A|not-B) — predict the probability of A given that B happens or doesn't\n- **Question groups**: Related questions bundled together\n\n### Conditional Pairs\n\nLaunched February 2023. A conditional pair links two questions: \"What is the probability of A given B?\" and \"What is the probability of A given not-B?\" This enables forecasters to express beliefs about causal relationships (or at least correlations) between events.\n\nExample: \"Will AI cause a major cybersecurity incident by 2027?\" conditional on \"Will frontier AI models be open-sourced by 2026?\"\n\nThis maps directly to argument decomposition. An argument's conclusion depends on its premises. Metaculus-style conditional pairs formalize: \"How likely is the conclusion if premise P is true?\" vs. \"How likely is the conclusion if premise P is false?\"\n\n### Question Groups\n\nCollections of related sub-questions that can be combined. A complex question like \"Will climate change cause >1M refugees by 2030?\" could decompose into sub-questions about temperature trajectories, agricultural impact, political stability, and migration patterns.\n\n### Mapping to Argument Decomposition\n\n| Metaculus Concept | Argument Graph Analogue |\n|---|---|\n| Binary question | Claim confidence |\n| Conditional pair | Premise-conclusion dependency |\n| Question group | Argument bundle (multiple premises supporting a conclusion) |\n| Numeric range | Magnitude of an effect claim |\n| Resolution criteria | Verifiability standard for a claim |\n\nThe key insight: Metaculus already implements a primitive form of \"argument decomposition via probabilistic conditioning\" — but without the explicit logical structure of an argument graph. Deliberus could add the argument graph as the organizing skeleton, with Metaculus-style probabilistic questions at each node.\n\n---\n\n## 5. Argument Markets: Existing Systems\n\n### argue.fun — \"The First Argumentation Market\" (2025-2026)\n\nBuilt on Base blockchain, powered by GenLayer. The most direct attempt at merging markets with argumentation.\n\n**Mechanism**:\n1. Users pick Side A or B in a debate, stake $ARGUE tokens (minimum 1 token)\n2. When placing bets, participants **must submit their reasoning** (not just a position)\n3. Opposing agents identify weaknesses and counter-stake with counter-arguments\n4. At market close, a **multi-LLM jury** (multiple models running different LLMs) evaluates argument quality\n5. Winners split the pool proportionally to stake\n\n**Critical design choice**: \"The jury evaluates argument quality, not pool size.\" This means argument quality is judged separately from market dynamics — the market incentivizes participation, but resolution depends on AI evaluation.\n\n**Limitations**: LLM jury creates a centralized trust bottleneck. Argument evaluation by current LLMs is not transparent or formally verifiable. The system is adversarial (two sides only) rather than structured (argument graph with support/attack relations).\n\n### Contro — \"Debate Markets\" (2025-2026)\n\nContro combines two technologies:\n1. **GLOB-powered prediction markets** (Gradual Limit Order Books — a novel mechanism that replaces instant orders with \"steady streams\" matching supply and demand at uniform prices, designed to prevent frontrunning and sniping)\n2. **AI-based judgment** that listens to every comment, extracts substance, fact-checks claims, and scores each contribution\n\nKey design principles:\n- Influence scales with **both stake size AND contribution quality**\n- \"Earned authority\": track record matters, not just current stake\n- Structured evaluation surfaces best arguments through merit, not volume\n\nContro explicitly frames itself as creating \"antifragile markets\" — markets that get stronger from disagreement rather than weaker. The GLOB mechanism is designed to be manipulation-resistant by eliminating speed advantages.\n\n**Status**: Conceptual/early-stage. Technical details on resolution mechanisms for normative debates promised but not yet published.\n\n### Controvis — Argument Maps (separate project)\n\nAn argument mapping tool (open beta) that creates visual networks of supporting, attacking, and citing arguments. Not market-backed, but demonstrates the visualization layer that could complement market mechanisms. Users can navigate interactive argument maps to understand positions on polarizing topics.\n\n### Persuasion Markets (Mike Elias, 2023)\n\nA theoretical proposal for prediction markets where **self-reported opinions serve as the settlement oracle**:\n\n- Bet on whether a named public figure will agree with a proposition by a deadline\n- Bet on whether >50% of a group holding belief X will also adopt belief Y\n- Settlement = the oracle's stated opinion at the deadline\n\n**Key innovation**: Extends prediction markets to \"matters of judgment or interpretation rather than only matters of data.\" This directly addresses the normative-claim limitation.\n\nFour simultaneous functions:\n1. **Trust measurement** — reveals which figures users actually trust\n2. **Performance tracking** — accountability records for oracles\n3. **Rapid iteration** — dynamic reputation adjustment\n4. **Financial incentives** — compensates oracles (~5% of wagers)\n\n**Incentive alignment**: The most profitable strategy is identifying unpopular conclusions with strong evidence, then betting that open-minded figures will adopt them. Profit motive aligns with truth-seeking.\n\n### Argumentation-Based Information Exchange in Prediction Markets (ArgMAS 2008)\n\nAcademic research by Hadoux et al. investigating how argumentation processes among agents affect prediction market outcomes. Key finding: when agents can argue (exchange reasoning, not just prices) within social networks, group judgment quality improves. Different social network topologies and data distributions affect the magnitude of improvement.\n\nThis is the most direct academic precedent for combining argumentation with prediction markets, though it treats argumentation as a communication layer between market participants rather than as the market structure itself.\n\n### ARGORA (2026)\n\nA multi-expert LLM system that organizes discussions into explicit argumentation graphs showing which arguments support or attack each other. By casting these graphs as **causal models**, ARGORA can systematically remove individual arguments and recompute outcomes, identifying which reasoning chains were necessary. Published at AAMAS 2026 (International Conference on Autonomous Agents and Multi-Agent Systems).\n\n---\n\n## 6. Reputation as Currency\n\n### Metaculus Track Records\n\nMetaculus tracks each forecaster's:\n- **Calibration** (do events they predict at 70% happen ~70% of the time?)\n- **Resolution** (accuracy on resolved questions)\n- **Peer score** (how much better/worse than the community median)\n- **Coverage** (number of questions predicted on)\n\nScoring is **hybrid**: part absolute accuracy (rewarded even if alone on a question), part relative (compared against other forecasters). This encourages both broad participation and deep specialization.\n\n**Perverse incentive** (identified by Scott Alexander): quick, lazy predictions across many questions can accumulate more points than deeply researched predictions on few questions, because the time investment per marginal point is lower for broad coverage.\n\n### Manifold Markets Mana\n\nManifold uses play money (Mana) as its currency. Free on signup, earned through accurate predictions, spendable on creating new markets. Mana can be purchased but never cashed out.\n\n**Reputation signal**: \"play money profits\" (how much Mana gained through trading) serve as a reputation proxy. The system separates signal (forecasting skill) from noise (time invested, luck).\n\n**Limitation**: \"Nobody has ever asked me my Metaculus score before deciding how much to trust me\" — Scott Alexander. These scores remain trapped within their platforms. External credibility transfer is unsolved.\n\n### The Reputation Problem for Deliberus\n\nSeveral design tensions:\n\n1. **Absolute vs. relative scoring**: Absolute rewards participation even in thin markets (good for early Deliberus). Relative scoring better surfaces genuine skill but discourages participation when few others engage.\n\n2. **Zero-sum vs. positive-sum**: Zero-sum imitates real markets but discourages exploration. Positive-sum (platform subsidizes accuracy rewards) encourages participation but costs money.\n\n3. **Time horizon**: Prediction market reputation accumulates over many resolved questions. Deliberus arguments may not \"resolve\" for years or ever. How do you score someone's reputation on claims that haven't been tested?\n\n4. **Skill transferability**: A user with a strong track record on geopolitics should carry some credibility to adjacent domains but less to unrelated ones. Domain-specific reputation scores would be more informative but harder to compute.\n\n### Possible Deliberus Design: Epistemic Capital\n\nA reputation currency earned through:\n- **Calibration**: Making well-calibrated probability estimates on factual claims\n- **Argument quality**: Having arguments that withstand scrutiny (low successful-attack rate)\n- **Predictive record**: Past claims that were subsequently verified\n- **Intellectual honesty**: Updating positions when presented with strong counter-evidence (measurable via position-change tracking)\n\nEpistemic capital could weight a user's influence in the argument graph — not their vote, but the default visibility and trust level of their contributions.\n\n---\n\n## 7. Limitations: Normative Claims\n\n### The Core Problem\n\nPrediction markets work because they have **resolution criteria** — objective events that either happen or don't. \"Will GDP grow by 3%?\" resolves against measured data. \"Should we prioritize economic growth over environmental protection?\" has no objective resolution.\n\nThis is the fact/value boundary that Deliberus must navigate.\n\n### Attempted Solutions\n\n**Futarchy's approach**: Separate values from beliefs entirely. Vote on what we want, bet on how to get it. This works when the disagreement is about means, not ends. It fails when people genuinely disagree about what counts as a good outcome.\n\n**Persuasion markets**: Use self-reported opinions as oracles. \"Will 60% of informed participants agree with X after reading the evidence?\" This marketizes persuasion, not truth. It's useful for measuring argument effectiveness but doesn't determine whether the argument is *correct*.\n\n**Conditional welfare markets**: \"If we adopt policy X, what will happiness index Y be in 5 years?\" This requires:\n- A pre-agreed welfare metric (the normative choice)\n- A measurable proxy (the empirical challenge)\n- Long resolution timelines (the patience problem)\n\n**AI judge systems** (argue.fun, Contro): Use LLMs to evaluate argument quality. This outsources the normative judgment to a model, raising questions about whose values the model encodes and whether its judgments are stable/consistent.\n\n### The Deeper Tension\n\nMany disagreements that appear factual are actually normative in disguise:\n- \"Climate change will cause X damages\" embeds choices about discount rates, value of statistical life, and intergenerational equity\n- \"Immigration increases crime\" depends on which definition of crime, which timeframe, and which population comparison\n\nDeliberus's argument graph could make these hidden normative assumptions explicit by decomposing claims into their factual and value components. The factual components get markets; the value components get deliberation.\n\n### What Cannot Be Marketed\n\n- Pure value judgments (\"Freedom matters more than equality\")\n- Aesthetic preferences (\"This policy is elegant\")\n- Identity claims (\"This is who we are as a community\")\n- Moral imperatives (\"Torture is always wrong regardless of consequences\")\n\nThese require different aggregation mechanisms: ranked voting, consensus building, deliberative democracy, or simply transparent disagreement.\n\n---\n\n## 8. Decentralized Models: Augur and Gnosis\n\n### Augur\n\nThe first major decentralized prediction market (launched 2015, v2 in 2020). Built on Ethereum.\n\n**Oracle mechanism**: No central authority resolves markets. Instead:\n1. **Reporters** stake REP (Reputation) tokens on the outcome they believe is correct\n2. Disputes escalate through progressively larger bonds\n3. If bonds reach a threshold, REP splits into multiple versions (one per outcome) — a \"fork\" that forces the entire community to choose\n4. Honest reporting is designed to be always more profitable than manipulation\n\n**Resolution speed**: v1 took ~7 days minimum; v2 reduced to ~24 hours for uncontested markets.\n\n**Relevance to Deliberus**: The staking-and-dispute mechanism is structurally similar to argumentation. Asserting an outcome = making a claim. Disputing = attacking with a counter-claim. Escalation = the argument becoming more important and requiring more evidence (higher stakes). The fork mechanism (community splits over an unresolvable dispute) is an extreme form of \"agreeing to disagree\" with financial consequences.\n\n### Gnosis\n\nAnother Ethereum-based prediction market platform, now evolved into a broader DeFi ecosystem. Gnosis developed the **conditional token framework** — ERC-1155 tokens that can represent positions in combinatorial prediction markets. This allows nesting conditions arbitrarily deep: tokens representing \"A given B given C.\"\n\n**Conditional token framework relevance**: This is essentially an on-chain implementation of combinatorial LMSR, enabling exactly the kind of \"argument graph with market prices at each node\" that Deliberus might want. Each node in the argument graph could be a conditional token whose value depends on the resolution of its parent nodes.\n\n### Decentralization vs. Deliberus's Goals\n\n**Advantages of decentralized design for Deliberus**:\n- Censorship resistance (no single entity can suppress arguments)\n- Transparent rules (smart contract logic is auditable)\n- Permissionless participation\n- Immutable argument history\n\n**Disadvantages**:\n- Transaction costs (gas fees add friction to every interaction)\n- Speed (blockchain finality is slower than centralized databases)\n- Complexity (smart contracts are hard to upgrade/fix)\n- User experience (wallet management, key custody)\n\n**Pragmatic recommendation**: Deliberus should NOT start with blockchain. The argument graph and market mechanisms can be built centrally first (faster iteration, better UX). Decentralization can be added later if censorship resistance becomes essential. The conditional token framework's *design patterns* are valuable regardless of whether they run on a blockchain.\n\n---\n\n## 9. Design Sketch: Prediction-Market-Backed Argument Graph\n\n### Architecture Overview\n\n```\n                    ┌─────────────────────┐\n                    │   CONCLUSION        │\n                    │   \"Policy X is      │\n                    │    effective\"        │\n                    │   [no market -      │\n                    │    computed from     │\n                    │    premises]         │\n                    └────────┬────────────┘\n                             │\n                   ┌─────────┴──────────┐\n                   │                    │\n          ┌────────▼────────┐  ┌────────▼────────┐\n          │  PREMISE 1      │  │  PREMISE 2      │\n          │  \"X reduces Y   │  │  \"Cost of X     │\n          │   by 15%\"       │  │   < $1B\"        │\n          │  [LMSR: 0.72]   │  │  [LMSR: 0.45]   │\n          └────────┬────────┘  └────────┬────────┘\n                   │                    │\n          ┌────────▼────────┐  ┌────────▼────────┐\n          │  EVIDENCE       │  │  ATTACK         │\n          │  \"Study by Z    │  │  \"CBO estimate  │\n          │   shows 18%     │  │   says $1.3B\"   │\n          │   reduction\"    │  │  [LMSR: 0.61]   │\n          │  [credibility:  │  └─────────────────┘\n          │   0.84]         │\n          └─────────────────┘\n```\n\n### Node Types and Market Rules\n\n**Factual claims** (empirical, verifiable):\n- Full LMSR market with real resolution criteria\n- Price = market's probability estimate that the claim is true\n- Resolution: when evidence becomes available (study published, data released, event occurs)\n- Example: \"Global average temperature will exceed 1.5C by 2030\"\n\n**Logical connections** (argument structure):\n- No separate market; strength computed from connected nodes\n- \"Premise P supports Conclusion C with strength S\" where S reflects both P's probability and the conditional relevance P(C|P)\n- Could use Metaculus-style conditional pairs: P(C|P) and P(C|not-P)\n\n**Normative claims** (value judgments):\n- No LMSR market (no objective resolution)\n- Instead: weighted voting, deliberative scoring, or persuasion-market mechanics\n- Weight = function of participant's epistemic capital and declared stakeholder status\n- Example: \"Environmental protection should be prioritized over short-term economic growth\"\n\n**Evidence nodes** (supporting data):\n- Credibility score derived from source reliability, methodology assessment\n- Could be LLM-assisted with human review\n- Links to factual claims with measured evidential strength\n\n### Market Mechanics\n\n**Subsidized LMSR for each factual claim**:\n- Platform sets initial liquidity parameter `b` (low for niche claims, higher for major ones)\n- Maximum subsidy cost per claim = `b * log(2)` ≈ `0.693 * b` (for binary outcomes)\n- Users trade by buying/selling YES/NO shares\n- Price reflects consensus probability\n\n**Conditional markets for argument links**:\n- For each \"Premise P supports Conclusion C\" link, create conditional pair:\n  - Market 1: P(C is true | P is true)\n  - Market 2: P(C is true | P is false)\n- The difference reveals the **evidential relevance** of P to C\n- If both conditional probabilities are similar, P is irrelevant to C regardless of its truth value\n\n**Argument strength computation**:\n- Conclusion confidence = f(premise probabilities, conditional relevances, attack strengths)\n- Could use Dung's semantics (accepted/defeated/undecided) weighted by market prices\n- Or a Bayesian network where market prices provide the conditional probability tables\n\n### Incentive Design\n\n**What participants earn**:\n- Profit from correctly-priced factual claims (standard prediction market gains)\n- Epistemic capital from track record of accurate probability estimates\n- Reputation from constructing arguments that withstand attack\n- Discovery rewards for identifying new relevant premises or evidence\n\n**What the platform costs**:\n- LMSR subsidy per claim (bounded and known in advance)\n- Computation for argument strength propagation\n- Moderation / quality control\n\n**Anti-manipulation**:\n- LMSR naturally resists manipulation (moving the price is expensive, and manipulation creates arbitrage opportunities)\n- Attack/support structure is public and auditable\n- Sock puppet detection through interaction graph analysis\n- Epistemic capital requires sustained accurate predictions (hard to fake)\n\n### Resolution Mechanisms\n\n| Claim Type | Resolution | Timeline |\n|---|---|---|\n| Near-term factual | Observable outcome | Days to months |\n| Long-term factual | Designated data source | Months to years |\n| Scientific | Expert consensus + replication | Years |\n| Normative | No resolution (ongoing deliberation) | Never |\n| Conditional | Resolves when parent resolves | Cascading |\n\nFor long-term claims, the market provides a living probability estimate that updates as new evidence arrives. The claim doesn't need to resolve for the market to be useful — the price trajectory itself is informative.\n\n---\n\n## 10. Implications for Deliberus\n\n### What Prediction Markets Offer\n\n1. **Calibrated probability aggregation** — turning scattered beliefs into coherent probability estimates\n2. **Incentive alignment** — rewarding accuracy, not persuasiveness or volume\n3. **Information revelation** — surfacing private knowledge through trading\n4. **Anti-noise** — expensive-to-maintain positions filter out casual opinions\n5. **Continuous updating** — prices reflect current best estimates, not historical snapshots\n\n### What Prediction Markets Cannot Do (and Deliberus Must Handle Differently)\n\n1. **Resolve normative questions** — markets need objective resolution criteria\n2. **Capture logical structure** — markets price individual propositions, not the relationships between them\n3. **Handle novel arguments** — markets react to information but don't generate new reasoning\n4. **Build understanding** — a price is a conclusion without an explanation; Deliberus needs the \"why\"\n\n### The Hybrid Thesis\n\nDeliberus should not be a prediction market. Deliberus should not ignore prediction markets. The argument graph is the primary structure; markets provide calibrated probability estimates at specific nodes.\n\n**Argument graph** = the reasoning structure (what implies what, what attacks what)\n**Markets** = the confidence scores on factual nodes (how likely is each premise)\n**Deliberation** = the human process of constructing, evaluating, and improving arguments\n**Voting/consensus** = how values and priorities are aggregated\n\n### Key Design Decisions for Deliberus\n\n1. **Should markets use real money or play money?**\n   Real money produces better calibration but creates regulatory burden and excludes participants. Play money (Manifold-style Mana) lowers barriers but produces weaker signals. A hybrid could work: play money for most claims, real money for high-stakes factual questions.\n\n2. **How do you bootstrap liquidity?**\n   LMSR solves this with bounded subsidies. The platform acts as initial market maker for every claim. Cost is predictable: `b * log(2)` per binary claim.\n\n3. **What resolves long-term factual claims?**\n   Possible approaches: designated expert panels, Augur-style staking disputes, AI-assisted evidence evaluation, or simply leaving claims as perpetual markets with running probability estimates.\n\n4. **How does reputation transfer across domains?**\n   Track calibration per topic area. A user's geopolitics track record should carry partial weight in related fields but minimal weight in, say, biochemistry.\n\n5. **How formal should the argument structure be?**\n   Enough structure to enable automated strength computation (support/attack relations, premise-conclusion links) but not so much that participation requires formal logic training.\n\n### Open Research Questions\n\n- Can LMSR be adapted for arguments with continuous (not binary) truth values?\n- How should argument strength propagate through long inference chains without compounding uncertainty to uselessness?\n- Can persuasion markets (opinion-as-oracle) produce calibrated estimates on normative questions, or do they just measure rhetoric?\n- What is the minimum viable community size for useful argument markets?\n- How do you prevent \"argument insider trading\" (publishing a study, then trading on its impact before others can evaluate it)?\n\n---\n\n## Sources\n\n### Prediction Market Mechanisms\n- [How Prediction Markets Work: Complete Guide for 2026](https://predictreport.io/blog/how-prediction-markets-work) — PredictReport\n- [How Prediction Markets Work: A Complete Guide](https://marketmath.io/blog/how-prediction-markets-work) — Market Math\n- [LMSR (Logarithmic Market Scoring Rule)](https://blog.gensyn.ai/lmsr-logarithmic-market-scoring-rule/) — Gensyn/Delphi\n- [Comparing Prediction Market Mechanisms](https://www.jasss.org/21/1/7.html) — JASSS (Klingert)\n- [Combinatorial Information Market Design](https://hanson.gmu.edu/combobet.pdf) — Robin Hanson (2003)\n\n### Futarchy and Decision Markets\n- [Shall We Vote on Values, But Bet on Beliefs?](https://hanson.gmu.edu/futarchy2007.pdf) — Robin Hanson (2007)\n- [Futarchy Details](https://www.overcomingbias.com/p/futarchy-details) — Robin Hanson (2024)\n- [Federal Futarchy](https://www.overcomingbias.com/p/federal-futarchy) — Robin Hanson (2025)\n- [Decision Selection Bias](https://www.overcomingbias.com/p/decision-selection-bias) — Robin Hanson (2024)\n- [Decision Conditional Prices Reflect Causal Chances](https://www.overcomingbias.com/p/decision-conditional-prices-reflect) — Robin Hanson (2025)\n- [Summary and Takeaways: Hanson's Futarchy](https://forum.effectivealtruism.org/posts/ijohdoDbPvdeXMpiz/) — EA Forum (Lizka, 2021)\n- [Decision Rules and Decision Markets](https://www.ifaamas.org/Proceedings/aamas2010/pdf/01%20Full%20Papers/13_03_FP_0224.pdf) — Othman & Sandholm (2010)\n\n### Calibration and Accuracy\n- [Prediction Markets Are Very Accurate](https://marginalrevolution.com/?p=91721) — Marginal Revolution (Tabarrok, 2025)\n- [Polymarket Prediction Accuracy: Track Record & Brier Score](https://www.fensory.com/intelligence/predict/polymarket-accuracy-analysis-track-record-2026) — Fensory (2026)\n- [Prediction Markets as \"Truth Machines\"](https://sites.lsa.umich.edu/mje/2026/03/14/prediction-markets-as-truth-machines/) — Michigan Journal of Economics (2026)\n- [Polls, Pundits, or Prediction Markets](https://researchers.one/article/2018-11-6) — Harry Crane (2018)\n- [Prediction Market FAQ](https://www.astralcodexten.com/p/prediction-market-faq) — Scott Alexander\n\n### Reputation Systems\n- [Play Money And Reputation Systems](https://www.astralcodexten.com/p/play-money-and-reputation-systems) — Scott Alexander (2022)\n- [Manifold Markets Review 2026](https://www.zogby.com/reviews/manifold-markets) — Zogby\n- [Can You Rationally Disagree with a Prediction Market?](https://www.brownjppe.com/nickwhitaker) — Brown JPPE (Whitaker, 2021)\n\n### Argument Markets and Related Systems\n- [argue.fun](https://www.argue.fun/docs) — Argumentation Markets (2025-2026)\n- [Contro: Put Up or Shut Up](https://contro.tech/blog/2025-03-14-debatemarkets) — Contro (2025)\n- [Controvis](https://www.controvis.com/) — Argument Maps (2024)\n- [Persuasion Markets](https://www.mikeelias.com/p/persuasion-markets) — Mike Elias (2023)\n- [Argumentation-Based Information Exchange in Prediction Markets](https://link.springer.com/chapter/10.1007/978-3-642-00207-6_11) — ArgMAS 2008\n- [ARGORA: Orchestrated Argumentation for Causally Grounded LLM Reasoning](https://arxiv.org/html/2601.21533v1) — AAMAS 2026\n\n### Decentralized Markets\n- [Augur v2 Whitepaper](https://ar5iv.labs.arxiv.org/html/1501.01042) — Peterson et al. (2024)\n- [Augur v2 Resolution System](https://augur.net/blog/v2-resolution/) — Augur (2020)\n- [Decentralized Prediction Markets](https://timroughgarden.github.io/fob21/reports/ZLRL.pdf) — Columbia CS (2021)\n\n### Superforecasting\n- [Good Judgment Open Training Resources](https://www.gjopen.com/training/)\n- [The Science of Superforecasting](https://goodjudgment.com/about/the-science-of-superforecasting/) — Good Judgment\n- [Beliefs as Hypotheses: The Superforecaster's Mindset](https://goodjudgment.com/superforecasters-toolbox-beliefs/) — Good Judgment (2024)\n\n### Limitations and Critical Analysis\n- [Use with Caution: The Hidden Risks of Prediction Markets](https://cer.econ.columbia.edu/news/use-caution-hidden-risks-prediction-markets) — Columbia Economic Review (2026)\n- [The Epistemology of Prediction Markets](https://medium.com/@devdollzai/the-epistemology-of-prediction-markets-81a79de49488) — Grossi (2026)\n- [Prediction Markets: Truth Engine or Casino?](https://whirligigbear.substack.com/p/prediction-markets-truth-engine-or) — Whirligig Bear (2025)\n"}