{"path":"research/epistemic-gamification.md","content":"# Epistemic Gamification: Rewarding Good Reasoning Without Rewarding Winning\n\n*Research compiled March 28, 2026. Sources cited inline.*\n\n---\n\n## Overview\n\nThe central design challenge for Deliberus is not technical — it is motivational. How do you reward truth-seeking, calibrated reasoning, and intellectual honesty without accidentally rewarding debate performance, tribal signaling, or volume? Every prior structured argumentation platform either avoided gamification entirely (and died of indifference) or gamified the wrong things (and amplified adversarial dynamics). This document surveys the mechanisms that work, the anti-patterns to avoid, and proposes a concrete reputation architecture for Deliberus.\n\n---\n\n## 1. Prediction Market Platforms: Gamifying Accuracy, Not Conviction\n\n### Metaculus\n\nMetaculus ([metaculus.com](https://metaculus.com)) is the most relevant model for Deliberus because it rewards calibration over conviction — the score is not \"did you bet more than others?\" but \"were your probability estimates well-calibrated?\"\n\n**The scoring mechanism.** Metaculus uses a variant of the logarithmic scoring rule, which has a unique mathematical property: it is *strictly proper*, meaning no rational actor benefits from misreporting their true belief. The score for a prediction $p$ on an outcome that resolves to $o \\in \\{0, 1\\}$:\n\n```\nlog_score = log₂(p) if o = 1\n           log₂(1 - p) if o = 0\n```\n\nA forecaster who genuinely believes the probability is 70% gets the highest expected score by reporting 70% — not 80% to seem bold, not 50% to hedge. This is the key property missing from all debate-style scoring.\n\n**Calibration vs. resolution.** Metaculus tracks two distinct metrics: *resolution* (raw accuracy on individual questions) and *calibration* (whether your stated 70% predictions come true about 70% of the time). Good calibration is harder to fake than good resolution because it requires maintaining accurate uncertainty estimates across hundreds of questions over time. Metaculus displays calibration curves that show your historical probability assignments plotted against actual outcome frequencies — a powerful mirror for overconfidence.\n\n**2024 redesign: medals and leaderboards.** In late 2023, Metaculus introduced medals (bronze/silver/gold/platinum) and domain-specific leaderboards, replacing a single global reputation number. A forecaster can be a Gold-level contributor on AI safety questions while being a Bronze on geopolitics. This domain-specificity is important: it prevents the \"debate champion\" dynamic where one person's aggregate score drowns out domain experts. The leaderboard shows relative Brier score against the crowd median — not absolute correctness, but skill relative to peers. ([Metaculus Introduces New Forecast Scores, EA Forum](https://forum.effectivealtruism.org/posts/FodvZaiKftDCHPTub/metaculus-introduces-new-forecast-scores-new-leaderboard-and))\n\n**What makes users return daily.** Three mechanisms drive daily engagement on Metaculus: (1) questions with imminent resolution deadlines create urgency, (2) the community forecast line updates in real time so your prediction visibly moves \"the needle,\" and (3) prediction streaks and tournament leagues create mild competitive pressure without win/lose framing. The absence of a \"wrong answer\" stigma (probabilities allow graceful degradation — 65% on a wrong outcome is better than 95%) keeps participation psychologically safe.\n\n### Manifold Markets\n\nManifold ([manifold.markets](https://manifold.markets)) uses play money (Mana) and has run the most extensive A/B testing of gamification mechanics among prediction platforms.\n\n**Leagues.** Seasonal leagues (monthly) group 25 users by activity level. Mana earned during the season determines whether you move up or down a tier. At peak engagement, the leagues tab was the third most visited page on the site — people checked their standing obsessively. Leagues increased retention primarily through competitive social comparison (the \"trophy on the shelf\" dynamic) without requiring real money stakes. ([Manifold Leagues](https://manifold.markets/leagues))\n\n**Streaks.** Daily prediction streaks start at M5 and increase by M5 per consecutive day, capped at M25/day. The critical design insight from Duolingo research applies here: streaks leverage *loss aversion* more than gain motivation. Users are more motivated by avoiding losing a streak than by accumulating the streak reward. Manifold's streaks require at least one prediction daily — but since predictions can be trivial, this risks rewarding activity over quality.\n\n**Calibration city.** Manifold's calibration visualization shows predictions as dots sized by number of questions, plotted against actual outcome rates — providing the same calibration curve as Metaculus but in a more playful visual format. It builds the habit of thinking about prediction accuracy rather than just prediction confidence. ([Manifold FAQ](https://docs.manifold.markets/faq))\n\n**Critical analysis.** EA Forum research found Manifold Markets has significant problems: the play-money economy can decouple from epistemic incentives (gaming Mana through social manipulation is possible), and the market creation feature floods the platform with low-value questions. The lesson: Mana being non-cash-out-able preserves epistemic incentives but creates a floor problem — why care at all if the currency is worthless? Manifold's answer is social status (leaderboard rank), which partially works but also creates vanity metric dynamics. ([Manifold markets isn't very good — EA Forum](https://forum.effectivealtruism.org/posts/EaR9xFxspmYRkm3eo/manifold-markets-isn-t-very-good))\n\n### Good Judgment Open / Superforecasters\n\nPhilip Tetlock's superforecaster research (IARPA tournament, $20M, 5 years) identified traits that predict high-calibration forecasting. The key gamification insight: **identify the top 2% by measured calibration, name them, and let that status be the reward.** The title \"superforecaster\" carries genuine epistemic prestige because it is earned through verifiable accuracy metrics, not self-promotion.\n\nGood Judgment uses the *Relative Brier Score* — comparing your score to the crowd median rather than measuring absolute accuracy. This keeps the metric meaningful even when questions are trivially easy or hard. Elite superforecaster teams achieve Brier score reductions of 0.05–0.10 (10–25% improvement) over the unfiltered crowd. ([Good Judgment Project Wikipedia](https://en.wikipedia.org/wiki/The_Good_Judgment_Project); [Evidence on good forecasting practices — EA Forum](https://forum.effectivealtruism.org/posts/W94KjunX3hXAtZvXJ/evidence-on-good-forecasting-practices-from-the-good-1))\n\n**Implications for Deliberus:** A \"Deliberus Superforecaster\" title earned through measured calibration on factual claims would be a more valuable signal than any badge or point total — because it is externally verifiable and hard to fake.\n\n---\n\n> **Companion added Aug 2026.** This doc covers reward *mechanisms* — what to gamify, what not to, how to resist gaming. [incentives-analysis.md](incentives-analysis.md) covers the layer beneath: which participants have a reason to act at all, scored without assuming altruism. Two findings from it bear directly on the mechanisms below. **Supports are self-serving while attacks on one's own claim are altruistic**, so any reward scheme that treats the two symmetrically will over-collect supports and inflate QBAF strength invisibly. And **Community Notes has already solved the hardest version of this problem in production** — publication is the reward, cross-tribal rater agreement is the gate, so contributors are rewarded precisely for writing what the other side will sign — which makes it the most valuable incentive design to study here, not merely a bridging precedent.\n\n## 2. Epistemic Virtues Worth Rewarding — and How to Measure Them\n\nPhilosophy identifies a core set of epistemic virtues: open-mindedness, calibration, intellectual humility, curiosity, precision, honesty, and constructiveness. ([Virtue Epistemology — Stanford Encyclopedia of Philosophy](https://plato.stanford.edu/entries/epistemology-virtue/)) The challenge is operationalizing these into measurable platform behaviors.\n\n| Epistemic Virtue | Observable Platform Behavior | Measurement Approach |\n|---|---|---|\n| **Calibration** | Probability estimates match resolution rates | Log scoring rule across resolved claims |\n| **Intellectual honesty** | Publicly revising stated position after counter-evidence | Explicit position-change events, visible edit history |\n| **Steelmanning** | Strengthening the opposing argument before attacking it | Peer review rating: \"Does this accurately represent the opposing view?\" |\n| **Evidence provision** | Attaching sources, data, citations to claims | Count + quality score of cited evidence nodes |\n| **Precision** | Making falsifiable, specific claims (not vague assertions) | Verifiability rating: \"Can this claim be meaningfully tested?\" |\n| **Acknowledgment** | Responding \"you've convinced me\" and explaining why | Delta system (see r/ChangeMyView) |\n| **Bridge-building** | Arguments that receive support across ideologically distinct groups | Cross-group endorsement ratio (Polis-style clustering) |\n| **Constructive disagreement** | Quality counter-arguments that engage the substance | Peer rating of counter-argument quality, not agreement |\n\n### The Delta System: The Best Existing Model for Intellectual Honesty Reward\n\nReddit's r/ChangeMyView (CMV) ([Wikipedia](https://en.wikipedia.org/wiki/R/changemyview)) invented a practical mechanism: the delta symbol (∆) awarded to commenters who genuinely changed the original poster's view. Rules are strict — a delta requires an explanation of what changed and why, preventing strategic delta-farming. Deltas are tracked in user flairs, creating a visible record of persuasive effectiveness.\n\nResearch on top delta-earners found they: provided external evidence, used morality-based reasoning, wrote longer more substantive responses, engaged in back-and-forth argumentation, and showed semantic overlap with the original poster's framing. These are exactly the epistemic virtues Deliberus wants to reward. ([Characteristics of persuasive deltaboard members — eScholarship](https://escholarship.org/uc/item/6h12t1vd))\n\nThe limitation: the delta system only measures whether *the original poster* was convinced, not whether the argument was *correct*. A skilled rhetorician could earn many deltas while being systematically wrong. Deliberus would want to layer calibration scoring (are the things you convinced people of actually true?) on top of persuasion scoring.\n\n### Intellectual Humility: The Hardest Virtue to Gamify\n\nRecent research on intellectual humility in online discourse ([ACL 2024, \"Modeling Intellectual Humility\"](https://aclanthology.org/2024.emnlp-main.327.pdf)) found that markers include:\n- Explicit acknowledgment of uncertainty (\"I'm not sure, but...\")\n- Hedging with evidence quality (\"The evidence suggests, though it's not definitive...\")\n- Responsiveness to criticism (\"That's a good point I hadn't considered\")\n- Willingness to revise (\"Given what you've said, I'd update my view on X\")\n\nLLMs can now detect these markers reliably in text, which means Deliberus can algorithmically identify intellectually humble contributions without relying solely on peer review. This opens the door to automatically surfacing contributions that model good epistemic practice — and rewarding them.\n\n---\n\n## 3. Anti-Patterns: What Not to Gamify\n\n### The Stack Overflow Warning\n\nStack Overflow provides the clearest case study in gamification going wrong. Research (Mazloomzadeh et al., 2021, \"Reputation Gaming in Stack Overflow\") identified four systematic fraud patterns: voting rings (groups that upvote each other across multiple threads), sockpuppet accounts (one user controlling multiple accounts), reciprocal acceptance (accepting each other's answers to gain reputation to upvote), and serial upvoting (systematically upvoting all of one person's posts). ([arXiv:2111.07101](https://arxiv.org/abs/2111.07101))\n\nThe root cause: Stack Overflow's reputation system conflates *quality* (is this answer correct?) with *quantity* (how many votes did it get?), and total reputation unlocks privileges. This creates compounding returns on early high-reputation users and incentivizes social manipulation over epistemic contribution.\n\nCritics now argue that Stack Overflow's gamification \"is out of control, with votes no longer related to accuracy and usefulness, and reputation becoming an end goal in itself.\" ([Has Stack Overflow Become An Antipattern? — DEV Community](https://dev.to/codemouse92/has-stackoverflow-become-an-antipattern-3icb))\n\n**Specific anti-patterns to avoid in Deliberus:**\n\n| Anti-Pattern | Why It Fails | Deliberus Alternative |\n|---|---|---|\n| Total argument count as reputation | Rewards volume over quality | Score only the top-K arguments per user per domain |\n| \"Winning\" a debate | Incentivizes performance, not truth-seeking | No debates; arguments either survive scrutiny or don't |\n| Follower/popularity counts | Creates status hierarchies that distort discourse | Domain-specific calibration scores, no global celebrity |\n| Engagement time on platform | Optimizes for addiction, not reasoning quality | No session length metrics in any scoring formula |\n| Like/upvote without semantic meaning | \"I agree\" ≠ \"well argued\" | Separate axes always (see Section 5) |\n| Badges for quantity milestones | \"50 arguments posted\" rewards volume | Badges only for qualitative behaviors |\n| Downvote without explanation requirement | Strategic downvoting of good opposing arguments | Downvote requires selecting a reason; contested downvotes are reviewable |\n| Global reputation leaderboard | Creates single \"debate champion\" dynamic | Domain-specific leaderboards only |\n\n---\n\n## 4. Concrete Gamification Mechanisms\n\n### The Badge Taxonomy\n\nBadges work when they reward *specific verifiable behaviors* rather than aggregate metrics. Each badge below rewards one epistemic virtue with a clear operationalization.\n\n**Category: Intellectual Honesty**\n- **Delta** — Publicly revised your stated position and explained the specific update (adapted from r/ChangeMyView). Requires: (1) previous stated position on record, (2) explicit update with citation to the argument that changed your mind, (3) brief explanation of what specifically changed.\n- **Calibrated** — 100+ resolved claims with calibration curve within ±5% of the diagonal. Rare, prestigious, domain-specific.\n- **Acknowledged** — Responded to a well-rated counter-argument with \"this addresses my point\" and awarded a quality rating. Monthly tracking.\n\n**Category: Evidence and Precision**\n- **Evidence Hunter** — Attached a primary source (peer-reviewed paper, official dataset, verified news source) to a claim that subsequently received high verifiability ratings from domain-credible reviewers.\n- **Citation Network** — Your cited evidence was independently linked to by 5+ other arguments (evidence node became a shared resource).\n- **Falsifiable** — Created a claim with explicit resolution criteria that was subsequently evaluated — either confirmed or refuted — by community evidence review.\n\n**Category: Constructive Engagement**\n- **Steelwoman/Steelman** — Strengthened an opposing argument in a way that was rated by that argument's supporters as \"fairly representing our position.\" Requires cross-group validation.\n- **Bridge Builder** — An argument you wrote received quality ratings from users who hold the opposing position on the root claim. Measured via position-clustering.\n- **Surgeon** — Identified and named a specific logical flaw (ad hominem, straw man, false dichotomy, appeal to authority) in an argument, with the identification subsequently confirmed by community review. Not a \"you're wrong\" badge — a diagnostic badge.\n\n**Category: Community Contribution**\n- **Decomposer** — Broke a vague or compound claim into atomic sub-claims that were subsequently adopted as the canonical decomposition. LLMs can assist but human review confirms.\n- **Linker** — Identified that two arguments in different threads were logically equivalent or directly contradictory, enabling the platform to connect them.\n\n**What the badge system does NOT include:** \"X arguments posted,\" \"Y hours active,\" \"Z votes received.\" These all reward quantity or popularity.\n\n### Calibration Dashboard\n\nModeled on Metaculus's calibration curve but adapted for claims rather than predictions. Every user has a personal dashboard showing:\n\n- **Calibration curve**: For all resolved claims where you stated a confidence level, how did your stated probabilities correlate with actual outcomes? (X-axis: stated confidence 0-100%, Y-axis: actual resolution rate). A perfectly calibrated user has a diagonal line.\n- **Domain breakdown**: Calibration score by topic area (AI, climate, economics, ethics, etc.). This prevents the \"I'm good at science but pretend to know economics\" problem.\n- **Update history**: All instances where you publicly revised a stated position, with the argument that prompted the revision. This history is a badge in itself — visible intellectual growth.\n- **Influence score**: How many other arguments cited yours as evidence or adopted your decomposition? This measures intellectual contribution without measuring popularity.\n\n### Argument Survival Score\n\nUnlike debate platforms that track \"wins,\" Deliberus can track *argument survival*: how many substantive challenges has an argument withstood without being successfully rebutted?\n\nAn argument receives:\n- +points when a counter-argument is rated \"fails to address the original\" by neutral reviewers\n- +points when new supporting evidence is added that increases the claim's community confidence estimate\n- +points when an opposing user awards it a quality rating (signal: \"I disagree but this is well-argued\")\n- -points when a counter-argument is rated \"successfully addresses a core premise\" by neutral reviewers\n- -points when supporting evidence is downgraded after closer examination\n\nArguments that survive many challenges become \"battle-tested\" — a visible marker on the claim indicating high scrutiny, not mere popularity.\n\n### Streaks: Depth, Not Activity\n\nDuolingo's streak system produces daily engagement but risks rewarding trivial activity. ([Duolingo Case Study 2025 — Young Urban Project](https://www.youngurbanproject.com/duolingo-case-study/)) For Deliberus, streaks should require *depth* not just presence:\n\n- **Evidence streak**: Attached a quality-reviewed source every day for N days. Does not reset if a day is missed but recent sources are lower quality.\n- **Review streak**: Rated at least one counter-argument substantively (selected \"addresses core premise\" or \"fails to address\" with explanation) every day for N days.\n- **Update streak** (rare): Publicly revised a stated position in response to new evidence — not a daily streak but a lifetime counter, displayed prominently.\n\nThe principle: streaks reward *the right things* when the qualifying action is specific and qualitative, not just \"log in and click something.\"\n\n---\n\n## 5. The \"Separate Agree from Well-Argued\" Innovation\n\nThis is the most fundamental design decision in the entire gamification architecture, and it has now been tested in the wild.\n\n### LessWrong's Implementation\n\nLessWrong activated two-axis voting on June 24, 2022. The standard karma axis (left) measures \"should this be visible/promoted?\" The new agree/disagree axis (right) measures \"do you think this is true?\" Both axes support strong-vote (click-and-hold). ([LessWrong Has Agree/Disagree Voting — LessWrong](https://www.lesswrong.com/posts/HALKHS4pMbfghxsjD/lesswrong-has-agree-disagree-voting-on-all-new-comment))\n\nThe EA Forum adopted the same system in September 2022. ([Agree/disagree voting — EA Forum](https://forum.effectivealtruism.org/posts/2JQQZevGENbChSA8k/agree-disagree-voting-and-other-new-features-september-2022))\n\n**Results and reception.** The system partially achieved its goal: karma became deconfounded from agreement. \"If a comment has high disagreement and high karma, the karma has been deconfounded — it seems much more likely that people have updated on it or otherwise thought the arguments have gone underappreciated.\" But several problems emerged:\n\n- **Cognitive overhead**: Two-click voting asks users to make a more complex judgment per comment. Many users skip one axis or the other, producing noisy data.\n- **Meaning ambiguity**: What does \"agree with this comment\" mean when the comment is a question, a personal anecdote, or a meta-comment about the thread? The \"agree\" axis is most meaningful for direct propositional claims.\n- **Visual confusion**: Two vote indicators per comment creates UI clutter, especially on mobile.\n\n**Implications for Deliberus:** The two-axis approach is directionally correct but needs cleaner design. In Deliberus, the problem is simpler than LessWrong because content is *structured as claims* — not free-form comments. \"Do you agree with this claim?\" and \"Is this argument well-reasoned?\" are both meaningful questions when applied to an explicit proposition. The UI should make this mandatory (not optional) for substantive engagement, not an afterthought on every piece of content.\n\n### The Three-Axis Framework\n\nFor Deliberus, three axes are more appropriate than two:\n\n1. **True/False confidence** (0-100%): \"How confident are you that this claim is true?\" This is the prediction-market input — feeds the calibration scoring system. Not agree/disagree (which is about values), but factual confidence.\n\n2. **Well-argued / Poorly-argued** (5-point scale): \"How well does this argument support its conclusion? Is the evidence quality high? Is the reasoning valid?\" Explicitly separated from agreement.\n\n3. **Important / Unimportant** (binary): \"Is this claim relevant and significant to the parent argument?\" Prevents irrelevant claims from cluttering the argument graph.\n\n**When a \"disagree but well-argued\" argument surfaces, this is the desired outcome.** It should be visually prominent — highlighted as a \"high-quality challenge\" — not buried because the majority disagrees. The whole point of the architecture is that this situation produces learning. The algorithm should actively surface arguments that score high on axis 2 while diverging from the crowd on axis 1.\n\n### Displaying Both Dimensions\n\nVisual design options:\n- **Color encoding**: Argument nodes colored by confidence estimate (warm = high confidence, cool = low confidence, gray = contested), bordered by argument quality (thick border = well-argued, thin = poorly rated).\n- **Quadrant display**: In argument clusters, show a 2D scatter of claims positioned by (community confidence estimate, argument quality rating). The high-quality / low-confidence quadrant is the \"worth investigating\" zone.\n- **Labels**: \"Contested: 48% confidence, high argument quality\" vs \"Uncontested: 91% confidence\" — explicit textual indicators rather than relying solely on color.\n\n---\n\n## 6. Preventing Gaming\n\n### Sybil Attacks and Fake Accounts\n\nThe graph structure of argumentation provides natural Sybil defense that pure voting platforms lack. Social-graph-based Sybil detection assumes that attackers cannot establish many connections to legitimate users. ([Sybil attack — Wikipedia](https://en.wikipedia.org/wiki/Sybil_attack)) In Deliberus, \"connections\" include: cited as evidence by credible users, responded to substantively by domain-credible contributors, awards received from established accounts.\n\nA new account can post claims but its contributions start at low-weight visibility. Weight increases through verified epistemic behaviors: resolved predictions that were calibrated, evidence citations that were independently corroborated, deltas received from established accounts. This is not a reputation gate (the claim is visible) but a *signal weight* gate (the claim starts in a lower-visibility tier).\n\n### Strategic Downvoting\n\nThe most dangerous gaming vector on argumentation platforms: users who systematically downrate well-reasoned opposing arguments to suppress visibility. Mitigations:\n\n1. **Reason requirement**: All low-quality ratings require selecting a category from a fixed list (straw man, missing evidence, non-sequitur, ad hominem, irrelevant to parent claim). This makes strategic downvoting slower and traceable.\n\n2. **Rating review**: Any argument with a large gap between quality-ratings-from-agreers and quality-ratings-from-disagreers is automatically flagged for neutral-party review. The pattern \"X has 80% agree among users who agree with the parent claim, 15% among users who disagree\" is a shill-voting signal.\n\n3. **Weighted ratings**: Quality ratings are weighted by the rater's calibration score in the relevant domain. Someone with a track record of poor epistemic behavior (calibration consistently off, deltas never received, own arguments consistently rated poorly) has lower weight in the quality-rating system.\n\n### Voting Rings\n\nStack Overflow's research on voting rings ([arXiv:2111.07101](https://arxiv.org/abs/2111.07101)) identified them through behavioral anomalies: serial upvoting across many threads, timing correlation between accounts, and acceptance patterns. Deliberus has additional detection signal: if two users consistently give each other high quality ratings without engaging substantively with each other's arguments (no counter-arguments, no evidence citations, no deltas in either direction), this is a coordination signal.\n\n### Quadratic Voting for Quality Assessment\n\nQuadratic voting (QV) — where the marginal cost of additional votes on the same item increases as the square of total votes cast — has theoretical advantages for expressing intensity of preference. However, QV is vulnerable to Sybil attacks: an attacker with N fake accounts can cast N single-votes for QV cost N, better than one account spending N² for the same effect. ([arXiv:2407.01844](https://arxiv.org/abs/2407.01844))\n\nFor Deliberus, the simpler weighted-voting approach (weight by calibration score and epistemic history) is more robust than QV, because it makes the \"fake accounts\" problem harder: fake accounts need to accumulate real epistemic track records to acquire weight, not just demonstrate identity.\n\n---\n\n## 7. Onboarding Gamification\n\n### The Core Principle: Value Before Identity\n\nDuolingo's major onboarding improvement came from reversing the sign-up gate: push the first lesson *before* account creation. Users who experience value first convert at dramatically higher rates than users who complete a registration form first. ([Duolingo's delightful user onboarding — Appcues](https://goodux.appcues.com/blog/duolingo-user-onboarding)) The lesson: \"commitment before identity\" works when the initial experience is genuinely engaging.\n\nFor Deliberus, the anonymous first interaction should show what the platform does and why it is different from a standard argument forum. This means the onboarding must demonstrate the *unique mechanic* — not just \"here is an argument tree\" but \"here is how we separate 'I agree' from 'this is well-argued,' and here's what happens when you try it.\"\n\n### The Five-Minute Hook\n\nA proposed first-five-minutes flow that establishes the core epistemic loop:\n\n1. **Present a single contested claim** (pre-seeded, chosen for the topic area the user selected during onboarding — climate, AI, politics, etc.). Not a trivial claim — something the user genuinely has a view on.\n\n2. **Ask: How confident are you that this is true? (0-100%)** — no explanation yet. This is the user's prior.\n\n3. **Show the strongest counter-argument** — not a strawman, but a steelmanned version. Attribute it to \"a well-calibrated contributor in this domain.\"\n\n4. **Ask: Has your confidence changed? (0-100%)** — explicitly framing belief revision as natural and expected.\n\n5. **Reveal the community's confidence estimate and calibration history for this claim.** \"The community has rated this claim at 61% confidence after 142 challenges. Here is the evidence graph.\"\n\n6. **Ask the user to rate the argument quality of the counter-argument** (well-argued / poorly-argued). Explain: \"This is separate from whether you agree with the conclusion.\"\n\nThis loop takes 2-3 minutes, demonstrates the core value proposition (calibrated belief revision with quality-separated evaluation), and establishes the behavior pattern before the user has posted anything.\n\n### First Contribution: Lowest-Friction Meaningful Action\n\nThe lowest-friction contribution that is still semantically meaningful on Deliberus is rating the quality of an existing argument — not creating a new one, not posting evidence, but simply applying the \"well-argued / poorly-argued\" rating to a claim presented during the onboarding flow. This requires judgment (not just clicking), is genuinely useful to the platform, and gives the user an immediate sense of participation.\n\nSequence:\n- **First action**: Rate an argument quality (2 clicks, takes 30 seconds)\n- **First feedback**: \"Your rating has been included in this claim's quality score. Here is how it changed.\" (immediate, visible impact)\n- **Second action**: Add a confidence estimate to a factual claim (one number, 10 seconds)\n- **Second feedback**: \"Your calibration dashboard has been started. You'll see how your estimates compare to outcomes over time.\"\n\nThis mirrors Duolingo's \"immediate success signal\" pattern: users who see their first contribution matter are more likely to return. ([Duolingo's Gamification Secrets — Orizon](https://www.orizon.co/blog/duolingos-gamification-secrets))\n\n### The Loss Aversion Hook\n\nDuolingo research shows streaks leverage loss aversion more effectively than reward accumulation. The 'Streak Freeze' feature reduced churn by 21%. For Deliberus, the equivalent is the *calibration streak*: your personal calibration curve, once visible, is something you want to maintain. The feedback is not \"you'll earn X points tomorrow\" but \"you have a 30-day calibration record — will you make it 31?\"\n\n---\n\n## 8. Proposed Reputation System\n\nSynthesizing the above into a coherent architecture. This is a design proposal, not a decision.\n\n### Core Principle: Multi-Dimensional, Non-Aggregated\n\nA single reputation number is corrupted by the first manipulation attempt. Deliberus should maintain at least four independent reputation signals that are displayed together but never combined into a single score:\n\n**Signal 1: Calibration Score (domain-specific)**\n- Measured using log scoring rule across all resolved claims where user stated a confidence level\n- Displayed as calibration curve (visual) plus percentile-in-domain (numeric)\n- Domain-specific: separate scores for AI, climate, economics, ethics, etc.\n- Only updates on claim resolution — cannot be gamed by activity volume\n\n**Signal 2: Argument Quality Record**\n- How often are your arguments rated \"well-argued\" by users who *disagree* with your conclusion? (Cross-adversarial quality rating — hardest to game)\n- How often do your arguments survive substantive challenges without being rebutted?\n- Normalized per-argument, not aggregate (no incentive to post more low-quality arguments)\n\n**Signal 3: Intellectual Honesty Record**\n- Count of public position revisions (deltas), with the arguments that prompted them\n- Count of quality ratings awarded to opposing arguments\n- This is a *narrative record*, not just a number — it shows your intellectual trajectory\n\n**Signal 4: Evidence Contribution**\n- Number of cited sources that were independently verified and subsequently used by other arguments\n- This rewards the research function of the platform — finding evidence, not just stating opinions\n\n### Display\n\nRather than a single number or a leaderboard, user profiles show:\n- A small calibration curve thumbnail (visual immediately communicates epistemic reliability)\n- Domain badges (earned through threshold calibration in specific areas)\n- Intellectual honesty record (delta count visible, with the option to view the revision history)\n- Argument quality percentile in the user's primary domain\n\nThe system does NOT show:\n- Total arguments posted\n- Total karma/points\n- Follower count\n- Global leaderboard rank\n\n### Privilege Unlocks\n\nRather than privileges unlocking via total reputation (Stack Overflow's approach), privileges in Deliberus unlock via *demonstrated behavior in the relevant category*:\n\n- Ability to create new root claims → requires at least 5 quality-rated arguments and a calibration baseline\n- Ability to rate claim verifiability → requires calibration score above 50th percentile in that domain\n- Moderator privileges for a topic area → requires argument quality record above 70th percentile in that area + intellectual honesty record of at least 3 public revisions\n- \"Verified contributor\" status → requires calibration score that persists above 60th percentile over 6 months\n\nThis prevents new accounts from immediately shaping the platform while keeping the barrier behaviorally specific rather than just time-based.\n\n---\n\n## Key Open Questions\n\n1. **How do you calibrate confidence for normative claims?** Factual claims resolve against observable outcomes. Value claims don't. Can the \"agree/disagree\" axis serve as a proxy for normative confidence, and can it be made meaningful without a resolution criterion?\n\n2. **How long is the calibration feedback loop?** Many important claims take years or decades to resolve. Users need shorter-feedback-loop mechanisms to develop epistemic skills — monthly \"claim resolution events\" or calibration tournaments on near-term questions could provide this.\n\n3. **Can LLMs reliably rate steelmanning quality?** The steelman badge requires cross-group validation (\"the opposing side agrees this fairly represents their position\"). LLMs could assist with this — but whose values does the LLM use to evaluate \"fair representation\"? This is a normative question embedded in what appears to be a technical one.\n\n4. **What is the minimum viable calibration record?** A calibration curve is meaningless with 5 data points. Users need to make ~50+ predictions before the curve is interpretable. How do you keep users engaged through the first 50, before their calibration data has epistemic meaning?\n\n5. **Does the multi-dimensional reputation system introduce too much cognitive complexity?** Four separate signals require users to understand what each means. The onboarding must clearly explain why there is no single score — that the absence of a single number is a feature, not a bug.\n\n---\n\n## Sources\n\n### Prediction Markets and Calibration\n- [A Primer on the Metaculus Scoring Rule — Metaculus](https://www.metaculus.com/notebooks/22486/a-primer-on-the-metaculus-scoring-rule/)\n- [Metaculus Help: Scores FAQ](https://www.metaculus.com/help/scores-faq/)\n- [Metaculus Introduces New Forecast Scores, New Leaderboard — EA Forum](https://forum.effectivealtruism.org/posts/FodvZaiKftDCHPTub/metaculus-introduces-new-forecast-scores-new-leaderboard-and)\n- [Aligning Incentives for Forecast Accuracy — Metaculus/Medium](https://metaculus.medium.com/aligning-incentives-for-forecast-accuracy-relevance-and-efficacy-a-new-paradigm-for-metaculus-26b0e79616cb)\n- [Manifold Markets FAQ](https://docs.manifold.markets/faq)\n- [Manifold Leagues](https://manifold.markets/leagues)\n- [Manifold markets isn't very good — EA Forum](https://forum.effectivealtruism.org/posts/EaR9xFxspmYRkm3eo/manifold-markets-isn-t-very-good)\n- [The Good Judgment Project — Wikipedia](https://en.wikipedia.org/wiki/The_Good_Judgment_Project)\n- [Good Judgment Open FAQ](https://www.gjopen.com/faq)\n- [Evidence on good forecasting practices — AI Impacts](https://aiimpacts.org/evidence-on-good-forecasting-practices-from-the-good-judgment-project/)\n- [Evidence on good forecasting practices — EA Forum](https://forum.effectivealtruism.org/posts/W94KjunX3hXAtZvXJ/evidence-on-good-forecasting-practices-from-the-good-1)\n\n### Epistemic Virtue Philosophy\n- [Virtue Epistemology — Stanford Encyclopedia of Philosophy](https://plato.stanford.edu/entries/epistemology-virtue/)\n- [Nurturing Virtues with Digital Democratic Innovations — Springer](https://link.springer.com/article/10.1007/s13347-025-00906-4)\n- [Modeling Intellectual Humility in Online Public Discourse — ACL 2024](https://aclanthology.org/2024.emnlp-main.327.pdf)\n- [The curious joy of being wrong: intellectual humility — The Conversation](https://theconversation.com/the-curious-joy-of-being-wrong-intellectual-humility-means-being-open-to-new-information-and-willing-to-change-your-mind-216126)\n\n### Two-Axis Voting Research\n- [LessWrong Has Agree/Disagree Voting On All New Comment Threads](https://www.lesswrong.com/posts/HALKHS4pMbfghxsjD/lesswrong-has-agree-disagree-voting-on-all-new-comment)\n- [Open Thread - Jan 2022 Vote Experiment — LessWrong](https://www.lesswrong.com/posts/ywpWMnJmqAkeaDtne/open-thread-jan-2022-vote-experiment)\n- [Agree/disagree voting — EA Forum September 2022](https://forum.effectivealtruism.org/posts/2JQQZevGENbChSA8k/agree-disagree-voting-and-other-new-features-september-2022)\n- [Two-factor voting for EA forum — EA Forum](https://forum.effectivealtruism.org/posts/e7rWnAFGjWyPeQvwT/two-factor-voting-two-dimensional-karma-agreement-for-ea)\n\n### r/ChangeMyView and Belief Revision\n- [r/changemyview — Wikipedia](https://en.wikipedia.org/wiki/R/changemyview)\n- [Characteristics of persuasive deltaboard members — eScholarship](https://escholarship.org/uc/item/6h12t1vd)\n- [Reddit's Change My View: a template for online discussion — The Next Web](https://thenextweb.com/socialmedia/2019/01/23/reddits-model-community-offers-a-prototype-for-controversial-discussions)\n\n### Stack Overflow Gamification and Anti-Patterns\n- [The Gamification — Coding Horror (Jeff Atwood)](https://blog.codinghorror.com/the-gamification/)\n- [A Dusting of Gamification — Joel on Software](https://www.joelonsoftware.com/2018/04/13/gamification/)\n- [Reputation Gaming in Stack Overflow — arXiv:2111.07101](https://arxiv.org/abs/2111.07101)\n- [Has Stack Overflow Become An Antipattern? — DEV Community](https://dev.to/codemouse92/has-stackoverflow-become-an-antipattern-3icb)\n\n### Sybil Attacks and Gaming Prevention\n- [Sybil attack — Wikipedia](https://en.wikipedia.org/wiki/Sybil_attack)\n- [An Efficient and Sybil Attack Resistant Voting Mechanism — arXiv:2407.01844](https://arxiv.org/abs/2407.01844)\n- [Quadratic Voting: How Mechanism Design Can Radicalize Democracy — AEA](https://ideas.repec.org/a/aea/apandp/v108y2018p33-37.html)\n- [Going Parabolic: Analyzing Sybil Resistance in Quadratic Voting — Stanford](https://purl.stanford.edu/hj860vc2584)\n\n### Onboarding Psychology\n- [Duolingo's delightful user onboarding — Appcues](https://goodux.appcues.com/blog/duolingo-user-onboarding)\n- [Duolingo Case Study 2025: How Gamification Made Learning Addictive](https://www.youngurbanproject.com/duolingo-case-study/)\n- [Duolingo's Gamification Secrets: Streaks & XP Boost Engagement by 60% — Orizon](https://www.orizon.co/blog/duolingos-gamification-secrets)\n- [How Gamification of Digital Tools Amplifies User Engagement](https://www.macrocreator.com/2025/05/22/how-gamification-of-digital-tools-amplifies-user-engagement/)\n"}