{"path":"research/saturation-and-the-long-tail.md","content":"# Saturation and the Long Tail\n\n**Date**: 2026-08-30 · **Type**: online research sweep, founder-commissioned — deepening the claim\n*\"distinct content saturates while text does not\"* from many directions\n([drowning-in-claims.md](drowning-in-claims.md) § 3). Eight directions swept; two genuine\nsharpenings of the original claim, one new computable instrument, and one deep caveat.\n\n---\n\n## 0. The claim, deepened in one paragraph\n\nThe original: distinct content saturates (the 72.5% ceiling) while text does not, so drowning\nconverts into a sameness problem. The research sharpens this into a **two-regime curve**: the\n**head** — the mass of commonly-made points — saturates fast and hard, exactly as the ceiling\nmeasured; but the **tail** of rare, distinct contributions decelerates *without ever closing*, and\n— the most important finding of the sweep — **the tail is where disproportionate value lives, and\nit is precisely what machine summarization destroys**. A claim graph is a tail-preserving\nstructure; summarization is tail-killing by construction. So the saturation story and the\nanti-flattening story turn out to be one story: hold the head compactly (sameness) and the tail\nlosslessly (addressability).\n\n## 1. Two mathematical families, and which one arguments follow\n\nTwo literatures model \"how many distinct types as observations grow,\" with different endings:\n\n- **Bounded richness** (ecology): species-accumulation curves flatten toward an asymptote, and the\n  Chao estimators bound the unseen remainder *from the counts of rare items alone* — \"all unseen\n  species are rare; estimators for unseen species depend only on rare-species data\" ([Chao,\n  species-richness estimation](https://www.uvm.edu/~ngotelli/manuscriptpdfs/ChaoEncyclopediaChapter.PDF);\n  [general framework](https://arxiv.org/pdf/2011.07270)).\n- **Unbounded-but-decelerating** (linguistics): Heaps' law — vocabulary grows as a sublinear power\n  of text length, V ≈ K·Lᵝ with β < 1, decelerating forever without an asymptote; hapaxes keep\n  arriving in gigaword corpora ([entropy and type-token ratio in gigaword corpora](https://arxiv.org/html/2411.10227v1)).\n\nThe argument evidence fits **both at once, by region**: key-point coverage saturates hard (the\n72.5% curve — [Bar-Haim et al.](https://aclanthology.org/2020.acl-main.371/)), which is\nChao-shaped; but the residual keeps producing genuinely new material, which is Heaps-shaped — in\nseveral argument datasets **more than half the noun phrases and entities are unique to\nlow-frequency arguments** ([diversity in argument summarization](https://arxiv.org/html/2402.01535v1)).\nSo: the head closes, the tail never quite does. Drowning-by-restatement is a head phenomenon;\ndiscovery is a tail phenomenon.\n\n## 2. The qualitative-research calibration: two saturations, not one\n\nInterview research measured exactly the layered saturation the nodes-to-edges shift predicts.\n**Code saturation** (the *list* of themes) arrives at ~9–12 interviews; **meaning saturation**\n(understanding the themes) needs 16–24 or more — *\"code saturation indicates we have heard it all;\nmeaning saturation indicates we understand it all\"*\n([Hennink, Kaiser & Marconi 2017](https://journals.sagepub.com/doi/10.1177/1049732316665344);\n[Guest, Bunce & Johnson 2006](https://skimle.com/blog/how-many-interviews-qualitative-research);\ncross-site replication needed 20–40 — Hagaman & Wutich 2017). Translated: **claim-level saturation\ncomes early; edge-, evidence- and understanding-level saturation comes much later** — independent\nsupport, with numbers, for mature growth shifting from new nodes to new edges and answers.\n\n## 3. The tail is where the value is — and machines kill it, measured twice\n\n- Key-point systems **systematically underperform on low-frequency key points** — coverage of\n  minority perspectives degrades across all tested approaches, while the long tail holds the\n  unique content ([diversity paper](https://arxiv.org/html/2402.01535v1)).\n- **\"Argument collapse\"**: LLMs summarizing long-form public debate flatten the spectrum toward\n  homogeneous middle-ground positions; tail arguments — minority, unconventional, dissenting —\n  are underrepresented or vanish entirely ([arXiv:2606.01736](https://arxiv.org/pdf/2606.01736)).\n\nThis joins the corpus's own run-3F finding (the crux of a real debate is typically *implicit* — a\ntail item by construction) and the DeepMind steering result into one picture: **frequency and value\nare anti-correlated exactly in the region machine compression drops.** The differentiation claim\nthat follows, sharper than before: a claim graph is the tail-preserving alternative — a rare claim\npersists as a first-class addressable node (and QEM's additivity protects pebbles from burial),\nwhere every summarizer, human or machine, would have dropped it. Anti-drowning and\nanti-flattening are the same design.\n\n## 3b. The founder's sharpening: the hazard is INTERNAL, at the main door\n\nFounder, on reading § 3 (2026-08-30): *\"this feels like a super important finding, also considering\nthat most people (or outside AI agents using a Deliberus MCP or some such) will likely mainly\ninteract with Deliberus through a truth-graph-summary or whatever.\"*\n\nThat closes the loop the sweep left open: argument collapse is not only what OTHER systems do to\ndebate — **every summary-shaped surface Deliberus itself serves is a tail-killing layer by the same\nmechanism**, and those surfaces (the synthesis endpoint, `Think with Deliberus`, the truth-graph\nanswer layer, the future read-only MCP) are exactly where most humans and nearly all machine\nreaders will meet the graph. The graph preserves the tail; **the door most visitors use is a\nsummarizer standing in front of it.** Consequences:\n\n1. **The inspectable-synthesis machinery upgrades from honesty feature to load-bearing\n   differentiator** — the citation gate and omissions ledger are what keep the door from silently\n   becoming the very argument-collapse layer the graph exists to beat.\n2. **A new dial: head/tail citation balance.** A synthesis can pass the citation gate citing only\n   head claims while the tail vanishes — the measured failure of every key-point system. The\n   omissions ledger should report what fraction of *rare* (low-match-count, minority-attributed)\n   claims in the retrieved subgraph reached the answer, alongside the existing conflict coverage.\n   Unruled, cheap.\n3. **For the MCP, tail access is the product**: an agent that wants the consensus middle can ask\n   any LLM; what only Deliberus holds is the addressable tail — the dissents, the implicit cruxes,\n   the minority renderings with their provenance. The agent surface should make tail retrieval a\n   first-class operation, not a side effect of summary.\n\n## 4. What mature commons actually experienced\n\n- **Wikipedia's growth is logistic, not exponential** ([Suh et al. 2009, \"The singularity is not\n  near\"](https://research.google/pubs/the-singularity-is-not-near-slowing-growth-of-wikipedia/)):\n  slowdown driven by *limited opportunities for novel contribution*, rising coordination overhead,\n  and **resistance to newcomers** — whose edits increasingly duplicate or conflict and get\n  reverted. **The design warning for Deliberus**: on Wikipedia, arriving after saturation feels\n  like rejection. In a claim graph it need not — a duplicate contribution should land as\n  *agreement with attribution* (your stance recorded, your voice added to a standing claim), which\n  the add-never-overwrite merge design already implies. Same saturation, opposite newcomer\n  experience — a UX decision, not a law of maturity.\n- **Stack Overflow**: duplicate growth is a named quality-decay driver in mature Q&A, and duplicate\n  *detection* recall decays as the corpus grows\n  ([reproducibility study](https://ieeexplore.ieee.org/document/8330262/)) — **sameness gets harder\n  exactly when it matters most**, which prices the claim-sameness work honestly: permanent, and\n  scaling adversely.\n\n## 5. The recombination era: external support for nodes → edges\n\nAt the research frontier, **new distinct ideas cost ever more** — research productivity falling\n~5–7%/year across fields ([Bloom et al., *Are Ideas Getting Harder to\nFind?*](https://web.stanford.edu/~chadj/IdeaPF.pdf); methodological critiques exist —\n[Guzey](https://guzey.com/economics/bloom/) — so hold the magnitude loosely). Meanwhile the\nhighest-impact work runs on **atypical combinations of existing knowledge**\n([Uzzi et al. 2013](https://www.kellogg.northwestern.edu/faculty/uzzi/htm/papers/science-2013-uzzi-468-72.pdf)):\nnovelty increasingly lives in *recombination* — in edges between existing nodes, not new nodes.\nThe graph is the right substrate for the era where connection is the growth frontier; the analogy\ndaemon is its automation. And Polis in production shows opinion-space dimensionality is tiny\nrelative to text volume: thousands of participants collapse to ~5 opinion groups and a dozen-odd\nconsensus statements ([vTaiwan field insights](https://arxiv.org/html/2502.05017v1)) — the small\nbasis again.\n\n## 6b. The second dial, and why the completion verdict is unavailable from inside (2026-09-10)\n\n**Founder, verbatim**, on whether engagement eventually closes the framing gaps: *\"I actually expect use of the system to keep contributing more and more considerations that broaden or clash with existing framings, as humans (and agents) engage more and more with it... And maybe there's an asymptotic closing of all conceivably relevant gaps. But I'm not sure how to ultimately make that judgment, I suspect it's impossible...\"*\n\n**Half of it the corpus already supports, and the halves differ by regime.** The head of a topic saturates\n(Chao-shaped, bounded, estimable from rare-item counts); the tail decelerates without closing (Heaps-shaped).\nA *framing* is a tail item by construction — each new position is rare, and a position nobody in the corpus\noccupies has a count of zero. So the curve's **shape** is as he describes, decelerating, and its **limit** is not\nzero: expect ever-slower arrival of genuinely new framings, never a last one.\n\n**Why the judgment is unavailable, which is his own suspicion given a reason.** A verdict that *all conceivably\nrelevant gaps are closed* is a claim about what lies outside the frame the graph was built in, rendered from\ninside it. That is the frame gap applied to the graph itself, and the corpus's own conviction says the inside\nview cannot make it (`what-human-judgment-is-for.md`; the terminus is a *fallible fixed point under the current\nmove-set*, never proven exhaustion — the same honesty one level down). So the honest object is not a completion\nverdict but a **rate with a provenance**.\n\n**The second dial, joined here for the first time.** § 6 pairs the novelty rate with the `does_not_fit` rate,\nwhich tells apart *nothing new is arriving* from *the taxonomy is forcing fits*. It cannot tell apart either\nfrom the third case: **the population has run out of frames, not the subject of material.** N users is not N\nframes — a thousand readers from one reading community are one frame sampled a thousand times\n([what-human-judgment-is-for.md](what-human-judgment-is-for.md)) — so a falling novelty rate over a narrow,\nself-selected contributor base is not saturation at all. Read three together:\n\n| Dial | Falling / low means | Cheap to compute? |\n|---|---|---|\n| Good-Turing novelty rate (§ 6) | new distinct content is arriving more slowly | yes, from mint-time dedup |\n| `does_not_fit` rate | the taxonomy is (not) forcing borderline material into wrong slots | **not yet — the channel is inert, see below** |\n| **contributor-frame diversity** (new) | the population producing that novelty is (not) still widening | needs a frame proxy: register, source cluster, declared lens, or arrival channel |\n\nSaturation is only claimable when novelty falls **while the third dial is still rising**. Novelty falling as the\nthird dial flattens is the ossification case wearing saturation's clothes, and it is the likelier one, because\narrival self-selects. Consequence for the reframe invitation ruled 2026-09-10: its response rate is itself a\nreading of the first dial at the frame level, and the positions of the people responding are a reading of the\nthird.\n\n**Which taxonomy, exactly** (founder asked 2026-09-11). Two closed sets in this system carry a `does_not_fit` confession value, and only these two: the **argument-scheme set** (`deliberus/extraction/schemes.py`, 46 labels, one of them `does_not_fit`, whose share the code's own comment calls *\"the taxonomy meeting a register it was not induced from\"*) and the **terminus/horizon type set** (`deliberus/terminus.py`, where `does_not_fit` additionally *requires* a free-text note naming the missing slot — the stronger design of the two). Claim type and epistemic status are closed enums with **no** confession value, so they cannot report a missing slot at all.\n\n**Measured live 2026-09-11** (`MATCH ()-[r]->() WHERE r.scheme = 'does_not_fit'`): **0 of 3,666 scheme-bearing edges**, and 0 of the 3 typed termini. By the project's own rule — *a guard that never fires looks exactly like one that works; a zero firing count on the live corpus is an alarm, never a pass* — that is an alarm, and the cause is the prompt rather than the world. The value reaches the model as a bare bullet under a `DEFAULT:` heading with **no instruction saying what it means or when to use it**, beside a second undifferentiated fallback (`default_inference`, 15 uses). Meanwhile the phenomenon it should catch is arriving unlabelled: **113 edges were stored with an empty scheme string**, which is silent — no note required, indistinguishable from a field that was never populated. ⚠ **Corrected 2026-09-11, same day: 113 is the wrong numerator.** Split by edge type, 39 of those are DEFINES or DECOMPOSES_INTO, which carry no inference and are *correctly* unlabelled; the genuine no-fits are the **74 argumentative edges** (SUPPORTS 39, ATTACKS 20, QUALIFIES 13, REFRAMES 2). Against 3,580 argumentative edges that is **2.1%** of edges arriving unlabelled. ⚠ **Corrected 2026-09-12: that is NOT a no-fit rate.** Re-running those same claims through the repaired prompt classified most of them cleanly, so a blank mostly meant the earlier model declining to choose, not the taxonomy missing a pattern. The true no-fit rate is unknown until the repaired prompt runs on a fresh source. The distinction is load-bearing for the repair: defaulting structural edges to the confession value would inflate its share and destroy it. Against the field-test baseline the corpus already cites (37% of real arguments fitting none of the 14 Walton schemes; our set defines 46 labels, 41 of them used), a true zero was never plausible.\n\n**Consequence for the three-dial read**: the second dial is currently unreadable, and repairing it is cheap — give the value an instruction line in the prompt, and count the empty-string scheme as *unclassified* rather than as nothing. Until then, a quiet second dial means the instrument is off, not that the taxonomy is fitting.\n\n**Registered as a prediction, so it can be scored**: if the reframe invitation ships and its take-up decays\nwhile contributor-frame diversity is flat, the corpus is ossifying rather than completing, whatever the\ncoverage numbers say.\n\n**Founder question 2026-09-13 — does the read need to know how many people are browsing the domain?**\nThe dials count *contributions*, not readers, and the reliability conditions, stated in one place: a\ncontribution count large enough to make each rate a rate; the second dial actually firing (repaired\n2026-09-12 — `scheme-set-exhaustiveness.md` § The three runs that made it fire); and the third dial still\nrising. Reader traffic is not a dial today. His suggestion, filed unruled: a per-domain\n**reader-to-contributor ratio** could separate *people looked and had nothing to add* from *nobody looked*,\nwhich the three dials cannot tell apart.\n\n## 6. A new instrument falls out: the Good-Turing novelty rate\n\nThe ecology math converts directly into a **computable per-topic saturation gauge**. Good-Turing:\nthe probability that the next observation is a *new* type ≈ the fraction of current types seen\nexactly once. Mint-time dedup already computes near-matches, so the **singleton rate is free** —\nand with singleton and doubleton counts, a Chao-style lower bound on *unseen distinct claims* per\ntopic is one formula away. This is the sixth dial for the drowning dashboard\n([drowning-in-claims.md](drowning-in-claims.md) § 6): **the measured probability that the next\nearnest contribution adds something distinct**. High → the topic is still opening; low → route\ncontributors toward connection, evidence and answering rather than new claims. Unruled, cheap, no\nLLM.\n\n**The deep caveat that must ride with it**: a saturating curve cannot distinguish *domain mapped*\nfrom **frame ossified**. If the extraction and sameness machinery force borderline contributions\ninto existing claims (the nearest-fit trap at corpus scale), measured coverage rises while real\ndistinctness is being suppressed — apparent saturation as frame lock-in's symptom. The paired\ndetector already exists: the `does_not_fit` rate. **A topic showing high estimated coverage AND a\nnear-zero does-not-fit rate is suspicious, not finished** — the two dials must be read together.\n\n## 7. Exposure shapes what arrives\n\nCrowd-ideation research: exposing contributors to **novel** ideas measurably raises the novelty of\nwhat they produce; exposure to common ideas does nothing\n([crowdsourced idea generation](https://www.researchgate.net/publication/323954018_Crowdsourced_idea_generation_The_effect_of_exposure_to_an_original_idea)).\nFeed consequence: what the surface shows is not neutral to what gets contributed — **a feed that\nsurfaces rare-but-strong claims breeds tail contributions; a feed that surfaces the popular head\nbreeds restatement.** One more independent argument for epistemic-value ordering over\nengagement ordering, and a design input for the triage feed.\n\n## 8. Honest limits\n\nThe 72.5% ceiling and the KPA results come from short, single-shot opinion corpora, not from\nlong-lived deliberation; no study was found that tracks argument-repertoire exhaustion over years\nin one living debate (a genuine gap — Deliberus's own graph will eventually be the better\ndataset). Bloom's magnitudes are contested. The qualitative-saturation numbers are about\ninterviews, transferred here by structural analogy. And everything in § 6 is proposed, not built.\n\n---\n\nCross-references: [drowning-in-claims.md](drowning-in-claims.md) (the parent question; the dial\njoins its dashboard) · [the-reuse-flywheel-prior.md](the-reuse-flywheel-prior.md) (the asymptote;\nthe ceiling's first appearance) · [the-residual-error-taxonomy.md](the-residual-error-taxonomy.md)\n(frame lock-in — § 6's caveat is its measurement face) ·\nthe standing threat model's convergence-illusion entry (argument collapse is the outside\nliterature naming our hazard) ·\n[the-ratification-record-and-the-triage-feed.md](the-ratification-record-and-the-triage-feed.md)\n(§ 7's exposure lever) · [taxonomy-gaps-and-the-closed-enum.md](taxonomy-gaps-and-the-closed-enum.md)\n(the does-not-fit pairing)\n"}