{"path":"open-problems.md","content":"# Open Problems\n\n*One sentence of context first: **Deliberus turns debates into maps** — a text goes in, and out comes a web of claims, evidence, and objections where every piece is visible, linked, and open to challenge, with a computed score showing how well-supported each claim currently is.*\n\n*This page lists the questions we have **not** solved. Each is written so you can attack it without knowing anything else about the project — and attacking them is exactly what this page is for. If you have a take on any of these, we want it. None of them are rhetorical: every one is genuinely open, and a good answer to several would change how the system is built.*\n\n## Two running examples, used throughout\n\n**Grandma's keys.** A family is arguing about whether grandma should stop driving. The claim on the table: *\"We should take grandma's car keys.\"* Mapped, it splits into three parts: a **fact** (*she's had two near-misses this year*), a **value** (*driving is her independence — her whole life out there*), and a **promise** (*we told dad we'd look after her*). The evidence pile: the insurance record of the near-misses, a neighbor's account of one of them, a doctor's note about her eyesight, and grandma's own testimony that she only drives familiar roads in daylight.\n\n**The expert witness.** Someone bolsters an argument with: *\"Dr. Smith is a credible expert on labor economics.\"* Mapped, that also splits into three parts: *she has a strong publication record*, *her field is actually labor economics* (not, say, finance), and *she's in good standing in her profession*. Notice these three work differently from grandma's three — which is itself one of the problems below.\n\n---\n\n# How meaning splits — and adds back up\n\n*A mapped claim is a whole built from parts: support flows down when evidence attaches to a part, and flows back up when the parts' scores combine into the whole's. These three problems are all conservation worries about that flow — splitting can spill meaning that never lands anywhere, and adding up can quietly mint support that was never contributed.*\n\n## 1. Does the score survive the slicing?\n\nWhen we map grandma's claim, every piece of evidence gets filed under the part it supports. But reasonable people file differently: does the doctor's note about her eyesight support only the near-miss *fact*, or does it also weigh on the *independence* question (worse eyesight means less freedom either way)? Does the neighbor's account belong under the fact, or is it too secondhand to file at all? Every filing choice nudges the part-scores, which nudges the overall score — and there is no objectively correct filing. So a skeptic can say: *your scores reflect your filing choices, not the evidence.* A neighboring scientific theory (assembly theory, which scores molecules by counting their building steps) was seriously attacked on exactly this point — \"your number is an artifact of how you sliced things\" — and had no answer, because it never ran the test on itself. **We have now run it on ourselves** (August 2026): all 87 decomposed claims in the live map, recomputed under four reasonable filings — as filed, never-decomposed, joins-read-the-other-way, shared-evidence-counted-once. The verdict survived every filing for 91% of them, the typical score did not move at all, and both genuinely sensitive claims turned out to hinge on one unrecorded choice (are the parts *jointly required* or *independently supporting*?) — a choice the map can record, which converts the sensitivity from a diffuse worry into a fixable state. Honest limits: today's map is machine-filed and lightly engaged, so the study ships as a re-runnable instrument rather than a certificate — the numbers will be re-checked as real people answer real questions. **Progress now looks like**: the sensitivity flag surfaced on the affected claims' own pages, and the study re-run standing as the corpus gains adversarial, human filings.\n\n## 2. What does \"adding up the parts\" mean?\n\nLook at the two running examples side by side. The expert witness's three parts are **all required** — if Dr. Smith's field turns out to be finance rather than labor economics, her publication record and good standing don't rescue her expertise. It's a stool: the weakest leg decides. Grandma's three parts are different — the fact, the value, and the promise **each independently push** toward taking the keys; losing one still leaves two real reasons. A scoring system must treat these two shapes differently (we recently taught ours the stool case), but two shapes are surely not enough: some considerations undermine each other, some only work in combination, and in questions of value, interaction between considerations may be the rule rather than the exception. **Progress looks like**: a small, tested vocabulary of ways parts can relate to their whole — added only where someone actually disputes how the parts combine.\n\n## 3. Can evidence be divided honestly at all?\n\nA century-old result in philosophy of science says evidence never supports a single isolated claim — it supports a claim *together with* a web of background assumptions. The doctor's note \"supports\" the near-miss fact only if you also trust doctors' eyesight assessments in general, this doctor in particular, and the note's relevance to driving — none of which anyone wrote down. If that's right, then \"which part does this evidence belong to?\" may sometimes have no true answer at all. Our response is to manage the problem rather than pretend to solve it: evidence gets *linked* to everything it genuinely bears on instead of filed in one drawer, \"can't tell\" is an allowed answer, evidence shared between parts is discounted so it isn't counted twice, and each score can say how much of it rests on shared evidence. Is honest management enough, or does the fuzziness run deeper? **Progress looks like**: the shared-evidence note published beside every score, and problem #1's study showing the management holds up.\n\n# What the map cannot yet see\n\n*These three are blindness problems: meaning hiding behind different words, support that doesn't come as steps, and value living in what's deliberately left unsaid.*\n\n## 4. When are two claims the same claim?\n\nOne economist writes *\"the minimum wage costs jobs.\"* Another writes *\"raising the wage floor reduces employment among the low-skilled.\"* Same claim — zero shared words. Everything a shared map promises — reusing work, spotting duplicates, catching someone attacking a caricature of an opponent instead of the real position — depends on recognizing sameness like that. Nobody has solved it. Word-similarity tools find matching *vocabulary*, not matching *meaning*: when we mapped two legal experts arguing opposite sides of the same question, the software found almost no shared claims between them, even though they plainly discussed the same things — each side simply says it in their own words. **Progress looks like**: a matching method that finds same-argument-in-different-words pairs reliably enough to build on.\n\n## 5. Is one style of thinking quietly favored?\n\nOur official form of examination is step-by-step: state the claim, list the evidence, answer the challenge questions, repeat. But grandma's own testimony — *\"I only drive familiar roads in daylight\"* — carries the weight of sixty years behind the wheel, and some knowledge comes exactly that way: testimony, long practice, lived experience. A midwife's judgment, refined over a thousand births, may be scrutinized by her community in its own rigorous way for generations — yet a step-by-step examination scores it \"unsupported\" simply because its support doesn't come as steps. **Progress looks like**: either an honest account of what step-by-step examination cannot see, or a second style of examination that doesn't secretly reduce to the first.\n\n## 6. What should never be taken apart — and who decides?\n\nSome agreements survive only because nobody spells them out. In grandma's family, the peace may quietly depend on never settling who has the final say about her life — spell *that* out, and you might win the argument but lose the Sunday dinners. Peace treaties work the same way: write down exactly what each side means by the key sentence, and the treaty collapses. A machine whose whole purpose is inviting things to be spelled out has no concept of *productive* vagueness — and no rule for who must consent before someone runs a treaty, a family compromise, or another person's cherished belief through the machinery. **Progress looks like**: a way for the map to mark \"this vagueness is doing work — handle with care,\" and a consent rule for dissecting texts whose stakeholders never asked.\n\n## 7. What does the machine leave out — and which way does it lean?\n\nSeveral leak-points we've already found come with watchdogs built or fixes designed: the extraction step flattening two opposed claims toward a mushy middle (watched by a dedicated instrument), a summary dropping one side of a conflict (counted, per answer), the crucial premise both sides assume but neither writes down (a detector exists, in propose-first form). Those aren't the question — we expect use of the system itself to keep closing them. The question is the leans we can't see from inside. Three we can name: the AI models doing the reading carry whatever tilts their training gave them; sharper, the instruments we use to *check* the machinery run on the same class of models as the machinery itself — a circularity nobody has broken; and before any mapping happens there's the quiet agenda-setter — *what never gets mapped at all*. Whose questions have no map is the oldest form of agenda power, and no instrument of ours can see an absence. Our stance for every lean we find: it must be visible, attributable — is this what people have contributed so far, or is it the machinery? — and treated as an invitation to contribute the missing side, never a hidden nudge. But that covers the leans we've found. **Progress looks like**: you naming the ones we haven't. This is the page's widest question, deliberately — if you can see a structural bias in how a system like this ingests, maps, or summarizes, we're all ears.\n\n# Can the commons stay alive?\n\n*These three are survival problems: who feeds the map, who governs it, and what stops it from quietly rotting.*\n\n## 8. What does an AI reader owe the map?\n\nSoftware may become the map's biggest reader: AI assistants answering people's questions could pull from it constantly. There's a cautionary tale — Stack Overflow, the programmers' question-and-answer site, watched its community wither once AI systems began consuming its answers at scale and giving nothing back: no visits, no credit, no new contributors. Requiring attribution by license is a floor, not an answer. **Progress looks like**: a worked-out answer to what machine readers owe the commons they read — before machine reading dominates.\n\n## 9. Who guards the guards?\n\nEvery measuring instrument in the system publishes its own failure modes — but the same person who builds the instruments decides what they measure. There is no published rulebook: no versioned public statement of the scoring rules and category systems, no procedure for changing them, and no governance beyond one founder. Our own analysis predicts this is exactly where the project would first be caught betraying its principles. **Progress looks like**: a public rulebook with an amendment procedure, and the right for contributors to take their contributions and leave — so that exit is real, not theoretical.\n\n## 10. Can the map learn to feel time?\n\nFor most of this project's life, nothing in the map aged: the doctor's note from three years ago scored today exactly what it scored the day it was written. The first machinery now exists — every claim carries the date it entered the map, a deterministic sweep proposes \"this support may have aged out\" against per-kind decay horizons, and a human can ratify that a piece of evidence has lapsed (the ratification is itself a claim anyone can challenge; lapsed evidence stops counting). But the hard parts are still open: the decay horizons are our guesses, not measured facts about how fast each kind of evidence really rots; a settled *verdict* still never expires on its own; and the best-documented way knowledge infrastructure dies is quiet neglect of maintenance — the proposals now exist, and nobody yet has a reason to be the person who reviews them. **Progress looks like**: horizons calibrated against reality instead of guessed, expiry reaching verdicts and not just evidence, and a maintenance role someone actually has a reason to fill.\n\n# And the bets underneath it all\n\n*If the numbers hold, the blind spots shrink, and the commons survives — two staked bets remain. These are not problems we can't solve; they are positions we've taken, with reasons, that our own instruments can score against us.*\n\n## 11. Does all this structure actually beat a well-kept wiki?\n\nThe honest rival isn't chaos — it's good prose with links, kept by careful editors, with an AI on top to answer questions. A thorough wiki page titled \"Should grandma stop driving?\" might serve the family well. Our bet is that *machine-readable* structure earns its extra cost: because the computer knows which claims attack which, an inconsistency can be *computed* rather than waiting for an editor to notice, and scores update automatically when something upstream changes — the doctor's note gets challenged, and every claim resting on it feels it at once. But at today's size that advantage is mostly still a promise, and at one layer (linking claims across different sources — problem #4) our approach is currently *failing* while a wiki would shrug. **Progress looks like**: the head-to-head — the same questions answered from raw documents, from a flat list of claims, and from the full map — scored honestly.\n\n## 12. Do values actually converge at the bottom?\n\nOur boldest bet: dig far enough beneath any two worldviews' disagreement and what remains is small, nameable, and largely shared — different *weightings* of a common basis, not alien bedrocks. We hold this with confidence, for reasons: across domain after domain, *structure* converges while *weightings* diverge. The same circular structure of human values replicates across dozens of countries while the rankings differ freely; the same handful of moral foundations shows up in everyone while political tribes load them differently; the same seven cooperative rules are judged *good* wherever they show up — across 60 hand-coded societies and 256 machine-read ones, not one treats them as bad; even in mathematics, strikingly different axiom systems recover the same core theorems. And the one time values were elicited *structurally* in an experiment — context by context, which-is-wiser by which-is-wiser — participants overwhelmingly converged on the directions without being told them. That is exactly the shape our own first mapped descents found: shared ground plus a small, precisely-typed residue of genuine difference. And the bet can still lose: our residue map counts every irreducible clash *against* it, and a map built from real descents that keeps finding alien bedrock would mean we were wrong — visibly, in public, on our own scoreboard.\n\n---\n\n*Every problem here has a longer version with the full reasoning in the [research corpus](../README.md). This page states them; the corpus argues them.*\n"}