{"path":"research/sameness-merge-and-split-across-fields.md","content":"# Sameness, Merge, and Split Across Fields\n\n**Date**: 2026-08-25\n**Type**: Literature grounding for the live sameness/auto-merge design loop (briefing-stack Loop 1). Ordered by the founder mid-loop: *\"Look up academic papers etc on all of this, think up potential related academic fields that might have published on closely related issues.\"* Also answers two founder questions from the same loop: whether concept-tagging solves the post-split re-homing problem, and how claim merging and concept merging relate.\n**Sits beside**: [soft-canonical-clustering-and-reversible-merge-semantics.md](soft-canonical-clustering-and-reversible-merge-semantics.md) (the April architecture), [claim-sameness-philosophical-readings.md](claim-sameness-philosophical-readings.md) (the philosophical cross-check), [conceptual-threads.md § Thread 4](../conceptual-threads.md) (the dedup problem).\n\n---\n\n## 0. The loop context this feeds (recorded so the doc stands alone)\n\nThe residual-error taxonomy named identity scattering as a systematic strength distorter; the founder ruled that sameness-detection and \"robust and wisely balanced auto-merge\" of claims and concepts must exist. Three founder contributions inside the loop, all load-bearing:\n\n1. **Cluster-carries-the-strength is his from ~2008** — *\"I've known this would be a weakness since when I first was sketching all this out.\"*\n2. **Atomization is crucial to comparability** — sameness is crisp only between claims at matched grain; a bundle matches a *part* of another claim (containment, not identity).\n3. **The one-step invariant** — if deduplication holds recursively below, double-counting in a group's consolidated evidence can only arise one edge-step away, because deeper structure is already shared. *\"Everything upstream/deeper will be the same.\"*\n\nHis proposed shape: no write-time suggestion friction — anyone types anything; careful auto-merge afterward; anyone can split when truly warranted. Conditions established in-loop: merges visible at the claim surface, split as an accountable act, author notified on merge, and the machinery measured (blind-labeled pair eval) before it switches on. Interim principle for the tier question: **the machine may act on the map; it may only propose about the territory** — with value-claim merges straddling the line, handled by bar-height rather than separate workflows.\n\n---\n\n## 1. Does concept-sense tagging solve the re-homing problem?\n\nThe problem: claims A and B are merged; during the merged period someone answers a question or attaches evidence *to the group*; the group is later split — which member was the contribution about?\n\n**The founder's instinct (tag which exact concept-senses each claim uses) is right, and it generalizes.** The generalization: **every legitimate split carries a discriminator** — the feature by which the members turned out to differ (different senses of a shared term, different scope, different time-frame, different asserter). That discriminator is *exactly the classifier to apply to during-merge contributions*: an answer that engages \"freedom as non-coercion\" re-homes to the member using that sense. The split's own justification generates the re-homing rule. Sense-tags are the most common discriminator type, and they serve **twice**: as a *pre-merge gate* (differing sense-tags on a shared term should block or heavily penalize the merge — preventing the most common false-merge class before it happens) and as the *post-split sorter*.\n\n**But it is necessarily more complex than tags alone — a three-outcome protocol, not a function:**\n\n1. **Discriminator-classified** → re-home to one member (the majority case when the split reason is crisp).\n2. **Genuinely shared** → a contribution engaging only what A and B have in common (\"the study's sample was small\") legitimately attaches to *both* — not ambiguity but common ground, which the split does not erase.\n3. **Irreducibly ambiguous** → a contribution whose own text doesn't touch the discriminator, made by someone who saw the merged view. Their mental referent was an object that turned out not to exist (the false unity). No classifier can recover an intention that was never differentiated. Honest handling: keep it attached to a record of the merged-period object, marked *made during merge, referent unresolved*, and ask the contributor — the author-notice pattern pointed the other way. A closed enum here would force a nearest fit; this is the confession-channel principle applied to identity history.\n\nOne dependency to respect: sense-tagging is itself machine output (the concept-usage edges), so the discriminator classifier must not share a mechanism with the merge machinery it audits — the two-detector independence condition, again.\n\n## 2. How claim merging and concept merging relate\n\n**Layered mutual dependency, with sameness flowing upward.** Claims are built from concepts; two claims can only be the same if their shared terms are used in the same senses — so concept-sense identity is an *input* to claim identity, exactly as claim identity is an input to question identity (the birth-certificate chain from Stage 1). But concept senses are themselves individuated by their use across claims — the sense IS the usage pattern. Circular but not vicious: it's an alternating refinement (bootstrap senses from claim usage, use senses to sharpen claim identity, repeat), the co-evolution of dictionary and speech.\n\n**The chiasmus — the default error runs in opposite directions per layer:**\n\n- **Claims want merging, with careful splitting.** The wild produces *duplicates*: many people restate one proposition. The workhorse operation is grouping; the failure to watch for is false unity.\n- **Concepts want splitting, with careful merging.** The wild produces *false unities*: one word covering incompatible senses. The workhorse operation is sense-discovery (the \"we used the same word differently — THAT's why we disagree\" moment is the product's aha); the failure to watch for is over-fragmentation.\n\n**And the stakes multiply differently.** A false claim-merge corrupts one group's accounting. A false concept-sense merge corrupts *every claim using the term* — concepts are shared infrastructure, so sense identity errors are the higher-blast-radius class. This is the April doc's \"concepts are more dangerous\" made mechanical.\n\n## 3. The fields, and what each has already solved\n\n### 3a. Biological taxonomy — the deepest kin, and it discovered the one-step insight independently\n\nTaxonomists have run a governed identity system for centuries: names (words) vs **circumscriptions** (what the name actually covers, per a given author) — and their crisis is exactly ours. Franz & Peet's concept-mapping language and the \"Names Are Not Good Enough\" study ([Franz et al., Semantic Web Journal](https://www.semantic-web-journal.net/system/files/swj623.pdf)) measured it: the name *Andropogon virginicus* denotes **six non-congruent concepts** across eight classifications, and names were reliable identity proxies for only ~60% of aligned regions. Their solutions, all transferable:\n\n- **The \"sec.\" convention**: a name is individuated *according to* whose circumscription — \"Andropogon virginicus **sec.** Weakley 2006.\" Identity is author-indexed by syntax. This is sense-attribution as first-class citizenship, and the precedent for claim-identity being indexed rather than absolute.\n- **Five graded articulations, not same/not-same** (Region Connection Calculus): *congruent, includes, is-included-in, overlaps, excludes*. The founder's containment-not-identity point has a formal vocabulary waiting: a bundle *includes* an atom; two overlapping bundles *overlap*. Cluster relations richer than binary sameness are a solved representation problem.\n- **Lump/split reconciliation through the children**: differences between a lumper's and a splitter's classifications are \"frequently reconcilable through addition (+) or subtraction (−) of lower-level concepts\" — i.e., **compare composites via their parts** — the one-step invariant, discovered independently, in production use for reconciling rival taxonomies, with logic reasoners (Euler/X) computing consistent multi-taxonomy alignments from expert-asserted part-relations.\n- **Intensional vs ostensive identity** (INT/OST): a concept compared by its *defining properties* vs by *what it points at* — the same split as a sense defined by a definitional claim vs by its usage examples. Two identity modes, kept separate in their notation.\n\n### 3b. Entity resolution / record linkage (databases, ~55 years)\n\nThe merge machinery literature. Directly transferable:\n\n- **Incremental linkage with repair** ([Gruenheid, Dong, Srivastava, VLDB 2014](http://www.vldb.org/pvldb/vol7/p697-gruenheid.pdf)): merge, split, and *move* as first-class incremental operations, and — the key result — **newly arriving records supply evidence that fixes previous linkage errors** (new members dilute a false cluster and suggest moving the odd one out). The Stage-5 alarm has an algorithmic ancestor: arriving evidence, not periodic audit, is the natural split-trigger.\n- **Materialized sub-partitions for future splits** (Whang & Garcia-Molina, VLDB J): keep finer-grained groupings alongside the operative one, so when the matching rule tightens, clusters split cheaply from stored parts instead of recomputing. Argues for storing *why* each member joined (which signals), which the April doc's \"cluster-entry reason\" already specifies.\n- **Clerical review** (Fellegi-Sunter tradition; [Detective Gadget 2024](https://iris.unibas.it/retrieve/02b98d44-dc54-4906-9e0a-adf7f5b64635/data-09-00139.pdf)): a match band, a non-match band, and a *borderline band routed to humans* — plus false-positive check functions that flag \"suspicious groups\" for expert eyes. The type-sharded bar-height design is this field's standard practice wearing our ontology.\n\n### 3c. The Semantic Web's sameAs crisis — the dark twin of the one-step invariant\n\nThe web's identity link (owl:sameAs) was strict logical identity, used sloppily at scale. Results ([Halpin et al. 2010](https://files.ifi.uzh.ch/ddis/iswc_archive/iswc/pps/web/iswc2010.semanticweb.org/pdf/261.pdf); [Raad et al. 2020 survey](https://www.semantic-web-journal.net/system/files/swj2430.pdf)): 3–20% of identity links erroneous, and because **identity is transitive, single bad links chained into catastrophes** — one closure falsely unified 177,000 names for different countries, cities, and people (\"mushy peas\"). The lesson for us is the founder's invariant inverted: **identity infrastructure propagates errors exactly as efficiently as it propagates benefits.** One false merge, once other merges chain through it, is not one error — it is a corridor. Their remedies map onto our ladder: graded identity predicates weaker than strict sameness (closeMatch, nearlySameAs — our similarity edge → cluster-candidate → cluster progression), **context-qualified identity** (sameAs holding only within a named graph — the scheme-indexed-sameness hypothesis from the philosophical readings has a Semantic Web precedent), and a whole subfield for *detecting erroneous identity links*. Design consequence worth pinning: **be conservative about transitive chaining across clusters** — cluster membership should not silently compose across merge decisions the way sameAs closures did.\n\n### 3d. Fact-checking claim matching — production-scale claim identity, multilingual\n\n\"Has this claim been checked before?\" is a production task with datasets of 206k claims across 27 languages ([SemEval-2025 Task 7](https://aclanthology.org/2025.semeval-1.323.pdf); Shaar et al. 2020; CLEF CheckThat). Their standard pipeline is Stage 3's recipe validated at scale: dense retrieval proposes, a stronger model (now LLMs) confirms, with claim *normalization* (their word for decontextualization) as the enabling preprocessing. Multilingual claim matching is a solved-enough problem that Swedish support is a model choice, not a research problem.\n\n### 3e. Key point analysis — the gray zone, measured\n\nIBM's key point analysis task (Bar-Haim et al.; [ArgMining 2021 shared task](https://aclanthology.org/2021.argmining-1.19.pdf)) matches many argument formulations to one concise \"key point\" — many-formulations-to-one-operative-claim as a benchmark task. The number that matters: in their annotated corpus, **22.8% of argument-to-key-point matches were ambiguous — trained human annotators could not agree**. The gray zone in claim sameness is not an artifact of weak machinery; it is a measured property of the judgment itself. A borderline band isn't a concession — it's descriptively correct.\n\n### 3f. Noted from knowledge, unverified this session (flagged per the read-depth discipline)\n\n- **Library science**: authority control (name records merged/split under governance for a century) and **FRBR's layered identity** — work / expression / manifestation / item — a mature ontology in which \"the same work in different expressions\" is exactly operative-claim vs formulations. The cluster-with-members design is FRBR-shaped.\n- **Law**: *res judicata*'s identity-of-claims doctrine (when is a new suit \"the same claim\"? — the transactional test) and case consolidation/severance as governed merge/split of proceedings.\n- **Collaborative editing**: group-undo is a known-hard problem (undoing an operation others have built on) — the during-merge re-homing problem is its cousin; the corpus's CRDT doc already touches the state-merging side.\n\n## 4. What transfers into the design, compressed\n\n1. **Graded identity relations, not binary** — congruent / includes / overlaps as cluster-relation vocabulary (taxonomy's RCC-5; solves bundle-vs-atom representation).\n2. **Arriving evidence as the split-trigger** — the incremental-repair result; Stage 5 should fire on new contributions that dilute a cluster, not only on downstream divergence.\n3. **Store the merge's evidence with the merge** — materialized reasons and sub-groupings make future splits cheap and explainable.\n4. **Borderline band by design** — 22.8% measured human ambiguity; route the band to humans (or to \"visible cluster-candidate, not merged\").\n5. **No silent transitive chaining** — the 177k-name catastrophe; cluster composition is itself a merge decision, never an inference.\n6. **Author-indexed identity has centuries of precedent** — \"sec.\" is the existence proof that identity-relative-to-a-perspective can be operational, not just philosophical.\n7. **The discriminator protocol for re-homing** (§1): split-reason generates the sorter; three outcomes; irreducible residue confessed and asked, never force-fit.\n\n## 5. \"Is this solved?\" — the triage, with measured ceilings (verified 2026-08-26)\n\nThe founder's question after the field sweep. The answer is not solved-vs-unsolved but **three buckets, and the line between them falls almost exactly where the type-sharded design already drew it** — which is the useful finding: our architecture is tracking a boundary several independent fields found by hitting it.\n\n### Bucket 1 — Solved engineering. Adopt, do not invent.\n\nMachine-minted duplicate identity (by construction, no research needed) · the merge/split/move data structure and incremental repair algorithms (entity resolution, ~55 years, production) · graded relation vocabulary (taxonomy's RCC-5) · provenance, audit, versioning (library authority control, taxonomy, version control) · candidate retrieval (shipped here at 0.31s corpus-wide). Nothing in this bucket is a research risk.\n\n### Bucket 2 — A measured ceiling: reliable, imperfect, and the number is knowable in advance.\n\n**Cross-document event coreference** (\"do these two texts describe the same event?\") is the closest measured analogue to claim sameness. State of the art 2025: **CoNLL F1 88.4% on ECB+ and 85.2% on GVC** ([ACCI, arXiv:2506.01488](https://doi.org/10.48550/arxiv.2506.01488)) — but on the harder, less-curated FCC corpus the same task sits around **70% for fine-tuned models, ~77% with the best collaborative method** ([ACL 2024](https://aclanthology.org/2024.acl-long.164.pdf)). So: high-80s in favorable conditions, low-to-mid-70s in unfavorable ones.\n\n**The architecturally decisive finding in that literature: a frontier LLM asked directly underperforms by nearly 10 CoNLL F1 points.** GPT-4 few-shot \"faces substantial adaptability challenges in directly predicting cross-document event coreference structures.\" The winning architecture is *collaborative*: the LLM **summarizes and contextualizes each mention**, and a dedicated matcher makes the pairwise/clustering decision. Consequence for us: **do not build \"ask the model whether these two claims are the same\" as the core** — use the model to normalize and contextualize each claim, then match. The extraction pipeline's decontextualization pass is already that first half, so the recommended architecture is half-built here by accident.\n\nCalibration that limits all of this: **the entity-resolution and coreference literatures are entropy-class** (dirty data, honest variation, no agent modelling the matcher). Our eventual setting is strategy-class — parties who want a claim merged or split for advantage. By the adversary-class rule, that is a **rebuild, not a port**: the maturity is real and its scope is narrower than it looks.\n\n### Bucket 3 — Genuinely stumped, and the numbers say so plainly.\n\n**Sense granularity is the hard wall, and it is worse than expected.** Measured human inter-annotator agreement on fine-grained WordNet senses is **56%**, against **91%** for coarse-grained OntoNotes senses on the same task ([Brown, LREC 2010](http://www.lrec-conf.org/proceedings/lrec2010/pdf/927_Paper.pdf)); granularity, not the number of senses on offer, is what drives the difference. Navigli puts \"a credible upper bound for unrestricted fine-grained WSD around 70%.\" And the sharpest result: Ng et al. (1999) ran a greedy merge to find coarser sense classes reaching κ ≥ 0.8 — for **96 of 191 words the only way to get there was collapsing every sense into one class.** For half the vocabulary there is no reliable middle granularity at all.\n\nAlso here: value-claim sameness (the traditions are unanimous; **22.8% of argument-to-key-point matches were human-ambiguous** in the key point analysis corpus), cross-frame identity (same under one framing, different under another — open everywhere), the during-merge re-homing residue (§ 1, information-theoretic rather than a research gap), and adversarial identity (barely studied in our shape).\n\n### The key that unlocks bucket 3, and it is not an algorithm\n\n**OntoNotes did not discover the correct sense granularity. They \"iteratively merged senses until 90% inter-annotator agreement was reached\"** — granularity was the *control variable*, reliability the *target*, and the loss was accepted deliberately. Lexicography's own answer to the impasse is therefore a governance decision, not a discovery. The same paper reports lexicographers' companion move: annotators need **the option to select a group of senses, or a single broader underspecified sense**, rather than being forced to choose — the confession channel, independently invented in another field.\n\n**Deliberus has a strictly better dial available than lexicography does.** A dictionary must fix granularity a priori for all future uses; we can make it **demand-driven** — two senses need separating exactly when claims using them diverge downstream. That is the pragmatist test, and it is *operational* here in a way it can never be in a dictionary. Consequence worth carrying into the build: **the downstream-divergence split alarm is not only a safety net, it is the granularity-setting mechanism**, and it is better grounded than an agreement threshold because it is tied to consequences rather than to annotator intuition.\n\n### The pattern across every stumped field\n\nTaxonomy answered with *sec.* (identity according to whom) plus expert-asserted articulations checked by a reasoner. Library science answered with authority records under governance. The Semantic Web, after its catastrophes, answered with graded predicates and context-qualified identity. Lexicography answered with an agreement target and an underspecified escape hatch. Conceptual engineering answered by changing the question from *what is this concept* to *what should it be, and who decides*. **Every field that hit the wall built process instead of algorithm** — and a process layer for contested judgment is what Deliberus already is. Where sameness is genuinely contestable, the native move follows: make the identity judgment itself a claim in the graph, supportable and attackable like any other (the *publish it* response from the ontology sorting test; the merge/split tug-of-war then becomes readable signal rather than a nuisance).\n\nThe honest cost of that answer: **the risk transfers rather than disappearing** — from \"no algorithm can do this\" to \"nobody shows up to do the judging,\" which lands squarely on the engagement gradient and on ratification decay (residual-taxonomy classes 6 and 7).\n\n**Cross-references**: [the-residual-error-taxonomy.md](the-residual-error-taxonomy.md) (identity scattering, class 3) · [soft-canonical-clustering-and-reversible-merge-semantics.md](soft-canonical-clustering-and-reversible-merge-semantics.md) · [claim-sameness-philosophical-readings.md](claim-sameness-philosophical-readings.md) (the traditions; the scheme-indexed hypothesis that § 3c gives a precedent for) · [semantic-disambiguation-and-concept-tracking.md](semantic-disambiguation-and-concept-tracking.md) (the concept layer § 2 rides on) · [retrieval-instruments-beyond-cosine.md](retrieval-instruments-beyond-cosine.md) (what proposes candidates) · [crdt-collaborative-graphs.md](crdt-collaborative-graphs.md) (state-merge cousin).\n"}