{"path":"research/assembly-theory-and-the-reuse-mechanism.md","content":"# Assembly Theory and the Reuse Mechanism\n\n**Date**: 2026-08-16\n**Status**: Research. Nothing here is decided. Two items are close to buildable and one is a live open problem shared with a physics research programme.\n**Prompted by**: the founder, on hearing the town-and-road framing of his own cost prediction — *\"This has deep connections to Assembly Theory by the way!\"* Source: Sara Imari Walker in conversation with David Kipping, *Cool Worlds* #22, 2025-05-18. Transcript held privately (auto-captions, third-party content); passages quoted here are checked against it.\n\n---\n\n## 1. Why this is not a loose analogy\n\nAssembly theory measures an object by **the minimum number of construction steps required to build it, where anything already built may be reused**. Reuse is not an optimisation in that definition. It *is* the definition. Walker states the mechanism directly:\n\n> \"you can look at the minimal number of recursive steps **where you're reusing parts** to make that molecule\" — [0:33:22]\n\nAnd then states the reason reuse is forced rather than merely convenient:\n\n> \"at a certain complexity threshold, the space is so large that there's **no exhaustive search possible**. So anything you build in that space has to be along a constrained space where you **build from parts that already exist**. … So when you get into the world of complex objects, it's **always depth**, and you always have to build it along **historically contingent pathways**.\" — [0:36:31]\n\nThat is the founder's registered prediction, arrived at independently and stated as physics. His version: the investment falls asymptotically as the graph fills, because new work terminates in material already mapped. Walker's version: above a complexity threshold, building from existing parts is **the only way anything gets built at all**, because the space of possibilities is unsearchable.\n\nThe two claims differ in force, and the physics version is stronger. The founder's says construction gets *cheaper*. Walker's says that past a threshold, construction from scratch is *not available* — a planet that tried to enumerate molecules at RNA-nucleotide complexity, one copy each, \"would collapse to a black hole\" [0:18:48]. Applied here: a corpus attempting to map contested reasoning by deriving every argument from bedrock, every time, is not slow. It is impossible, and its apparent slowness is the impossibility showing up as a budget.\n\n---\n\n## 2. The mapping, term by term\n\n| Assembly theory | Deliberus |\n|---|---|\n| Elementary building blocks | Atomic claims, bedrock premises |\n| **Assembly index** — minimum construction steps *with reuse* | Decomposition steps to reach a claim, reusing already-mapped claims |\n| **Assembly pool** — what has already been made | The corpus of mapped claims and concepts |\n| **Copy number** — how many instances exist | How many arguments draw on a given premise (`usage_count`) |\n| **Assembly** *A* — exponential in index × linear in copy number | An argument both deeply structured *and* widely reused |\n| Historically contingent pathway | Provenance: `EXTRACTED_FROM`, `DECOMPOSES_INTO` |\n| Joint assembly space over a whole sample | The graph as one object rather than a bag of claims |\n\nThe provenance row is not decoration. Assembly theory's central commitment is that a high-assembly object **carries its own construction history**, and that this history is what distinguishes it from a merely complicated object. A claim graph retains exactly that. The graph is an assembly space over arguments, and it is one of the few artefacts in the world that is.\n\n---\n\n## 3. What today's measurement looks like in these terms\n\nThe most directly transferable idea in the conversation is **joint assembly** — not the index of one object, but of a whole coexisting set:\n\n> \"if you think about all of the molecules you detect above an abundance threshold, you can construct an assembly space for **the entire set** of molecules, not an individual molecule. … it's about the **minimal causation and constraint to actually have these molecules coexist**.\" — [1:03:10]\n\nAnd the discriminating result, which is the sentence to keep:\n\n> \"**Earth really stands out** … because it has a high diversity of molecules in high abundance that have **the same bonds**, but many different molecules. And you get like a **Jupiter-like** atmosphere, you'll have a lot of diversity of compounds, but they'll be made with **different bonds**. And so there's something about the **reuse of bonds** … you can tell it was a high assembly atmosphere **because of the depth of the space rather than a flat one** just because you had a high configuration.\" — [1:04:12]\n\nDeliberus measured its own version of this on 2026-08-16 without knowing it had a name.\n\n**At the claim layer the corpus is Jupiter.** 4,769 claims, high diversity, and almost nothing shared: 56 `SIMILAR_TO` edges across the whole graph, 16 implicit-premise claims of which exactly one is linked to anything, and `decompose_claim` performing no lookup of existing claims at all. Diversity without shared bonds. A flat assembly space.\n\n**At the concept layer it is Earth.** 87% of concepts recur, 479 usages, a reuse distribution flatter than Zipfian. Same bonds across many different objects. Depth.\n\nThis is a better framing of the same finding than the one the corpus already had, because it says *why the number matters*. Concept reuse at 87% against claim reuse near zero is not two statistics. It is one statement: **the corpus has built an assembly pool at one layer and not at the other**, and the layer without a pool is the layer where every construction still starts from raw material.\n\n---\n\n## 4. The hardest problem in assembly theory is the hardest problem here\n\nThis is the deepest convergence and it does not flatter either side.\n\n> \"**the copy number is the hardest part — it is actually identifying the boundaries of the selected units.** This is a known problem in biology, like what's the boundary of an individual unit of selection, and assembly theory is trying to generalise that concept to any material.\" — [0:42:55]\n\nCopy number requires knowing when two things are **the same thing recurring**. In chemistry that is nearly free; a molecule's identity is given. In Deliberus it is the open philosophical fork the corpus has already flagged as load-bearing, and which now blocks three separate pieces of work: the straw-man detector, the reuse lookup in decomposition, and any assembly-style measure imported from here.\n\n**So assembly theory does not hand over a solution. It hands over the news that the difficulty is principled.** Walker's programme, with a decade of work and a chemistry substrate where identity is nearly given, still names unit-of-selection boundaries as its hardest part. That reframes claim-sameness from a local embarrassment into the known hard problem of a research field — which is a reason to expect it to be hard, and a reason not to expect a clean answer by trying harder.\n\n---\n\n## 5. Three ideas worth taking, beyond the reuse mechanism\n\n**Downward causation, which argues for building the argument layer first.**\n\n> \"most of the molecules in cells exist **because a cell exists** … global constraints from cellular architecture that **allow these molecules to even be possible**.\" — [0:17:47]\n\nThe inverse of the mother-claim relation. A mapped argument does not merely *contain* its premises; its existence makes certain premises articulable that nobody would otherwise state. That is an argument against the intuition that one should build bedrock upward, and it fits the observed behaviour of the corpus, where premises surface because an argument called for them.\n\n**The cost thesis has a stronger form than the founder's own.**\n\n> \"the biosphere … is trying to **compactify as much possibilities in a small volume of space as possible**. … if you're organised you have a lot more **access to structures you can generate** … **Life is the physics that makes it possible for some things to exist.**\" — [0:55:27], [0:56:47]\n\nThe prediction says a mature corpus makes reasoning *cheaper*. This says the deeper effect is that it makes reasoning *possible* — it enlarges the adjacent possible rather than discounting the existing one. Worth considering as an upgrade to the registered prediction rather than a gloss on it, since it is a different and more testable claim: not \"descents get shorter\" but \"descents get made that would not have been attempted.\"\n\n**A criterion for whether the instruments are real.**\n\n> \"if you had an explanation for life, it would be like an explanation for gravity. **You can design other experiments, you understand suddenly all this other stuff** that you didn't understand.\" — [0:29:04]\n\nWalker's distinction between a *signature* (flags a thing) and an *explanation* (opens new questions) is a usable test for the honesty instruments. The completeness oracle, the hinge score and the stance detector are currently signatures. Whether any of them has opened a question that could not previously be asked is a check worth running against the record.\n\n---\n\n## 6. What does not transfer, and the ways this could mislead\n\n**Copy number is not a quality signal.** Walker's crystal counterexample [0:38:42–0:40:52] is that high assembly index is achievable without life, and the resolution turns on identifying the genuinely repeated unit. The corresponding hazard here is sharper and it runs the wrong way for us: **the most-reused claim in a corpus may be the least examined one.** A cliché has enormous copy number precisely because nobody decomposes it. If reuse ever becomes a ranking signal, it will rank platitudes first. Any adoption needs reuse crossed with decomposition depth, never reuse alone.\n\n**Our regime is the shallow one.** Walker notes that molecular and mineral assembly spaces have abrupt thresholds while planetary atmospheres have \"a very shallow assembly space … it's not a sharp boundary\" [1:04:12]. Twenty-five sources is the shallow regime. Expect gradients, not a threshold, and treat any sharp number as suspicious.\n\n**Assembly theory is contested, and the literature check run the same day found it more heavily contested than this paragraph first said.** The assembly index has been argued in a formal proof its authors have not overturned to be **mathematically equivalent to Shannon entropy rate via dictionary-based compression**, an LZ-family scheme (*PLOS Complex Systems* 2024; *npj Complexity* July 2026); the MA > 15 life-detection threshold was **independently found in error by two separate groups**; and a sympathetic reviewer concludes it \"has merit but is not nearly as novel or revolutionary as claimed\" (*Journal of Molecular Evolution* 2024). **So its authority is not available to borrow — no document here may write \"assembly theory shows that…\", least of all a funding application.** What survives is the mechanism, which is not original to it: construction-with-reuse in a space too large to search belongs to Kauffman's adjacent possible and Arthur's combinatorial evolution, and those are what the timing argument should cite. Full state of the dispute and what it changes: [complexity-transitions-and-why-now.md](complexity-transitions-and-why-now.md) §0.\n\n**One criticism of it is the most useful thing here.** The critics' central technical charge is that the index \"only takes into consideration **identical copies** rather than the full spectrum of causal operations to which an object may be subject\". That is Deliberus's own measured problem said by someone else about something else — run 6 found structurally kindred claims at ~0.60 cosine, and claim-level reuse is near zero because similarity measures vocabulary while kinship is structural. **The design consequence arrives before the build: a reuse lookup in decomposition must match structurally, never by identity.**\n\n**And one parallel above needs a caveat.** Downward causation is a good idea about Deliberus, but the critics argue the formalism cannot actually measure top-down causation, only bottom-up. Carry it as an intuition with independent justification, never as something assembly theory demonstrates.\n\n**And the direction of borrowing matters.** Nothing here licenses the claim that a claim graph *is* an evolving system or that reasoning *is* alive. The transferable object is narrow and precise: a formalism for construction-with-reuse in a space too large to search, plus the finding that identifying repeated units is the binding difficulty.\n\n---\n\n## 7. One unclaimed opening, stated plainly\n\nWalker names the application herself and says she has not done it:\n\n> \"I've thought about studying the **evolution of mathematics through an assembly theoretic lens**, because there's a lot of people that study evolution of theorem proving and **how theorems are built on other theorems** and how mathematicians navigate that space, and I think it's an **incredibly rich space to study evolutionary principles, but I haven't done any of that work myself. It's on my bucket list, but there's only so many projects you can do.**\" — [0:48:08]\n\nTheorems built on theorems is one step from arguments built on arguments, and it is the analogy this corpus already runs at length in [lean-deliberus-analogies.md](lean-deliberus-analogies.md). Deliberus holds something her stated bucket-list project would need and mathematics does not readily provide: a growing, machine-readable assembly space over **contested** reasoning, with provenance retained and construction steps recorded.\n\nThat is a research opening, not a claim about anyone's interest. What it justifies is a note in the funding and outreach material that the reuse mechanism has an independent formal treatment in another field, and that the corpus is an unusual substrate for it.\n\n---\n\n## Cross-references\n\nThe cost thesis this supplies mechanism for: [lowering-the-cost.md](lowering-the-cost.md) §5, §8. The measurement it reframes: [self-similar-decomposition-and-claim-ontology.md](self-similar-decomposition-and-claim-ontology.md) § The mechanism audited. The theorem-building analogy: [lean-deliberus-analogies.md](lean-deliberus-analogies.md). Why identifying repeated units is load-bearing here: [dogfood-run-6-israel-palestine-cross-domain.md](dogfood-run-6-israel-palestine-cross-domain.md). The scale ladder that Walker's multiscale selection [0:16:39] rhymes with: [fractal-scales-and-temporal-frame.md](fractal-scales-and-temporal-frame.md).\n\n\n## The code question — shared-code qualifiers read through Assembly Theory (2026-08-18, founder-prompted)\n\nThe founder heard the shared-code unlock criteria (convictions-across-scales.md § 2.1) and recognized Assembly Theory in them. The intuition is right, and it sharpens both sides.\n\n**What AT requires for high-assembly objects to exist in abundance**: a pool of reusable parts, a memory that persists parts across time, combinatorial composition, and copy number. Map these onto the six code criteria and they interlock rather than duplicate: *composability* is assembly stepwise-construction; *cheap write/read* is what makes copy number reachable; *stability under transmission* is what makes abundance evidence of selection rather than noise; and *claim identity* is AT's own hardest problem arriving in our substrate (Walker verbatim: identifying the boundaries of repeated units — already this doc's keeper #2). The code IS the memory format: what a genetic code does for proteins, the claim/relation/scheme/provenance stack does for arguments.\n\n**The AT-form of the founding conviction**: unaided discourse has a low assembly ceiling — an essay rebuilds its argument from scratch every time, so there is no persistent parts-pool, and arguments deeper than working memory cannot exist in copy number. The shared code's function, stated in AT terms, is to **raise the assembly ceiling of arguments that can persist and replicate**. That is \"opacity is a cost\" and the falling-cost prediction in one sentence of physics-adjacent vocabulary.\n\n**A possible instrument falls out**: AT detects selection by TWO axes at once — assembly index × copy number. A cliché is high-copy/low-assembly (this doc's standing warning); a private insight is high-assembly/copy-one (Antikythera); **epistemic selection is high-assembly AND high-copy across independent sources** — an idea that survived many minds' reconstruction. Both axes are computable from the graph (descent depth with reuse discount; cross-source recurrence), which would turn \"never use reuse alone as a quality signal\" into a two-axis rule with a measurement. Propose-only if built, per house rules; and AT's authority stays unborrowed (the index-equals-compression critique stands — never \"AT shows that…\").\n\n**The exhaustiveness correction (same conversation)**: calling the Walton scheme set \"the codon table — a closed set of legal inference moves\" over-claimed. The genetic code is closed by physics (64 codons); Walton's schemes are an **empirical taxonomy induced from practice** — counts vary across his own compendia, and the corpus's own three-runs finding applies (*every taxonomy carries exactly the distinctions of the register it was induced from and breaks on the next*). No completeness theorem exists for defeasible inference the way one exists for deductive logic. So the honest statement: a **curated open set with a closed-enum implementation**, which by the corpus's standing rule about closed enums needs a **does-not-fit confession value** on the scheme classifier — needs verification whether one exists; if not, that is a named gap. Two live readings of what the set is converging toward: the periodic-table reading (open but with predictable gaps) and the generative reading (schemes may decompose into a small closed basis of primitive moves plus the CQ mechanism, the way 64 codons come from 4 bases) — the second is a genuine research question the self-similar principle invites, since the scheme set itself is \"currently undecomposed, invitable deeper.\"\n\n**And the meta-conviction, flagged by its holder**: \"theory, analogies, and fractal evolutionary priors can enlighten the corpus\" is itself a conviction the convictions-across-scales method applies to — its check is that every analogy must pass mechanism-identity (check 2) or be demoted to organizing metaphor, which is exactly the discipline this doc's \"do not borrow AT's authority\" rule already enforces.\n\n\n**Update (2026-08-21)**: the boundary conundrum gained its third face — the founder connected it to the join-semantics problem (evidence division + conjunction/corroboration), and the connection cuts both ways: boundaries say what the units are, join semantics say what combining them means, and AT holds only the first half while our architecture's escape (contestable, provenance-carrying boundaries) is the move its formalism lacks. [evidence-division-and-the-foundation.md](evidence-division-and-the-foundation.md) §2c.\n"}