{"path":"research/n-of-1-uplift-protocols.md","content":"# Measuring Personal Capability Uplift: Two n-of-1 Protocols\n\n*Methods note, July 8, 2026. Motivated by the funder question \"what would demonstrate that this actually increases ability to solve hard problems?\" applied at n=1 (the founder). Neither protocol is committed yet; this documents the designs so they can be picked up deliberately rather than improvised. Companion: [dogfood-run-2-orthogonal-experiments.md](dogfood-run-2-orthogonal-experiments.md), the group-evaluation design in the funding applications.*\n\n> **Status note added 2026-08-17 — read this before citing the doc as the project's evaluation plan.** These protocols measure **personal capability uplift**, and on 2026-08-16 the founder rejected that as the unit against which the project is *justified*: a commons is not evaluated by what one reader gains from it. The measures that fit are corpus-shaped ([lowering-the-cost.md](lowering-the-cost.md) §8). Neither protocol has been run.\n>\n> That does not retire them, and the reason is a distinction worth keeping. **Justification and failure-detection point in different directions.** A library is not justified by whether one borrower reads better — and a library that reliably made its borrowers worse would still be a bad library. Individual measurement is unsuited to the first job and well suited to the second. So what these protocols are *for* has narrowed rather than vanished: Protocol A remains the only design here that could return an honest negative from the platform's most skilled user, and Protocol B remains the strongest available friction log. Neither is the project's answer to \"why should this exist\".\n\n## Protocol A: the Forecast-Descent Protocol (rigorous, ground-truth-anchored)\n\nThe credibility problem with any self-experiment is fourfold: no counterfactual, grader bias, question selection, and the time confound (structured work takes longer, and time alone helps). This design kills each one.\n\n- **Questions**: ~50 short-horizon resolvable forecasting questions (Metaculus/Manifold class, resolving within 6-10 weeks), selected by a pre-registered rule (not hand-picked).\n- **Randomization**: each question assigned by coin-flip to Unaided or Descent condition (~25/25).\n- **Unaided condition**: 30 timeboxed minutes; forecast + written rationale + a list of what the question hinges on (logged in Fatebook or equivalent).\n- **Descent condition**: the same 30 minutes, but spent building the argument graph first (extraction, decomposition, hinge check), then forecasting.\n- **Primary metric**: Brier score per condition once reality resolves the questions. Ground truth removes graders entirely.\n- **Secondary metric (Deliberus-native)**: crux coverage — did the hinge instrument surface load-bearing considerations the unaided pass missed, and did those considerations predict resolution?\n- **Pre-registration**: metric, question-selection rule, timeboxes, and analysis plan published publicly before data collection. All graphs public by construction.\n- **Honest limits, stated up front**: n=1; the subject is the builder (maximum skill AND maximum motivation); ~25 questions per condition powers only fairly large effects, so the outcome may honestly be \"suggestive, underpowered\". The design's virtue is that it can come out negative — and a null result from the platform's most skilled user is strong evidence, published the same way.\n- **Cost**: ~25 hours over two months; near-zero money.\n\n## Protocol B: the Deep-Descent Diary (naturalistic, qualitative)\n\nThe founder's own instinct, preserved as its own genre: take one question dear to the heart, or one real pending life decision, or one topic where one's own position is unsettled, and spend a few honest hours walking it through the platform end to end — extraction of the relevant sources plus one's own stated position, decomposition of the load-bearing claims, the weighing descent where values conflict, hinge reading, residue typing. Then write the diary: what did the descent surface that unaided reflection had not? Did the position change, sharpen, or decompose? Where did the tooling fight the thinking?\n\nThis does not clear a skeptic's uplift bar (no counterfactual, no blinding) and should never be presented as if it does. What it produces instead: (1) the strongest possible *friction log* (real stakes expose real UX failures — dogfood runs 1-2 used others' debates; this uses one's own); (2) first-person phenomenological evidence of the kind that makes demos and essays compelling; (3) the honest precursor data for designing Protocol A's descent condition well. The two protocols are complements: B generates the hypotheses and the stories, A generates the numbers.\n\n## Sequencing note\n\nProtocol B can happen any afternoon and is worth doing before the friends round (it rehearses the guided path on the person who will guide it). Protocol A wants a deliberate start with the pre-registration published first, and its results window (6-10 weeks) means a summer start yields data before late-autumn funding decisions.\n"}