{"path":"research/second-opinion-mvp-strategy.md","content":"# Second Opinion: What Would the MVP Actually Prove?\n\n**Date**: March 28, 2026\n**Model**: GPT-5.2 (max thinking mode)\n**Triggered by**: Fredrik's question: \"What would the MVP prove that we cannot already be reasonably convinced of?\"\n\n---\n\n## The Reframing\n\nThe MVP doesn't prove a hypothesis. It proves the project has **crossed the phase transition from *true statements about a product* to *a product that changes someone's behavior*.** That's the only proof that matters for:\n- (a) A platform with high cognitive activation energy\n- (b) A founder with a 15-year pattern of conceptual refinement without shipping\n\n## Five \"Proofs\" Examined\n\n| What MVP supposedly proves | Already known? | Cheapest test | But what the MVP ACTUALLY tests |\n|---------------------------|---------------|---------------|-------------------------------|\n| LLMs extract arguments | Yes (Claimify, F1 scores) | Script + 10 texts (days) | N/A — this is settled |\n| Users tolerate correction | Analogous evidence exists | Wizard of Oz (n=20) | **Whether the correction loop is intrinsically tolerable as a self-service workflow** — Wizard of Oz lies about this because operator patches holes |\n| Maps help decisions | Decades of evidence | Manual maps shared | N/A — this is settled |\n| Retention | Genuine unknown | Manual service tracking | **Whether people return when THEY do the work** — concierge inflates retention |\n| \"Aha moment\" exists | Unknown | Figma prototype | **Time-to-value in real environment**: does paste→draft→edit reach satisfying state in 5-10 min for first-time user? |\n\n**The real unknowns**: workflow tolerability (#2) and time-to-value (#5). These cannot be reliably tested without building. Wizard of Oz systematically over-estimates both.\n\n## The Publishing-First Trap\n\n**Upsides**: validates framing, seeds credibility for grants, attracts niche community.\n\n**Trap**: \"Resonance is not demand. Posts get agreement from people who will never change their behavior. Publishing rewards *coherence*, not *adhesion*.\" For a 15-year researcher, publishing becomes productive procrastination because the ideas are genuinely good and socially rewarded.\n\n**Rule**: Publishing is only a net win if constrained: ONE flagship essay, strictly time-boxed, CTA = \"join waitlist to test prototype.\" No essay series until prototype exists.\n\n## The Recommended Hybrid Path\n\n| Week | Action |\n|------|--------|\n| 0-1 | Run claim extraction experiment (validates technical core, already planned) |\n| 1-8 | Build thin-slice prototype (ugly UI OK) with ONE correction action |\n| Parallel (strict cap) | Publish ONE essay whose sole CTA is \"test the prototype\" — no multi-part series |\n| 4-8 | Wizard of Oz ONLY to design correction UX, not \"validate demand\" |\n| 8 | Real usage tests with 10-20 people. Measure: did they come back without being asked? |\n\n## The Diagnostic Question\n\n**What is the single measurable behavior you want within 8 weeks?**\n\nThis answer determines everything downstream. Examples:\n- \"User completes first map + makes 5 corrections in under 10 minutes\"\n- \"3 users import a second URL within 14 days\"\n\n## Key Quotes\n\n> \"The evidence you get [from pre-product testing] will skew toward optimism on the exact dimension Deliberus is most likely to fail on: sustained use of a cognitively demanding workflow.\"\n\n> \"A working prototype is a shareable object that travels. Essays travel too, but they attract 'people who like ideas.' Prototypes attract 'people who want a tool now.'\"\n\n> \"You don't need a 4-6 month MVP before learning. You need a 6-8 week thin slice that is real, and weekly demos as an execution forcing function.\"\n\n## Source\nGPT-5.2 via PAL MCP, March 28, 2026. Full analysis with file references to consensus-path-forward.md, steelmanned-critiques.md, single-player-utility.md, vision.md.\n"}