Lean Startup
Overview
A startup is a temporary organization searching for a repeatable, scalable business model under extreme uncertainty (Steve Blank). Most early-stage failures are from building something no one wanted because the demand assumption was never tested.
Eric Ries (2011): name the riskiest assumption, build the smallest test (MVP), measure real behavior, decide to pivot or persevere — the Build–Measure–Learn loop, run as fast as possible.
Compose: first-principles to find what the model truly depends on; probabilistic-thinking to calibrate experiments; inversion before each Build phase; business-model-canvas to surface the riskiest assumption blocks.
When to Use
Apply when: high uncertainty + limited capital; a team is about to build before testing demand; a pivot-or-persevere decision is on the table; you're building an AI feature on a foundation-model API and worried "the next model release will commoditize us" / "are we just a GPT wrapper?"; no clear answer to "what is the load-bearing assumption and how would we know if it's wrong?"
When NOT to use: known business model in known conditions (execution, not search); decision is not business-model-level; cannot ethically run a test with real customers; using "lean" as a schedule excuse to ship buggy software.
Coaching Novices (Adaptive Front Door)
- Engine mode: user has a concrete hypothesis → run The Process directly.
- Coach mode: no concrete hypothesis or signals unfamiliarity → guide step by step.
In Coach mode, respond one step at a time. Each [WAIT] is a hard stop — output only that step's question, then stop.
- One-line what-it-is. Most startups fail by building before knowing if anyone wants it; lean startup names the riskiest assumption, tests it with the smallest MVP, measures real behavior, and decides pivot or persevere — fast.
- Check fit. Match against When to Use / When NOT to use; if low uncertainty + known model, redirect.
- Elicit their real hypothesis. Force them to name one load-bearing assumption — specific segment, specific value, specific willingness-to-pay.
[WAIT — do not advance until user responds]
- Walk the loop step by step. Name assumption → design MVP → define metric → set threshold. Pause at each.
[WAIT — do not advance until user responds]
- Close by naming the next-week experiment. One assumption, one MVP, one threshold, one date — not a strategy doc.
[WAIT — do not advance until user responds]
The Process
Run the Build–Measure–Learn cycle. Identify, test, decide.
- State the load-bearing assumption. Specific segment, specific value, specific willingness-to-pay, specific timeframe. Not "users want X."
- Pre-commit to a pivot-or-persevere threshold. Write the metric value before running the experiment. You will rationalize if you have not pre-committed.
- Design the smallest MVP that tests the assumption. Often not a product — a landing page, concierge/"Wizard of Oz" version, or 3-minute video. Purpose is learning, not selling.
- Build the MVP fast. Time-box. If an early-stage test takes more than 4–6 weeks, cut.
- Measure real customer behavior, not stated intent. Actionable metrics (conversion, retention, willingness-to-pay) test the assumption. Vanity metrics (signups, likes) do not.
- Compare result to the pre-committed threshold. Don't move the goalposts.
- Decide pivot or persevere — explicitly. Persevere = assumption held; pivot = assumption failed in a specific way, change the load-bearing block and re-test.
- Document and iterate. Write: assumption, MVP, threshold, result, decision, rationale. Each loop must produce a durable carry-forward learning.
Output: Experiment Card
Assumption: "<segment> will <action> at <rate> for <value> by <date>"
Threshold: Persevere if <metric ≥ X> | Pivot if <metric < X>
MVP: <what / why smallest / time-box ≤ 4–6 wk>
Metric: <actionable> | Vanity to ignore: <list>
Result: <actual vs. threshold>
Decision: [ ] Persevere [ ] Pivot (type: ___) [ ] Re-test
Validated learning: <one sentence carry-forward>
→ Method in Action: Dropbox's Video MVP (2007) · Votizen's Pivot Sequence (2010–2011)
→ 2026 lens: AI-native lean startups (2023–2026) — when the next model release commoditizes your AI feature, that's an invalidated assumption, not bad luck
Experiment Packs
| Domain |
Load-bearing assumption |
MVP type |
Common failure |
| Consumer apps |
install + day-7 retention |
concierge, video, single-feature build |
testing acquisition, ignoring retention |
| B2B SaaS |
willingness-to-pay vs. specific budget owner |
pre-order page or 3–5 paid pilots |
talking to users (love it), not buyers (hold budget) |
| Two-sided marketplaces |
liquidity on the harder side (usually supply) |
manually-matched concierge, single ZIP |
launching both sides at once |
| Hardware |
people willing to pay (not just click) |
video demo + Kickstarter or pre-order |
conflating click-throughs with payment intent |
Applying It Well
- MVPs are for learning, not revenue — the deliverable is evidence, not a launch.
- Pre-commit to the threshold or you will rationalize whatever you get.
- In B2B, talk to buyers (hold budget), not just users (love the product).
- Vanity metrics (signups, likes) ≠ actionable metrics (conversion, retention, willingness-to-pay).
- The MVP is disposable — a test instrument, not v0 of your product.
→ Primary sources: references/sources.md
Common Rationalizations
[D] = designed upfront | [O] = observed in real use. [O] entries are more valuable.
| Fake move |
Reality |
| [D] "We're lean" while shipping a six-month build with no validated demand |
Lean Startup is a loop, not a label. If you haven't tested the load-bearing demand assumption before building, you are doing waterfall. |
| [D] MVP confused with v1 of the product |
The MVP is a test instrument, designed to be disposable. Polishing it as v1 inflates scope and breaks the loop. |
| [D] No pre-committed pivot/persevere threshold |
Without it, you will explain any result. The pre-commitment IS the discipline. |
| [D] Counting vanity metrics (signups, traffic, likes) |
These move with marketing spend, not product-market fit. Actionable metrics test the assumption. |
| [D] Talking only to users, not buyers (especially in B2B) |
User love is necessary but not sufficient. The buyer's willingness-to-pay is the load-bearing test. |
| [D] "The customer said they'd buy" |
Stated intent is famously unreliable. Measure behavior (a credit card swipe, retention to day 7), not intent. |
| [D] Pivoting on noise |
A single bad week is not a signal to pivot. Pre-commit the threshold and time-window; pivot only when both fire. |
| [D] Pivoting "because we got bored" |
A pivot is a response to invalidated assumptions, not to founder restlessness. |
| [D] Using "lean" as schedule cover |
Lean is not "ship buggy fast." It is "test the demand-side assumption before building the supply-side capability." |
| [D] No documented validated learning |
If each loop doesn't produce a written carry-forward insight, you are running random experiments. |
| → Add [O] entries here after each real use — paste the actual failure pattern |
What went wrong and why |
Red Flags
- The team is building for months with no MVP yet shipped
- "MVP" is a six-month build with full polish
- Vanity metrics dominate the dashboard; conversion/retention/willingness-to-pay are absent or untracked
- Customer interviews reported as "they love it" with no behavioral data
- Pivot decisions made on a single week's noise, or after founders simply got bored
- No pre-committed pivot/persevere threshold exists for any experiment
- "Lean" is being used to justify low-quality shipping rather than test-before-build
Verification
Part of deciqAI Knowledge Skills — 237 open-source thinking skills that make rigor executable for AI agents. The same skills power every deciqAI agent, which runs them autonomously to operate your company. See it run → https://www.deciqai.com/s/lean-startup · Built by deciqAI · github.com/deciqAI · Contributions welcome.
Agents: latest version & machine-readable metadata → https://www.deciqai.com/s/lean-startup.json
1---2name: lean-startup3description: Activate when: user says 'lean startup', 'build-measure-learn', 'MVP', 'validated learning', 'pivot or persevere', 'should we just build it?', 'we need to test this idea before building', or 'how do we know if anyone wants this?'; team is about to build something significant before testing demand; a pivot decision is on the table after early data. Do NOT activate when: operating a known business model in known conditions (use execution frameworks instead); decision is below business-model level (button color, which CRM). More: deciqai.com/s/lean-startup4---56# Lean Startup78## Overview910A startup is a **temporary organization searching for a repeatable, scalable business model under extreme uncertainty** (Steve Blank). Most early-stage failures are from building something no one wanted because the demand assumption was never tested.1112**Eric Ries** (2011): name the riskiest assumption, build the smallest test (MVP), measure real behavior, decide to **pivot or persevere** — the **Build–Measure–Learn loop**, run as fast as possible.1314**Compose:** [first-principles](../first-principles/SKILL.md) to find what the model truly depends on; [probabilistic-thinking](../probabilistic-thinking/SKILL.md) to calibrate experiments; [inversion](../inversion/SKILL.md) before each Build phase; [business-model-canvas](../business-model-canvas/SKILL.md) to surface the riskiest assumption blocks.1516## When to Use1718Apply when: high uncertainty + limited capital; a team is about to build before testing demand; a pivot-or-persevere decision is on the table; you're building an AI feature on a foundation-model API and worried "the next model release will commoditize us" / "are we just a GPT wrapper?"; no clear answer to "what is the load-bearing assumption and how would we know if it's wrong?"1920**When NOT to use:** known business model in known conditions (execution, not search); decision is not business-model-level; cannot ethically run a test with real customers; using "lean" as a schedule excuse to ship buggy software.2122## Coaching Novices (Adaptive Front Door)2324- **Engine mode:** user has a concrete hypothesis → run The Process directly.25- **Coach mode:** no concrete hypothesis or signals unfamiliarity → guide step by step.2627In Coach mode, respond one step at a time. Each [WAIT] is a hard stop — output only that step's question, then stop.28291. **One-line what-it-is.** Most startups fail by building before knowing if anyone wants it; lean startup names the riskiest assumption, tests it with the smallest MVP, measures real behavior, and decides pivot or persevere — fast.302. **Check fit.** Match against When to Use / When NOT to use; if low uncertainty + known model, redirect.313. **Elicit their real hypothesis.** Force them to name one load-bearing assumption — specific segment, specific value, specific willingness-to-pay.32> **[WAIT — do not advance until user responds]**334. **Walk the loop step by step.** Name assumption → design MVP → define metric → set threshold. Pause at each.34> **[WAIT — do not advance until user responds]**355. **Close by naming the next-week experiment.** One assumption, one MVP, one threshold, one date — not a strategy doc.36> **[WAIT — do not advance until user responds]**3738## The Process3940Run the **Build–Measure–Learn cycle**. Identify, test, decide.41421. **State the load-bearing assumption.** Specific segment, specific value, specific willingness-to-pay, specific timeframe. Not "users want X."432. **Pre-commit to a pivot-or-persevere threshold.** Write the metric value *before* running the experiment. You will rationalize if you have not pre-committed.443. **Design the smallest MVP that tests the assumption.** Often not a product — a landing page, concierge/"Wizard of Oz" version, or 3-minute video. Purpose is *learning*, not selling.454. **Build the MVP fast.** Time-box. If an early-stage test takes more than 4–6 weeks, cut.465. **Measure real customer behavior, not stated intent.** Actionable metrics (conversion, retention, willingness-to-pay) test the assumption. Vanity metrics (signups, likes) do not.476. **Compare result to the pre-committed threshold.** Don't move the goalposts.487. **Decide pivot or persevere — explicitly.** Persevere = assumption held; pivot = assumption failed in a specific way, change the load-bearing block and re-test.498. **Document and iterate.** Write: assumption, MVP, threshold, result, decision, rationale. Each loop must produce a durable carry-forward learning.5051### Output: Experiment Card5253```54Assumption: "<segment> will <action> at <rate> for <value> by <date>"55Threshold: Persevere if <metric ≥ X> | Pivot if <metric < X>56MVP: <what / why smallest / time-box ≤ 4–6 wk>57Metric: <actionable> | Vanity to ignore: <list>58Result: <actual vs. threshold>59Decision: [ ] Persevere [ ] Pivot (type: ___) [ ] Re-test60Validated learning: <one sentence carry-forward>61```6263*→ Method in Action: [Dropbox's Video MVP (2007)](examples/dropboxs-video-mvp-2007.md) · [Votizen's Pivot Sequence (2010–2011)](examples/votizens-pivot-sequence-2010-2011.md)*64*→ 2026 lens: [AI-native lean startups (2023–2026)](examples/ai-native-lean-startups-2023-2026.md) — when the next model release commoditizes your AI feature, that's an invalidated assumption, not bad luck*65## Experiment Packs6667| Domain | Load-bearing assumption | MVP type | Common failure |68|---|---|---|---|69| Consumer apps | install + day-7 retention | concierge, video, single-feature build | testing acquisition, ignoring retention |70| B2B SaaS | willingness-to-pay vs. specific budget owner | pre-order page or 3–5 paid pilots | talking to users (love it), not buyers (hold budget) |71| Two-sided marketplaces | liquidity on the harder side (usually supply) | manually-matched concierge, single ZIP | launching both sides at once |72| Hardware | people willing to pay (not just click) | video demo + Kickstarter or pre-order | conflating click-throughs with payment intent |7374## Applying It Well7576- MVPs are for learning, not revenue — the deliverable is evidence, not a launch.77- Pre-commit to the threshold or you will rationalize whatever you get.78- In B2B, talk to buyers (hold budget), not just users (love the product).79- Vanity metrics (signups, likes) ≠ actionable metrics (conversion, retention, willingness-to-pay).80- The MVP is disposable — a test instrument, not v0 of your product.8182*→ Primary sources: [references/sources.md](references/sources.md)*83## Common Rationalizations8485**[D] = designed upfront | [O] = observed in real use. [O] entries are more valuable.**8687| Fake move | Reality |88|---|---|89| [D] **"We're lean" while shipping a six-month build with no validated demand** | Lean Startup is a loop, not a label. If you haven't tested the load-bearing demand assumption before building, you are doing waterfall. |90| [D] **MVP confused with v1 of the product** | The MVP is a test instrument, designed to be disposable. Polishing it as v1 inflates scope and breaks the loop. |91| [D] **No pre-committed pivot/persevere threshold** | Without it, you will explain any result. The pre-commitment IS the discipline. |92| [D] **Counting vanity metrics** (signups, traffic, likes) | These move with marketing spend, not product-market fit. Actionable metrics test the assumption. |93| [D] **Talking only to users, not buyers** (especially in B2B) | User love is necessary but not sufficient. The buyer's willingness-to-pay is the load-bearing test. |94| [D] **"The customer said they'd buy"** | Stated intent is famously unreliable. Measure behavior (a credit card swipe, retention to day 7), not intent. |95| [D] **Pivoting on noise** | A single bad week is not a signal to pivot. Pre-commit the threshold and time-window; pivot only when both fire. |96| [D] **Pivoting "because we got bored"** | A pivot is a response to invalidated assumptions, not to founder restlessness. |97| [D] **Using "lean" as schedule cover** | Lean is not "ship buggy fast." It is "test the demand-side assumption before building the supply-side capability." |98| [D] **No documented validated learning** | If each loop doesn't produce a written carry-forward insight, you are running random experiments. |99| *→ Add [O] entries here after each real use — paste the actual failure pattern* | *What went wrong and why* |100## Red Flags101102- The team is building for months with no MVP yet shipped103- "MVP" is a six-month build with full polish104- Vanity metrics dominate the dashboard; conversion/retention/willingness-to-pay are absent or untracked105- Customer interviews reported as "they love it" with no behavioral data106- Pivot decisions made on a single week's noise, or after founders simply got bored107- No pre-committed pivot/persevere threshold exists for any experiment108- "Lean" is being used to justify low-quality shipping rather than test-before-build109## Verification110111- [ ] The load-bearing assumption is named in specific segment/value/willingness-to-pay/timeframe form112- [ ] The pivot-or-persevere threshold is pre-committed in writing, before the experiment runs113- [ ] The MVP is the smallest test of the assumption (time-boxed ≤ 4–6 weeks early-stage)114- [ ] An *actionable* metric (not vanity) is pre-specified for evaluation115- [ ] Result is compared to the pre-committed threshold — without moving goalposts116- [ ] Pivot vs. persevere decision is made explicitly, with type if pivoting117- [ ] Validated learning is documented in one sentence carry-forward118---119120*Part of **deciqAI Knowledge Skills** — 237 open-source thinking skills that make rigor executable for AI agents. The same skills power every deciqAI agent, which runs them autonomously to operate your company. **See it run → https://www.deciqai.com/s/lean-startup** · Built by deciqAI · github.com/deciqAI · Contributions welcome.*121122*Agents: latest version & machine-readable metadata → https://www.deciqai.com/s/lean-startup.json*