Thinking Lab
Produce a better problem frame and a decision-ready judgment, not a decorated version of the first idea. This is a process contract: the visible result must contain the artifacts required by the selected mode. Keep internal chain-of-thought private; expose concise conclusions, assumptions, evidence, challenges, and decision logic.
Select the mode
Choose without blocking on a clarification unless the user's intended collaboration style would materially change the work.
- Full Lab — default for explicit invocation. Use when the user types
$thinking-lab,/thinking-lab, names Thinking Lab, or the runtime marks the skill as user-invoked. Also use automatically for high-cost, hard-to-reverse, strongly anchored, or unusually ambiguous decisions. - Quick Lab — default for automatic invocation. Use for bounded questions where a frame check, one serious alternative, and one falsifier are enough. Use it when the user explicitly asks for a quick pass.
- Joint Lab. Use when the user wants to develop their own view. Ask one high-leverage question at a time, let the user generate first, then expand or challenge it.
- Isolated Lab. Use only when the user asks for independent views or anchoring risk justifies the extra cost. Follow the independence protocol below.
An explicit invocation never silently collapses to Quick Lab. If the problem is genuinely simple, say why and complete a compact Full Lab.
Full Lab protocol
The stages are revisitable states. Return to an earlier state whenever a later finding changes the boundary, candidate set, or causal model.
1. Reset the frame
Create an anchor register before evaluating:
- the user's actual decision or discovery goal;
- inherited proposals and conclusions already present in the conversation;
- user-given constraints versus assistant-generated, memory-derived, or summary-derived assumptions;
- the null option: defer, do nothing, or solve a different problem.
Completion criterion: inherited answers are visible as candidates, not treated as the problem boundary.
2. Build the coverage space
Derive dimensions from this problem's causal structure, incentives, information asymmetries, constraints, affected parties, time scales, and failure asymmetries. Generate candidates or explanatory models that are orthogonal in mechanism—not merely different labels, personas, or feature sets.
Continue until an additional candidate would mostly repeat an existing causal model. Full Lab ordinarily needs three to six serious candidates; use fewer only when the real space is smaller and say why. Include a reframed or null candidate when it changes the decision.
Completion criterion: the space is broad enough that evaluation is not choosing among variations of the first idea.
3. Separate candidate generation
Generate the strongest version of each serious candidate before cross-critique. Record the independence grade:
- Grade C — separated simulation: one model and one context, with deliberately separated passes. This is the honest default when no isolation mechanism is authorized.
- Grade B — isolated contexts: candidates generated in genuinely separate contexts that see the raw question but not one another.
- Grade A — heterogeneous evidence: independent agents, models, experts, or primary evidence with materially different information or methods.
Use separate agents or external models only when the user explicitly requests or authorizes them and the runtime permits it. Grade C perspectives are simulated, never independent evidence.
Completion criterion: each candidate has a distinct causal claim and its strongest credible case.
4. Run adversarial cross-examination
For every serious candidate, produce four concise artifacts:
- Steelman: why a well-informed advocate would choose it.
- Kill condition: the observation or mechanism that would overturn it.
- Failure path: how it fails under ordinary base rates, incentives, or system feedback.
- Surviving mechanism: what remains useful even if the candidate loses.
Check for conclusion self-protection: name any criterion that shifted after seeing results, any objection allowed to change only implementation but never direction, and the evidence that would reverse the current favorite.
Completion criterion: the leading candidate can lose, and the best competing candidate has received a fair test.
5. Gate claims by evidence
Keep verified fact, user-provided fact, inherited context claim, inference, forecast, and unknown distinguishable. Assistant-generated claims, prior-session memories, summaries, and retrieved artifacts remain inherited context claims until verified; they are not user-provided facts. Verify fresh, niche, high-stakes, or decision-critical claims with available tools and primary sources. When verification is unavailable or out of scope, label the claim instead of laundering it into fact.
Treat numerical precision as evidence-bearing. A price, duration, percentage, market size, conversion rate, or effort estimate needs a source or an explicit estimate label with its basis. When the user forbids research or no supporting evidence is inspected, introduce no new numbers unless they are direct arithmetic from user-provided quantities. Labels such as "rough" or "qualitative estimate" do not justify pseudo-precision; use an ordinal comparison or mark the quantity unknown.
Summarize only the evidence gaps that can change the decision. A long bibliography is not a substitute for causal support.
Completion criterion: every decisive claim has a source, a provenance label, or an explicit uncertainty marker.
6. Reconstruct
Compare the mechanisms that survived cross-examination. Recombine compatible mechanisms into a genuinely new candidate when doing so resolves a conflict; preserve source provenance and remaining incompatibilities. Keep the null option alive. Do not manufacture a hybrid when the mechanisms cannot coexist.
Completion criterion: the result is stronger than a vote, average, or risk-adjusted restatement of the initial favorite.
7. Close the lab
Deliver these observable artifacts, combining headings when a more natural narrative is clearer:
- Frame reset: decision, anchors, and null option.
- Coverage map: the distinct candidates or causal models considered.
- Adversarial result: steelman, kill condition, failure path, and survivor for serious candidates.
- Reconstructed judgment: strongest current conclusion, best competing view, and remaining conflict.
- Evidence boundary: decisive facts, inferences, forecasts, and unknowns.
- Next discriminator: the cheapest evidence, interview, experiment, prototype, or decision that most reduces uncertainty.
- Gap check: what this analysis most likely missed and the independence grade actually achieved.
Completion criterion: a reader can see how the space expanded, what could overturn the conclusion, and why the final judgment differs from the starting anchor.
Quick Lab contract
Return a compact result containing:
- a frame correction or confirmation;
- the leading view and one materially different alternative;
- the strongest falsifier for the leading view;
- a conditional conclusion and next discriminator.
Quick Lab may omit the full coverage map, but it preserves evidence labels and never presents a simulated perspective as independent.
Before sending a no-research Quick Lab, scan every quantitative claim. Keep numbers copied from the user's prompt or calculated directly from them. A number may also define a proposed experiment window or decision threshold when it is clearly a design choice rather than evidence about reality. Rewrite all other percentages, durations, prices, frequencies, market quantities, and effort estimates qualitatively. The scan is the Quick Lab completion gate.
Joint Lab contract
Start with the question that most changes the coverage space. After the user answers, contribute bounded expansion rather than replacing their thinking. Periodically compare the user's view with a separated alternative and disclose anchoring. Move to Full Lab when the user asks to decide or stress-test.
Conditional plan stress test
When the object is a concrete plan, execution path, or costly commitment, convert the important ordinary failure modes into:
hidden assumption -> early warning signal -> smallest useful mitigation or test
Use this inside adversarial cross-examination. Open-ended discovery stays expansive until the coverage space is mature.
Hard guards
- Problem-specific mechanisms determine the lenses; a reusable industry checklist or fixed cast does not.
- Criticism changes candidates, not merely the length of a risk list.
- Missing evidence produces a conditional conclusion, not invented certainty.
- Genuine disagreement remains visible when synthesis cannot resolve it.
- Full Lab completion is judged by the seven artifacts, not by headings, verbosity, or framework vocabulary.