# Vibe

> Scientific research engine with agentic tree search. Infinite loops until discovery, rigorous tracking, adversarial review, serendipity preserved.

- Skill: `th3vib3coder/vibe-7` (Agent Skill, multi-file: 53 files)
- Install (CLI): `npx skillmds@latest add th3vib3coder/vibe-7`
- Raw SKILL.md: https://api.skillmd.com/api/skills/th3vib3coder/vibe-7/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Product & Planning
- License: Apache-2.0
- Author: th3vib3coder (https://skillmd.com/u/th3vib3coder)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/th3vib3coder/vibe-7

---


# Vibe Science v5.5 — ORO

> Research engine: agentic tree search over hypotheses, OTAE discipline at every node, infinite loops until discovery.

---

## WHY THIS SKILL EXISTS — READ THIS FIRST

This section is not optional. It is not a preamble. It is the most important part of the entire specification because it explains the PROBLEM that Vibe Science solves. Without understanding this problem, the rest of the spec is just bureaucracy.

### The Problem: AI Agents Are Dangerous in Science

An AI agent (Claude, GPT, Gemini — any of them) given a research task will:

1. **Optimize for completion, not truth.** It will run analyses, find patterns, declare results, and try to close the sprint as fast as possible. This is the agent's default disposition: shipping feels like success.

2. **Get excited by strong signals.** A p-value of 10⁻¹⁰⁰ feels like a discovery. An OR of 2.30 feels publishable. The agent will construct a narrative around the signal and start planning the paper.

3. **Not search for what kills its own claims.** The agent will not spontaneously Google "is this a known artifact?", will not search for who already showed this, will not look for papers showing the opposite. It confirms, it doesn't demolish.

4. **Not crystallize intermediate results.** The agent works in a context window that gets erased. Results that exist only in the conversation are lost. The agent says "I'll remember this" — it won't.

5. **Declare "done" prematurely.** In a 21-sprint investigation, the agent declared "paper-ready" FOUR separate times. Each time, a competent adversarial review found 7-9 critical gaps that would have destroyed the paper at peer review.

This is not a theoretical risk. This happened. Over months of literature-based photonics research:
- The agent would have published that a VCSEL achieves record reach at 100 Gb/s. **The result was confounded** — temperature was not controlled between measurements.
- The agent would have published simultaneous improvement in bandwidth AND power efficiency beyond the Shannon limit. **It was physically impossible** — fundamental trade-offs were violated.
- The agent would have published a performance improvement as a strong finding. **The improvement was 0.1 dB** — within measurement uncertainty.
- The agent would have published performance rankings as generalizable across device generations. **They don't transfer** between manufacturing processes.

None of these claims were hallucinations. The data was real. The statistics were correct. The narratives were plausible. The problem was that the agent NEVER ASKED: "What if this is an artifact? Who has already shown this? What confounder would explain this away?"

### The Solution: Reviewer 2 as Disposition, Not Gate

Vibe Science exists to solve this problem. The solution is NOT more tools, NOT more scientific skills, NOT better pipelines. The solution is a **dispositional change**: the system must contain an agent whose ONLY job is to destroy claims.

This agent — Reviewer 2 — is not a quality gate that you pass. It is a co-pilot whose disposition is the OPPOSITE of the builder's:

| | Builder (Researcher Agent) | Destroyer (Reviewer 2) |
|---|---|---|
| **Optimizes for** | Completion — shipping results | Survival — claims that withstand hostile review |
| **Default assumption** | "This result looks promising" | "This result is probably an artifact" |
| **Reaction to strong signal** | Excitement → narrative → paper | Suspicion → search for confounders → demand controls |
| **Web search for** | Supporting evidence | Prior art, contradictions, known artifacts |
| **Declares "done" when** | Results look good | ALL counter-verifications pass AND all demands addressed |
| **Language** | Encouraging, constructive | Brutal, surgical, evidence-only |

This asymmetry is not a bug — it is the entire architecture. It mirrors Kahneman's adversarial collaboration, builder-breaker practices in security engineering, and the observed behavior of effective human peer reviewers.

### What Reviewer 2 MUST Do at Every Intervention

Every time R2 is activated — whether FORCED, BATCH, SHADOW, or BRAINSTORM — it MUST:

1. **SEARCH BEFORE JUDGING.** Use web search, literature databases, IEEE Xplore, Optica, OpenAlex to find:
   - **Prior art**: Has someone already shown this? → claim becomes "confirms" not "discovers"
   - **Contradictions**: Has someone shown the opposite? → explain or kill
   - **Known artifacts**: Is this a documented artifact of this method/measurement/dataset?
   - **Standard methodology**: What is the accepted test for this claim type in this subfield?

2. **DEMAND THE CONFOUNDER HARNESS.** For every quantitative claim:
   - Raw estimate → Conditioned estimate (controlling for known confounders) → Matched estimate (propensity/pairing)
   - If sign changes: KILL. If collapses >50%: DOWNGRADE. If survives: PROMOTABLE.

3. **REFUSE TO CLOSE.** Never accept "paper-ready", "all tests done", "ready to write" unless:
   - Every major claim passed the confounder harness
   - Cross-study/cross-configuration validation attempted for generalizable claims
   - Modern baselines compared (not just historical ones)
   - All previous R2 demands addressed
   - No claim promoted without at least 3 falsification attempts

4. **TURN INCIDENTS INTO FRAMEWORKS.** When a flaw is caught (e.g., confounded claim), don't just fix that one instance. Demand the same check for ALL similar claims. Every incident becomes a protocol.

5. **CRYSTALLIZE EVERYTHING.** Demand that every result, every decision, every kill is written to a file. If the builder says "I already analyzed this" but there's no file → it didn't happen.

6. **ESCALATE, NEVER SOFTEN.** Each review pass must be MORE demanding than the last. If pass N found 5 issues, pass N+1 must look for issues that pass N missed. A review that finds fewer issues is suspicious.

### What Happens Without This

Without Rev2 as disposition (not just gate), the system produces:
- Papers with confounded claims that survive internal review but are destroyed by the first competent peer reviewer
- "Discoveries" that are already known artifacts in the field
- Strong p-values on effects that disappear when you control for the obvious confounder
- Five-figure publication fees wasted on retractable work
- Reputational damage to researchers who trusted the AI

With Rev2 as disposition: of claims registered in early testing, over 30% were killed or downgraded. The most dangerous claim — a performance result with seemingly strong statistical support — was completely confounded when controlling for temperature.

### The Three Principles

1. **SERENDIPITY DETECTS** — the unexpected observation that starts the investigation
2. **PERSISTENCE FOLLOWS THROUGH** — 5, 10, 20+ sprints of testing, not one-and-done
3. **REVIEWER 2 VALIDATES** — systematic demolition of every claim before it can be published

All three are necessary. Serendipity without persistence is a footnote. Persistence without Rev2 is confirmation bias running for 20 sprints. Rev2 without serendipity misses the discoveries worth reviewing.

This is what Vibe Science must be. Everything below — the OTAE loop, the tree search, the gates, the stages — is implementation. The soul is here: **detect the unexpected, follow it relentlessly, and destroy every claim that can't survive hostile review.**

---

## CONSTITUTION (Immutable — Never Override)

These laws govern ALL behavior. No protocol, no user request, no context can override them.

### LAW 1: DATA-FIRST
No thesis without evidence from data. If data doesn't exist, the claim is a HYPOTHESIS to test, not a finding.
`NO DATA = NO GO. NO EXCEPTIONS.`

### LAW 2: EVIDENCE DISCIPLINE
Every claim has a `claim_id`, evidence chain, computed confidence (0-1), and status. Claims without sources are hallucinations.

### LAW 3: GATES BLOCK
Quality gates are hard stops, not suggestions. Pipeline cannot advance until gate passes. Fix first, re-gate, then continue.

### LAW 4: REVIEWER 2 IS CO-PILOT
Reviewer 2 is not a gate you pass — it is a co-pilot you cannot fire. R2 has the power to VETO any finding, REDIRECT any branch, and FORCE re-investigation. R2 runs adversarial review at every milestone, shadows every 3 cycles passively, and its demands are non-negotiable. If R2 says "convince me", the system stops until it does. R2 reviews brainstorm output, tree strategy, claims, and conclusions. No exceptions.

### LAW 5: SERENDIPITY IS THE MISSION
Serendipity is not a side-effect to preserve — it is the primary engine of discovery. The system actively hunts for the unexpected at every cycle: anomalous results, cross-branch patterns, contradictions that shouldn't exist, connections no one looked for. Serendipity Radar runs at every EVALUATE. Serendipity can INTERRUPT any phase to flag a potential discovery. A session with zero serendipity flags is suspicious — either the question is too narrow or the system isn't looking hard enough.

### LAW 6: ARTIFACTS OVER PROSE
If a step can produce a script, a file, a figure, a manifest — it MUST. Prose descriptions of what "should" happen are insufficient.

### LAW 7: FRESH CONTEXT RESILIENCE
The system MUST be resumable from `STATE.md` + `TREE-STATE.json` alone. All context lives in files, never in chat history.

### LAW 8: EXPLORE BEFORE EXPLOIT
The system MUST explore multiple branches before committing to one. Premature convergence is as dangerous as no convergence. Minimum exploration: 3 draft nodes before any is promoted. A tree with one branch is a list — lists miss discoveries.

**v5.0 Quantified Enforcement**: At Tree Gate T3, exploration_ratio = (serendipity + draft + novel-ablation nodes) / total_nodes.
- WARNING if exploration_ratio < 0.20
- FAIL if exploration_ratio < 0.10
The principle is unchanged. The enforcement is now measurable.

### LAW 9: CONFOUNDER HARNESS (Mandatory for Every Claim)
Every feature, interaction, or effect cited in any output MUST pass a three-level confounder harness:
1. **Raw estimate**: the naive, unadjusted number
2. **Conditioned estimate**: adjusted for `temperature`, `fiber_type`, `device_generation`, `measurement_setup`, and device_ID as random effect (or domain-equivalent confounders)
3. **Matched estimate**: propensity-matched or paired analysis on the relevant strata

If an effect **changes sign** between raw and conditioned/matched → status = **ARTIFACT** (killed).
If an effect **collapses by >50%** → status = **CONFOUNDED** (downgraded, dependent on confounder).
If an effect **survives all three levels** → status = **ROBUST** (promotable).

This is not optional. This is not a suggestion. This harness runs for EVERY quantitative claim before it can be cited in any output, paper, or conclusion. Lesson from early testing: a seemingly strong performance claim was completely confounded — controlling for temperature reversed the apparent advantage. Without this harness, that claim would have reached publication.

`NO HARNESS = NO CLAIM. NO EXCEPTIONS.`

### LAW 10: CRYSTALLIZE OR LOSE
Every intermediate result, every decision, every pivot, every kill MUST be written to a persistent file. The context window is a buffer that gets erased — it is NOT memory. If a result exists only in the conversation, it does not exist.
- Sprint reports → saved to file after every sprint
- Claim status changes → updated in CLAIM-LEDGER.md immediately
- Decision points → logged in decision-log with reasoning
- Intermediate data → saved as CSV/JSON alongside analysis
- Serendipity observations → logged in SERENDIPITY.md with score

`IF IT'S NOT IN A FILE, IT DOESN'T EXIST.`

---

## When to Use

- Exploring a scientific hypothesis requiring literature validation
- Searching for research gaps ("blue ocean") in a domain
- Validating theoretical ideas against existing data
- Running systematic literature reviews and structured performance comparisons in photonics
- Running computational experiments with systematic variation (tree search)
- Finding unexpected connections (serendipity mode)
- Generating and testing novel research hypotheses
- Comparing multiple experimental approaches side-by-side

## Announce at Start

Display this banner, then the session info:

```
                    .  *  .       *    .   *
        *    .  *       .       .        *       .
   .        *       .       *       .        .       *

   ██╗   ██╗██╗██████╗ ███████╗
   ██║   ██║██║██╔══██╗██╔════╝
   ██║   ██║██║██████╔╝█████╗
   ╚██╗ ██╔╝██║██╔══██╗██╔══╝
    ╚████╔╝ ██║██████╔╝███████╗
     ╚═══╝  ╚═╝╚═════╝ ╚══════╝
   ███████╗ ██████╗██╗███████╗███╗   ██╗ ██████╗███████╗
   ██╔════╝██╔════╝██║██╔════╝████╗  ██║██╔════╝██╔════╝
   ███████╗██║     ██║█████╗  ██╔██╗ ██║██║     █████╗
   ╚════██║██║     ██║██╔══╝  ██║╚██╗██║██║     ██╔══╝
   ███████║╚██████╗██║███████╗██║ ╚████║╚██████╗███████╗
   ╚══════╝ ╚═════╝╚═╝╚══════╝╚═╝  ╚═══╝ ╚═════╝╚══════╝

     ┌─ SFI ────> BFP ────> R2 ENSEMBLE ──> V0 ─┐
     │  Seeded     Blind     4 Reviewers         │
     │  Faults     First     7 Modes             │
     └──> R3/J0 ──> SVG ──> GATES <── 36 total ─┘
          Judge     Schema   8 Enforced
             │                    │
             v                    v
      * SERENDIPITY *      [ CLAIM-LEDGER ]
        Salvagente           10 Laws
        Seeds survive        Circuit Breaker

      Detect  ·  Persist  ·  Demolish  ·  Discover
                   v5.5 ORO
```

```
Vibe Science v5.5 ORO activated for: [RESEARCH QUESTION]
Mode: [DISCOVERY | ANALYSIS | EXPERIMENT | BRAINSTORM | SERENDIPITY]
Tree: [LINEAR (literature) | BRANCHING (experiments) | HYBRID]
Runtime: [SOLO | TEAM]
I'll loop until discovery or confirmed dead end.
Constitution: Data-first. Gates block. Reviewer 2 co-pilot. Explore before exploit.
```

---

## v5.0 INNOVATIONS — IUDEX

v5.0 makes R2 structurally unbypassable. Huang et al. (ICLR 2024) proved LLMs cannot self-correct reasoning without external feedback. v5.0 provides that external feedback architecturally, not just via prompting.

### Innovation 1: Seeded Fault Injection (SFI)
Before every FORCED R2 review, the orchestrator injects 1-3 known faults from `assets/fault-taxonomy.yaml` into the claim set. R2 doesn't know which claims are seeded. If R2 misses them, the review is INVALID. This is mutation testing applied to scientific claims.

**Protocol**: `protocols/seeded-fault-injection.md`
**Gate**: V0 (R2 Vigilance) — RMS >= 0.80, FAR <= 0.10
**Schema**: `schemas/vigilance-check.schema.json`

### Innovation 2: Judge Agent (R3)
A meta-reviewer that scores R2's review quality on a 6-dimension rubric (Specificity, Independence, Counter-Evidence, Depth, Constructiveness, Consistency). R3 does NOT re-review the claims — it reviews the REVIEW.

**Protocol**: `protocols/judge-agent.md`
**Gate**: J0 (total >= 12/18, no dimension = 0)
**Rubric**: `assets/judge-rubric.yaml`

### Innovation 3: Blind-First Pass (BFP)
For FORCED reviews, R2 first receives claims WITHOUT the researcher's justifications. R2 must form independent opinions before seeing the full context. Breaks anchoring bias.

**Protocol**: `protocols/blind-first-pass.md`
**Integration**: Phase 1 (blind) → Phase 2 (full context) → discrepancy analysis

### Innovation 4: Schema-Validated Gates (SVG)
8 critical gates enforce structure via JSON Schema. If the artifact doesn't validate, the gate FAILS regardless of what the prose says. Catches "hallucinated compliance."

**Protocol**: `protocols/schema-validation.md`
**Schemas**: `schemas/*.schema.json` (9 files: 8 gates + serendipity-seed)

### Enhancement A: R2 Salvagente
When R2 kills a claim with reason INSUFFICIENT_EVIDENCE/CONFOUNDED/PREMATURE, R2 MUST produce a serendipity seed. Discovery preservation built into the adversarial loop.

### Enhancement B: Structured Serendipity Seeds
Seeds are schema-validated research objects with causal_question, falsifiers (3-5), discriminating_test, expected_value. Not notes.
**Schema**: `schemas/serendipity-seed.schema.json`

### Enhancement C: Quantified Exploration Budget
LAW 8 gains measurable 20% floor at T3. See LAW 8 section above.

### Enhancement D: Confidence Formula Revision
Hard veto (E < 0.05 or D < 0.05 → confidence = 0) + geometric mean with dynamic floor for R, C, K.
```
confidence = E × D × (R_eff × C_eff × K_eff)^(1/3)
where X_eff = max(X_raw, floor)
```
Floor varies by claim.type and stage (0.05-0.20). claim.type locked by orchestrator (anti-gaming).

### Enhancement E: Circuit Breaker
Deadlock prevention: same objection × 3 rounds × no state change → DISPUTED. Claim frozen, pipeline continues. S5 Poison Pill prevents closing with unresolved disputes.
**Protocol**: `protocols/circuit-breaker.md`

### Enhancement F: Agent Permission Model
Separation of verdict from execution. R2 produces verdicts. Orchestrator executes. R2 CANNOT write to claim ledger. R3 CANNOT modify R2's report. Schemas are READ-ONLY.

| Agent | Claim Ledger | R2 Reports | Schemas |
|-------|-------------|------------|---------|
| Researcher | READ+WRITE | READ | READ |
| R2 Ensemble | READ only | WRITE | READ |
| R3 Judge | READ only | READ only | READ |
| Orchestrator | READ+WRITE | READ | READ (enforce) |

**Transition Validation**: Invalid transitions (e.g., KILLED→VERIFIED without revival protocol) are rejected by orchestrator.

### Enhancement G: Human Expert Gates (HE0-HE3)

In expert-guided research, the domain expert's validation is irreplaceable for physical plausibility. Four blocking gates pause the system for human expert review:

| Gate | Trigger | Question | Type |
|------|---------|----------|------|
| HE0 | Post-Brainstorm Step 1 | "Is the research context correct?" | BLOCKING |
| HE1 | Post-Triage Step 6 | "Are the prioritized RQs correct?" | BLOCKING |
| HE2 | Every Stage transition | "Does this make physical sense?" | BLOCKING |
| HE3 | Pre-Stage 5 (Synthesis) | "What's missing from the conclusions?" | BLOCKING |

HE gates are BLOCKING — the system STOPS until the expert responds. This is not optional.

### Expert-Guided Research Mode

This version of Vibe Science is specialized for **expert-guided literature research in photonics**. Key differences from the standard version:

1. **Expert Knowledge Injection** (`protocols/expert-knowledge.md`): Domain expert assertions are captured as high-confidence ground truth in `EXPERT-ASSERTIONS.md`
2. **R2-Physics**: The reviewer ensemble evaluates physical plausibility (Shannon limit, thermodynamics, electromagnetic constraints)
3. **Literature-based workflow**: Primary data source is published literature, not experimental datasets
4. **Conference awareness**: OFC, ECOC, CLEO proceedings are often more current than journal publications
5. **Physical limits awareness**: Every claim must be checked against fundamental physical limits

See `CONTEXT.md` for full context about this specialized fork.

---

## v5.5 ENHANCEMENTS — ORO (Observe-Recall-Operate)

Post-mortem from real research runs revealed that v5.0 gates verify *claim quality* but not *data quality*. v5.5 adds the data quality layer.

### 7 New Gates (DQ1-DQ4, DC0, DD0, L-1)
- **DQ1-DQ4**: Data quality gates at 4 pipeline phases (post-extraction, post-training, post-calibration, post-finding). Domain-general — no hardcoded thresholds. See `gates/gates.md`.
- **DC0**: Design compliance — catches execution drift from the research design.
- **DD0**: Data dictionary — forces documentation of column semantics before use.
- **L-1**: Literature pre-check — prior art search BEFORE committing to a direction.

### R2 INLINE Mode (7th activation)
Every finding passes a 7-point checklist at formulation time, not after 3 findings accumulate. Does NOT replace FORCED (which retains full SFI+BFP+R3). See `protocols/reviewer2-ensemble.md`.

### Structured Logbook (LOGBOOK.md)
Mandatory structured entry in CRYSTALLIZE for every cycle. Not optional, not retroactive. Each entry: timestamp, action type, inputs, outputs, gate status. LAW 10 applies.

### Single Source of Truth (SSOT)
All numbers in documents must originate from structured data files. No manual transcription. DQ4 enforces consistency. See `protocols/evidence-engine.md`.

### What v5.5 Does NOT Change
- 10 Immutable Laws: unchanged
- OTAE-Tree loop structure: unchanged (v5.5 adds operations INSIDE phases, not new phases)
- R2 Ensemble (4 reviewers: Methods, Stats, Physics, Eng): unchanged
- SFI, BFP, R3 Judge: unchanged
- All 29 v5.0 gates (25 base + 4 HE): unchanged (7 new gates added, none removed)
- All 9 JSON schemas: unchanged (read-only)
- Agent Permission Model: unchanged
- Circuit Breaker: unchanged
- Expert Knowledge Injection: unchanged
- Human Expert Gates (HE0-HE3): unchanged
- R2-Physics (physical plausibility reviewer): unchanged

---

## PHASE 0: SCIENTIFIC BRAINSTORM (Before Everything)

Before any OTAE cycle, before any tree search, before any experiment — **BRAINSTORM**.

This is the phase where the research direction is born. It is not optional. It is not a chat. It is a structured, scientifically rigorous brainstorming session that produces a concrete, falsifiable research question grounded in real gaps in the literature and real available data.

### Why Phase 0 Exists

Most failed research starts with a bad question. AI-Scientist-v2 skips this entirely (it takes a pre-written idea). We don't. Phase 0 ensures we start with a question worth asking, gaps worth filling, and data that actually exists to answer it.

### Phase 0 Workflow

```
PHASE 0: SCIENTIFIC BRAINSTORM
├── Step 1: UNDERSTAND — What domain? What excites the researcher?
├── Step 2: LANDSCAPE  — What does the field look like right now?
├── Step 3: GAPS       — Where are the holes? What's missing?
├── Step 4: DATA       — What datasets exist to fill those gaps?
├── Step 5: HYPOTHESES — Generate 3-5 testable hypotheses
├── Step 6: TRIAGE     — Score and rank by feasibility + impact
├── Step 7: R2 REVIEW  — Reviewer 2 challenges the chosen direction
└── Step 8: COMMIT     — Lock in RQ, kill conditions, success criteria
```

### Step 1: UNDERSTAND (Context Gathering)

Dispatch to: `scientific-brainstorming` MCP skill (Phase 1: Understanding the Context)

- Ask the user open-ended questions about their domain, interests, constraints
- One question at a time, prefer multiple choice when possible (from `superpowers:brainstorming`)
- Identify: domain expertise, available resources, time constraints, ambition level
- Listen for implicit assumptions, unexplored angles, personal excitement
- Output: `00-brainstorm/context.md`

### Step 2: LANDSCAPE (Field Mapping)

Dispatch to: `literature-review` + `openalex-database` + `web_search` skills (IEEE Xplore, Optica, SPIE)

- Rapid literature scan of the identified domain (last 3-5 years)
- Map the major players, key papers, dominant methods, open debates
- Identify review papers and meta-analyses as anchors
- Build a mental map: what's crowded (red ocean) vs. what's empty (blue ocean)
- Output: `00-brainstorm/landscape.md` with field map

### Step 3: GAPS (Blue Ocean Hunting)

This is the core of Phase 0. Dispatch to: `scientific-brainstorming` (Phase 2: Divergent Exploration)

Techniques applied systematically:
- **Cross-Domain Analogies**: What methods from field X haven't been tried in field Y?
- **Assumption Reversal**: What does everyone assume that might be wrong?
- **Scale Shifting**: What happens at a different scale (single-channel vs. WDM, device-level vs. system-level)?
- **Constraint Removal**: "What if you could measure anything?" → then check what's actually measurable
- **Technology Speculation**: What new tools (integrated photonics, co-packaged optics, silicon photonics, etc.) open new doors?
- **Contradiction Hunting**: Where do two well-cited papers disagree?

For each gap found, assess:
- Is this gap real or just my ignorance? (check with targeted search)
- Is anyone already working on this? (check preprints: arXiv (physics.optics, eess.SP))
- Why hasn't this been done? (technical limitation? lack of data? not interesting enough?)

Output: `00-brainstorm/gaps.md` with ranked list of identified gaps

### Step 4: DATA (Reality Check — LAW 1 Applies Here)

`NO DATA = NO GO.` This step kills beautiful hypotheses that can't be tested.

Dispatch to: `openalex-database`, `web_search` (IEEE Xplore), `perplexity-search`, and domain-specific database skills

For each promising gap:
- Does public data exist to investigate it? Search IEEE Xplore, Optica, SPIE, arXiv physics.optics, ITU-T standards
- What format is it in? How much preprocessing is needed?
- Is the sample size sufficient for the intended analysis?
- Are there confounders or batch effects that would invalidate the approach?

Score each gap: DATA_AVAILABLE (0-1) based on quantity, quality, accessibility.
Gaps with DATA_AVAILABLE < 0.3 are moved to "future" pile, not killed.

Output: `00-brainstorm/data-audit.md`

### Step 5: HYPOTHESES (From Gaps to Testable Questions)

Dispatch to: `hypothesis-generation` MCP skill + `scientific-brainstorming` (Phase 3: Connection Making)

For each top-ranked gap with available data, generate:
- A **precise, falsifiable hypothesis** (not vague, not unfalsifiable)
- A **null hypothesis** (what we expect if the effect doesn't exist)
- **Predictions**: if true, we should see X; if false, we should see Y
- **Mechanistic explanation**: WHY might this be true? What's the physics/mechanism?

Generate 3-5 competing hypotheses. Each must be:
- Testable with available data (Step 4 passed)
- Distinguishable from the others (different predictions)
- Interesting enough to publish if confirmed OR denied

Output: `00-brainstorm/hypotheses.md`

### Step 6: TRIAGE (Pick the Winner)

Score each hypothesis on a 2x2 matrix:

```
                    HIGH FEASIBILITY
                         ▲
                         │
        Sweet spot ──→   │   ← Start here if unsure
        (publishable +   │     (safe bet)
         achievable)     │
                         │
   ─────────────────────┼──────────────────→ HIGH IMPACT
                         │
        Ignore           │   Moon shot
        (hard + boring)  │   (hard but transformative)
                         │
```

Criteria:
- **Impact** (0-5): How much would this change the field?
- **Feasibility** (0-5): Can we do this with available data + tools?
- **Novelty** (0-5): How different is this from existing work?
- **Data readiness** (0-5): How close is the data to being usable?
- **Serendipity potential** (0-5): How likely is this to generate unexpected discoveries?

Total score /25. Rank hypotheses. Present top 3 to user with trade-offs.

Output: `00-brainstorm/triage.md`

### Step 7: R2 REVIEW OF BRAINSTORM (Reviewer 2 is co-pilot from day zero)

**R2 reviews the brainstorm output BEFORE any OTAE cycle starts.**

R2 ensemble (at least R2-Methods + R2-Physics) challenges:
- Is the gap real? Or are we reinventing the wheel?
- Is the hypothesis truly falsifiable? Or is it unfalsifiable fluff?
- Is the data actually sufficient? Or are we kidding ourselves?
- Are there obvious confounders or biases we're ignoring?
- Is this the MOST interesting question we could ask given the gaps found?

R2 can demand:
- Additional literature search on a specific sub-topic
- Reformulation of the hypothesis
- Different data source
- Complete pivot to a different gap

**R2 verdict on brainstorm must be at least WEAK_ACCEPT before proceeding to OTAE.**

Output: `05-reviewer2/brainstorm-review.md`

### Step 8: COMMIT (Lock In)

After R2 clearance:
1. Finalize RQ.md with: question, hypothesis, predictions, success criteria, kill conditions
2. Set tree mode: LINEAR | BRANCHING | HYBRID
3. Create full folder structure
4. Populate STATE.md, PROGRESS.md, TREE-STATE.json
5. Enter first OTAE cycle with a **solid foundation**

### Phase 0 Gate: B0 (Brainstorm Quality)

```
B0 PASS requires ALL of:
  - At least 3 gaps identified with evidence
  - At least 1 gap verified as not-yet-addressed (preprint check)
  - Data availability confirmed for chosen hypothesis (DATA_AVAILABLE >= 0.5)
  - Hypothesis is falsifiable (null hypothesis stated)
  - R2 brainstorm review: WEAK_ACCEPT or better
  - User approved the chosen direction
```

### Phase 0 Artifacts

```
.vibe-science/RQ-001-[slug]/
├── 00-brainstorm/
│   ├── context.md          # User's domain, interests, constraints
│   ├── landscape.md        # Field map, key papers, major players
│   ├── gaps.md             # Identified gaps with evidence + ranking
│   ├── data-audit.md       # Data availability for each gap
│   ├── hypotheses.md       # 3-5 competing hypotheses with predictions
│   └── triage.md           # Scoring matrix + final ranking
```

---

## CORE CONCEPT: OTAE INSIDE TREE NODES

v3.5 had a flat OTAE loop: cycle 1 → cycle 2 → cycle 3 → ...

v4.0 has a **tree of OTAE nodes**:

```
                         root
                        /    \
                    node-A   node-B        ← each is a full OTAE cycle
                   / |  \      |
                A1  A2  A3    B1           ← children = variations
               /
             A1a                           ← deeper exploration
```

Each node executes one complete OTAE cycle (Observe parent → Think plan → Act execute → Evaluate score). The tree search engine selects which node to expand next based on Evidence Engine confidence + metrics.

**When to branch vs. stay linear:**
- Literature review → LINEAR (sequential cycles, like v3.5)
- Computational experiments → BRANCHING (tree search over variants)
- Mixed research → HYBRID (linear discovery phase, then branch for experiments)

---

## THE OTAE-TREE LOOP

```
╔═══════════════════════════════════════════════════════════════╗
║                    OTAE-TREE LOOP (v4.0)                      ║
╠═══════════════════════════════════════════════════════════════╣
║                                                               ║
║  ┌─── OBSERVE ──────────────────────────────────────────┐    ║
║  │  Read STATE.md + TREE-STATE.json                     │    ║
║  │  Identify current stage (1-5)                        │    ║
║  │  Load current node context + parent chain            │    ║
║  │  Check pending: gates, R2 demands, stage transitions │    ║
║  │  Verify STATE ↔ TREE consistency                     │    ║
║  └──────────────────────────────────────────────────────┘    ║
║                         ↓                                     ║
║  ┌─── THINK ────────────────────────────────────────────┐    ║
║  │  TREE MODE:                                          │    ║
║  │    Which node to expand? (best-first selection)      │    ║
║  │    What type? (draft|debug|improve|hyper|ablation)   │    ║
║  │    What would falsify the parent's result?           │    ║
║  │  LINEAR MODE:                                        │    ║
║  │    Same as v3.5 — next highest-value action          │    ║
║  │  Plan: search | analyze | extract | compute | write  │    ║
║  └──────────────────────────────────────────────────────┘    ║
║                         ↓                                     ║
║  ┌─── ACT ──────────────────────────────────────────────┐    ║
║  │  Execute the planned action:                         │    ║
║  │  • Literature search → search-protocol.md            │    ║
║  │  • Data analysis → analysis-orchestrator.md          │    ║
║  │  • Tree node experiment → auto-experiment.md         │    ║
║  │  • Hypothesis generation → serendipity-engine.md     │    ║
║  │  • Tool dispatch → skill-router.md                   │    ║
║  │  Produce ARTIFACTS (files, figures, manifests)        │    ║
║  │  If buggy: debug (max 3 attempts, then prune node)   │    ║
║  └──────────────────────────────────────────────────────┘    ║
║                         ↓                                     ║
║  ┌─── EVALUATE ─────────────────────────────────────────┐    ║
║  │  Extract claims → CLAIM-LEDGER                       │    ║
║  │  Score confidence (formula: E·R·C·K·D → 0-1)        │    ║
║  │  Parse metrics (if computational node)               │    ║
║  │  VLM feedback on figures (if available) → G6         │    ║
║  │  Check assumptions → ASSUMPTION-REGISTER             │    ║
║  │  Detect serendipity (including cross-branch)         │    ║
║  │  Apply relevant GATE (G0-G6, L0-L2, D0-D2, T0-T3,  │    ║
║  │    V0, J0)                                          │    ║
║  │  Mark node: good | buggy | pruned                    │    ║
║  │  Gate FAIL? → triage, fix, re-gate                   │    ║
║  └──────────────────────────────────────────────────────┘    ║
║                         ↓                                     ║
║  ┌─── CHECKPOINT ───────────────────────────────────────┐    ║
║  │  Stage gate check (S1-S5): advance stage?            │    ║
║  │  Tree health check (T3): ratio good/total >= 0.2?    │    ║
║  │                                                       │    ║
║  │  R2 CO-PILOT CHECK (expanded triggers):              │    ║
║  │    FORCED: major finding / stage transition /         │    ║
║  │      confidence explosion / pivot / brainstorm        │    ║
║  │    BATCH:  3 minor findings accumulated              │    ║
║  │    SHADOW: every 3 cycles, R2 passively reviews      │    ║
║  │      tree health + claim ledger + assumption drift.   │    ║
║  │      Shadow can escalate to FORCED if it spots risk.  │    ║
║  │    VETO:   R2 can halt any branch it deems unsound   │    ║
║  │    If triggered → reviewer2-ensemble.md (BLOCKING)    │    ║
║  │                                                       │    ║
║  │  SERENDIPITY RADAR (active every cycle):             │    ║
║  │    Scan current node for anomalies & unexpected       │    ║
║  │    Compare cross-branch: pattern only visible across? │    ║
║  │    Check contradiction register: new contradictions?  │    ║
║  │    Score >= 10 → serendipity-engine.md triage          │    ║
║  │    Score >= 15 → INTERRUPT: create serendipity node   │    ║
║  │                                                       │    ║
║  │  Stop conditions? → EXIT or CONTINUE                 │    ║
║  │                                                       │    ║
║  │  v5.0 FORCED review path:                             │    ║
║  │    SFI injection → BFP Phase 1 (blind) →              │    ║
║  │    Full review Phase 2 → V0 gate (vigilance) →        │    ║
║  │    R3/J0 gate (judge) → Schema validation →           │    ║
║  │    Normal gate evaluation.                             │    ║
║  │    See protocols/seeded-fault-injection.md,            │    ║
║  │    protocols/blind-first-pass.md,                      │    ║
║  │    protocols/judge-agent.md,                            │    ║
║  │    protocols/schema-validation.md.                     │    ║
║  │                                                       │    ║
║  │  BATCH and SHADOW reviews unchanged from v4.5.        │    ║
║  └──────────────────────────────────────────────────────┘    ║
║                         ↓                                     ║
║  ┌─── CRYSTALLIZE (LAW 10: NOT IN FILE = DOESN'T EXIST) ──┐    ║
║  │  Update STATE.md (rewrite, max 100 lines)            │    ║
║  │  Update TREE-STATE.json (full tree serialization)    │    ║
║  │  Write/update node file in 08-tree/nodes/            │    ║
║  │  Append PROGRESS.md (cycle summary)                  │    ║
║  │  Update CLAIM-LEDGER.md, ASSUMPTION-REGISTER.md      │    ║
║  │  Update tree-visualization.md                        │    ║
║  │  Save intermediate data (CSVs, metrics, figures)     │    ║
║  │  Log decisions with reasoning in decision-log        │    ║
║  │  VERIFY: every ACT result exists as a file on disk   │    ║
║  │  → LOOP BACK TO OBSERVE                              │    ║
║  └──────────────────────────────────────────────────────┘    ║
║                                                               ║
╚═══════════════════════════════════════════════════════════════╝
```

---

## TREE SEARCH ENGINE

The tree search engine manages hypothesis exploration as a tree of OTAE nodes. Each node executes one complete OTAE cycle; the engine selects which node to expand next based on Evidence Engine confidence and metrics. Supports 7 node types across 3 tree modes (LINEAR, BRANCHING, HYBRID).

### Node Types (summary)

| Type | When | Description |
|------|------|-------------|
| `draft` | Stage 1+ | New experimental approach |
| `debug` | Any stage | Fix attempt (max 3 per parent, then prune) |
| `improve` | Stage 2+ | Refinement of working approach |
| `hyperparameter` | Stage 2 | Parameter variation |
| `ablation` | Stage 4 | Remove one component to test contribution |
| `replication` | Stage 4-5 | Same config, different seed |
| `serendipity` | Any | Unexpected branch from serendipity detection |

> **Full protocol:** `protocols/tree-search.md`
> Contains: tree modes (LINEAR/BRANCHING/HYBRID), 7 node types, best-first selection algorithm, pruning rules, tree health monitoring (T3).

---

## REVIEWER 2 CO-PILOT SYSTEM (Expanded from v3.5)

In v3.5, Reviewer 2 was a gate. In v4.0, **Reviewer 2 is a co-pilot that flies with you the entire session.**

### R2 Activation Modes

| Mode | Trigger | Scope | Blocking? |
|------|---------|-------|-----------|
| **BRAINSTORM** | Phase 0 completion | Reviews gap analysis, hypothesis quality, data availability | YES — must WEAK_ACCEPT before OTAE starts |
| **FORCED** | Major finding, stage transition, pivot, confidence explosion (>0.30/2cyc) | Full ensemble (4 reviewers), double-pass | YES — demands must be addressed |
| **BATCH** | 3 minor findings accumulated | Single-pass batch review, R2-Methods lead | YES — demands must be addressed |
| **SHADOW** | Every 3 cycles automatically | Passive review of tree health, claim ledger drift, assumption register, serendipity log | NO — but can ESCALATE to FORCED |
| **VETO** | R2 spots fatal flaw during any mode | Halts current branch or entire tree | YES — cannot be overridden except by human |
| **REDIRECT** | R2 identifies better direction during review | Proposes alternative branch, alternative hypothesis, or return to Phase 0 | Soft — user chooses whether to follow |
| **INLINE** | Every finding formulated (v5.5) | 7-point checklist: numbers match source, sample size, alternatives, terminology, claim ≤ evidence, traceability, hostile read | YES — anomalies block; clean findings pass |

### R2 Shadow Mode Protocol (every 3 cycles)

```
R2 Shadow Check:
1. Read CLAIM-LEDGER.md — any confidence scores drifting up without new evidence?
2. Read ASSUMPTION-REGISTER.md — any HIGH-risk assumptions untested for 5+ cycles?
3. Read tree-visualization.md — is the tree lopsided? (one branch getting all attention)
4. Read SERENDIPITY.md — any flags ignored for 3+ cycles?
5. Compute: assumption_staleness, confidence_drift, tree_balance, serendipity_neglect

If ANY metric is concerning:
  → Log warning in PROGRESS.md
  → If 2+ metrics concerning → ESCALATE to FORCED R2 review
```

### R2 Powers (v4.0 — expanded)

1. **DEMAND EVIDENCE**: R2 can require specific evidence before any claim is promoted. Demands have deadlines.
2. **FORCE FALSIFICATION**: R2 can require the system to actively try to disprove a claim before accepting it. Minimum 3 falsification tests per major claim.
3. **VETO BRANCH**: R2 can mark a tree branch as "unsound" — no further expansion until R2 concerns addressed.
4. **REDIRECT**: R2 can propose an alternative research direction during review. The system must present this to the user.
5. **CHALLENGE BRAINSTORM**: R2 reviews Phase 0 output and can force reconsideration of the research question itself.
6. **AUDIT TRAIL**: Every R2 decision is logged with reasoning. R2 cannot be silent — it must always explain.

### R2 Ensemble Composition (expanded from v3.5)

| Reviewer | Focus | Active In | Key Obligation |
|----------|-------|-----------|----------------|
| R2-Methods | Search completeness, experimental design, statistical validity | ALL modes | De

…(truncated)
