Grounded Review
This skill performs the quality-gating and repair stage of the middle layer.
It reads:
- the grounded note,
- the literature result,
- the report draft,
- optional supporting research artifacts,
and produces:
- a review diagnosis and score report
- a review state file
- a final deliverable research report
This skill is not a broad search skill and not an output rendering skill. It sits between:
grounded-summary- and the final output-formatting/rendering stage
Position in the pipeline
The intended pipeline is:
grounding -> grounded-research-lit -> grounded-summary -> grounded-review -> output rendering
This means:
grounded-summarywrites the main report draftgrounded-reviewscores, diagnoses, repairs, re-checks, and finalizes that draft- output skills then present that final report as pdf / docx / md / slides / audio / charts / fancy formats
grounded-review should therefore behave like a quality gate with bounded repair, not like a second compression stage.
Purpose
The purpose of this skill is to turn a substantial report draft into a final deliverable research report through:
- structured evaluation,
- weakness diagnosis,
- bounded repair,
- re-evaluation,
- finalization.
It should:
- preserve the core structure and substance of the report draft
- improve evidence alignment
- catch overclaiming
- restore missing important technical details if the draft still lost them
- improve section balance and readability
- strengthen explicit linkage between grounded findings and literature evidence
- remove search-process noise from the report body
- make the report suitable for downstream output-format skills
- act as a quality gate, not just a surface editor
It should not:
- perform an unrestricted new literature search
- invent new claims not grounded in the existing inputs
- radically change the research direction
- rewrite a rich report into a shorter but thinner version
- act as a presentation-formatting skill
- run an unbounded self-improvement loop
When to Use
Use this skill when:
grounded.mdalready existslit.mdalready existssummary.mdalready exists- you want a final deliverable research report for the current grounded item
Do not use this skill when:
- the report draft has not been written yet
- the literature result has not been written yet
- you only want a quick synthesis draft rather than a final report
- you only want formatting or export to pdf/docx/slides/audio
Inputs
How to get ground_id
Read ground_id.txt from the grounding bundle to get the stable pipeline identifier:
data/grounded_notes/<ground_id>/ground_id.txt
Do NOT generate a new ground_id. All downstream directories reuse the same ground_id.
This skill assumes the following files already exist:
data/grounded_notes/<ground_id>/grounded.mddata/lit_results/<ground_id>/lit.mddata/report_inputs/<ground_id>/summary.md
Optional but strongly recommended supporting inputs:
data/lit_inputs/<ground_id>/opened_paper_notes.jsonldata/lit_inputs/<ground_id>/search_results.jsondata/lit_downloads/<ground_id>/manifest.jsondata/lit_inputs/<ground_id>/refine_coverage.jsondata/lit_inputs/<ground_id>/lit_initial.md
Input reading rule
At minimum, grounded.md, lit.md, and summary.md must be read.
Do not review the report draft in isolation.
The grounded note and literature result must be used as the source of truth for:
- evidence support
- scope
- uncertainty
- open questions
- next-step recommendations
- technical detail recovery when the draft is still too thin
When available, use opened_paper_notes.jsonl, manifest.json, and refine_coverage.json to validate:
- downloaded-paper coverage
- paper-level evidence sufficiency
- whether the draft has become thinner than the research-stage output
The report draft should be treated as:
- the main draft to be reviewed
- the default structural base to preserve
- not as unquestionable truth
Outputs
The review stage uses a round-based archival directory structure to preserve full review history for later comparison.
Directory Structure
data/review_outputs/<ground_id>/
├── round_0/
│ ├── review_report.md # Round 0 review diagnosis
│ └── review_state.json # Round 0 state snapshot
├── round_1/
│ ├── review_report.md # Round 1 review diagnosis
│ └── review_state.json # Round 1 state snapshot
├── round_2/ # (optional, if loop runs to round 2)
│ ├── review_report.md
│ └── review_state.json
├── round_3/ # (optional, if loop runs to round 3)
│ ├── review_report.md
│ └── review_state.json
├── round_4/ # (optional, if loop runs to round 4)
│ ├── review_report.md
│ └── review_state.json
├── round_5/ # (optional, if loop runs to round 5)
│ ├── review_report.md
│ └── review_state.json
├── review_history.json # Summary of all rounds (canonical reference)
├── review_report.md # Symlink to latest round (for backward compatibility)
└── review_state.json # Symlink to latest round (for backward compatibility)
data/reports/<ground_id>/
└── research_report.md # Current final report (updated after each repair round)
Required Per-Round Outputs
For each review round, write these files to data/review_outputs/<ground_id>/round_<N>/:
review_report.md— human-readable review diagnosis (per-round)review_state.json— machine-readable state snapshot (per-round)
Required Final Outputs
After the loop completes, update these files in data/review_outputs/<ground_id>/:
review_history.json— summary of all rounds (canonical reference for downstream skills)review_report.md— symlink/copy of the latest round's reportreview_state.json— symlink/copy of the latest round's state
And finalize the report in:
data/reports/<ground_id>/research_report.md— the final deliverable report
Optional supporting file if genuinely useful:
data/reports/<ground_id>/review_notes.md
If review_notes.md is not necessary, do not create it.
Output roles
review_report.md
A human-readable review diagnosis containing:
- rubric scores
- pass/fail judgement
- specific weaknesses
- repair priorities
- what changed across rounds
This file should be persisted by the parent/main agent.
If the reviewer subagent is read-only, the reviewer may return a structured diagnosis payload in chat, and the parent agent must write that payload into review_report.md.
review_state.json
A machine-readable state file containing at least:
ground_idroundscoresweighted_totalverdictneeds_repairrepair_actionspassedused_reviewer_roleused_writer_rolereviewer_agent_pathwriter_agent_pathreviewer_model_hintwriter_model_hintreviewer_independencefinal_report_path
This file should also be persisted by the parent/main agent. A read-only reviewer may populate its contents indirectly by returning structured review results to the parent agent.
research_report.md
The final deliverable research report for the current grounded item. This is the version that downstream output-format skills should use.
Reviewer Role and Independence
The review stage should be performed by a dedicated reviewer role / subagent when available.
Preferred behavior:
- writer role produces the draft
- reviewer role scores and diagnoses the draft
- writer role performs bounded repairs
- reviewer role re-checks the repaired draft
When your environment supports reviewer subagents with their own model configuration, prefer a reviewer model that is different from the writer model.
If a separate reviewer model is not available, the review stage may still run with the same base model, but it must explicitly record in review_state.json that reviewer independence is limited.
Do not let the same role simply "declare success" without rubric-based justification.
Execution Flow
The review stage must execute through a bounded loop with dedicated reviewer and writer subagents, not as a single monolithic pass in the parent context.
Required execution pattern
For every review task, the parent agent must follow this loop. The reviewer is either the external API (if reviewer_api_config is present) or the local reviewer subagent (if not present). The writer is always the local writer subagent via Task tool.
- Call reviewer (Round 0):
- If
reviewer_api_configpresent: callcall_ext_api.pyvia Shell → external API - If not present: launch reviewer subagent via
Tasktool
- If
- Parent agent archives the previous round (if round > 0: copy current
round_n→round_{n-1}/) - Parent agent persists
review_report.mdandreview_state.jsontoround_<N>/directory - If verdict is
repairand round < 5: launch writer subagent viaTasktool to apply repairs, then go to step 6 - If verdict is
pass: launch writer subagent viaTasktool to apply reviewer's suggested quality improvements (light repair), then skip to step 9 - Call reviewer again (Round N+1) — same method as step 1:
- If external mode: call
call_ext_api.py - If local mode: launch reviewer subagent via
Tasktool
- If external mode: call
- Parent agent archives the previous round → copy current
round_n→round_{n-1}/ - Parent agent persists the new round's
review_report.mdandreview_state.jsontoround_<N>/ - Repeat steps 4–8 for at most 5 repair rounds (pass verdicts jump from step 5 to finalize)
- After reviewer passes (or loop exhausted): writer applies light repair → update
review_history.json, symlinks, and finalizeresearch_report.md
External Reviewer Mode
When reviewer_api_config is provided in the invocation, the review stage uses the external LLM API as the reviewer for all rounds of the bounded loop (Round 0, Round 1, ..., Round N). The external reviewer is not a one-shot first-round tool — it is the primary reviewer throughout. The same external API is called at each round after each writer repair, until the report passes or the loop is exhausted.
When to Use
Use external reviewer mode when:
- You want to use a specific external model for review (e.g., a cheaper or more capable model)
- You need to integrate with existing API infrastructure
- You want more control over the reviewer's model parameters
Execution Pattern
When external reviewer config is available, it is used for every review round throughout the bounded loop:
┌─────────────────────────────────────────────────────────────────────┐
│ External Reviewer — All Rounds Loop │
├─────────────────────────────────────────────────────────────────────┤
│ │
│ Round 0: │
│ 1. Parent Agent receives reviewer_api_config │
│ 2. Parent Agent reads: grounded.md, lit.md, summary.md │
│ 3. Parent Agent constructs system + user prompts │
│ 4. Parent Agent calls call_ext_api.py via Shell (Round 0) │
│ 5. Parse JSON → persist round_0/review_report.md + review_state.json│
│ 6. If verdict=pass: launch writer for light repair → finalize │
│ 7. If verdict=repair: launch writer for repair → continue │
│ │
│ Round N (N >= 1): │
│ 8. Parent Agent reads: revised research_report.md + round_N-1/ │
│ 9. Parent Agent constructs prompts with Round N-1 context │
│ 10. Parent Agent calls call_ext_api.py via Shell (Round N) │
│ 11. Parse JSON → persist round_N/review_report.md + review_state.json│
│ 12. If verdict=pass: launch writer for light repair → finalize │
│ 13. If verdict=repair and N < 5: launch writer → continue │
│ 14. If verdict=repair and N >= 5: loop exhausted → finalize │
│ │
└─────────────────────────────────────────────────────────────────────┘
Key invariant: whenever reviewer_api_config is present, the parent agent must use call_ext_api.py for every review round — not just Round 0. Switching to the local reviewer subagent mid-loop violates the external reviewer contract.
External Reviewer API Script
Use the standalone Python script call_ext_api.py for direct API calls.
Script location: .cursor/skills/grounded-review/external-reviewer/call_ext_api.py
Usage:
python .cursor/skills/grounded-review/external-reviewer/call_ext_api.py \
--provider openai \
--model gpt-4o-mini \
--api-key "sk-xxx" \
--base-url "https://api.example.com/v1" \
--temperature 0.2 \
--max-tokens 4096 \
--system-prompt "Your reviewer role definition..." \
--user-prompt "Review this report..."
Script output: JSON with { content, usage } or { error, status: "failed" }
Invocation method: Use the Shell tool to run the script and parse JSON output.
Prompt Construction for External Reviewer
System prompt (from .cursor/agents/reviewer.md):
- The reviewer role definition
- The scoring rubric (6 dimensions with weights)
- The verdict threshold rules
- The output format requirements
User prompt — Round 0 (initial review):
Construct from:
- The full content of
grounded.md - The full content of
lit.md - The full content of
summary.md(current draft) - Instructions to score and return structured JSON
User prompt — Round N (N >= 1, re-check after writer repair):
Construct from:
- The full content of
grounded.md - The full content of
lit.md - The current
research_report.md(post-repair draft) data/review_outputs/<ground_id>/round_<N-1>/review_report.md(previous round diagnosis)- Instructions: "This is Round N re-check. Round N-1 verdict was [repair/pass]. The writer applied [list repair actions]. Please score the revised draft and return structured JSON."
Response Parsing
The external API response is a JSON string. Parse it to extract:
{
"content": "<model's response text>",
"usage": {
"prompt_tokens": 1000,
"completion_tokens": 500,
"total_tokens": 1500
}
}
The content should be parsed as a structured review payload with:
scores: 6 dimension scores with raw/weight/subtotalweighted_total: computed weighted scoreverdict: "pass" | "repair"needs_repair: booleanrepair_actions: array of {action, details}main_weaknesses: array of weakness descriptions
External vs Local Reviewer
| Aspect | External Reviewer (Shell) | Local Reviewer (Task) |
|---|---|---|
| Model | Fully customizable | Configured in reviewer.md |
| Invocation | Shell + Python script | Task tool + subagent |
| Independence | Guaranteed (different API) | Depends on model config |
| Latency | Network call required | Faster (in-process) |
| Cost | External API pricing | Cursor subscription |
| Stability | No MCP connection issues | Stable (in-process) |
Contract Compliance for External Reviewer
When using external reviewer, the following fields must still be recorded in review_state.json:
used_reviewer_role: must betrueused_writer_role: must betrue(if repair was needed)reviewer_independence: must be"high"(external model is fully independent)reviewer_agent_path:null(no local reviewer agent used)reviewer_model_hint: the external model name from configexternal_reviewer_config: the full config used (for audit)
Error Handling
If external API call fails:
- If
fallback_to_local: true: retry with local Cursor reviewer - If
fallback_to_local: false: report error and stop review
review_history.json format
After each round, append the round's summary to data/review_outputs/<ground_id>/review_history.json:
{
"ground_id": "<ground_id>",
"total_rounds": 2,
"rounds": [
{
"round": 0,
"weighted_total": 76.0,
"verdict": "repair",
"needs_repair": true,
"scores": {
"topic_alignment": {"raw": 4, "weight": 15, "subtotal": 12.0},
"coverage_completeness": {"raw": 4, "weight": 20, "subtotal": 16.0},
"evidence_specificity": {"raw": 4, "weight": 20, "subtotal": 16.0},
"analytical_depth": {"raw": 3, "weight": 20, "subtotal": 12.0},
"structure_and_narrative_coherence": {"raw": 4, "weight": 15, "subtotal": 12.0},
"deliverability": {"raw": 4, "weight": 10, "subtotal": 8.0}
},
"repair_actions": [{"action": "...", "details": "..."}],
"reviewer_agent_path": ".cursor/agents/reviewer.md",
"writer_agent_path": null,
"reviewer_model_hint": "<model from reviewer.md frontmatter>",
"writer_model_hint": "<model from writer.md frontmatter>",
"reviewer_independence": "high",
"timestamp": "2026-04-06T14:54:00Z"
},
{
"round": 1,
"weighted_total": 93.0,
"verdict": "pass",
"needs_repair": false,
"scores": {
"topic_alignment": {"raw": 4, "weight": 15, "subtotal": 12.0},
"coverage_completeness": {"raw": 5, "weight": 20, "subtotal": 20.0},
"evidence_specificity": {"raw": 5, "weight": 20, "subtotal": 20.0},
"analytical_depth": {"raw": 4, "weight": 20, "subtotal": 16.0},
"structure_and_narrative_coherence": {"raw": 5, "weight": 15, "subtotal": 15.0},
"deliverability": {"raw": 5, "weight": 10, "subtotal": 10.0}
},
"repair_actions": [],
"reviewer_agent_path": ".cursor/agents/reviewer.md",
"writer_agent_path": ".cursor/agents/writer.md",
"reviewer_model_hint": "<model from reviewer.md frontmatter>",
"writer_model_hint": "<model from writer.md frontmatter>",
"reviewer_independence": "high",
"timestamp": "2026-04-06T15:11:00Z"
}
]
}
After the loop completes, also update the symlinks/copies at the root of review_outputs/<ground_id>/ so existing tooling that reads review_report.md and review_state.json directly continues to work.
Choosing Between External and Local Reviewer
| Scenario | Recommended Approach |
|---|---|
reviewer_api_config provided |
External reviewer for all rounds via Shell + call_ext_api.py |
reviewer_api_config provided, Round N API call fails + fallback_to_local: true |
Retry with local reviewer for that round only; resume external for subsequent rounds |
reviewer_api_config provided, Round N API call fails + fallback_to_local: false |
Report error and stop review |
reviewer_api_config NOT provided |
Local reviewer via Task tool (all rounds) |
Critical rule: When reviewer_api_config is provided, the parent agent must not switch to the local reviewer subagent after Round 0. The external API is the designated reviewer for every round. Using the local reviewer mid-loop violates the contract, regardless of whether the local reviewer could produce a valid score.
CRITICAL: The reviewer subagent must be created via Task tool with subagent_type="generalPurpose". The prompt must start with an agent role declaration block to trigger Cursor's agent auto-matching, which will load .cursor/agents/reviewer.md and its frontmatter model field. The reviewer is read-only — it must not write files directly.
Use the Task tool with subagent_type="generalPurpose". The prompt must start with:
---
You are operating as the **REVIEWER AGENT** (grounded-review-reviewer).
Read and follow .cursor/agents/reviewer.md now to load your reviewer configuration and model.
---
Then the prompt must include:
- The exact
ground_idfor the current grounded unit - Instruction to read
.cursor/agents/reviewer.mdfirst - Instruction to read the canonical inputs:
grounded.md,lit.md,summary.md, and available supporting artifacts - Instruction to return a structured review payload
Example reviewer Task invocation:
---
You are operating as the **REVIEWER AGENT** (grounded-review-reviewer).
Read and follow .cursor/agents/reviewer.md now to load your reviewer configuration and model.
---
Use the grounded-review skill to review the current report draft.
ground_id: <the current grounded unit ID, e.g. "meeting_001_topic01">
round: <current round number, e.g. 0>
Inputs to read:
- data/grounded_notes/<ground_id>/grounded.md
- data/lit_results/<ground_id>/lit.md
- data/report_inputs/<ground_id>/summary.md
- data/reports/<ground_id>/research_report.md (if it already exists — read before scoring)
- data/review_outputs/<ground_id>/round_<N-1>/review_report.md (if round > 0 — for repair progress context)
- data/lit_inputs/<ground_id>/opened_paper_notes.jsonl (if exists)
- data/lit_downloads/<ground_id>/manifest.json (if exists)
- data/lit_inputs/<ground_id>/refine_coverage.json (if exists)
Task: Score the current draft using the rubric in grounded-review/SKILL.md,
check hard gates, diagnose weaknesses, produce minimum repair actions, and return a structured
review payload in your response.
IMPORTANT: Do NOT write files directly. Return a structured diagnosis
payload that the parent agent will persist to disk at data/review_outputs/<ground_id>/round_<N>/.
The payload must include:
- scores (all 6 dimensions with raw/weight/subtotal)
- weighted_total
- verdict (pass / repair)
- needs_repair (boolean)
- repair_actions (array of {action, details})
- passed (boolean)
- reviewer_independence: "high" (because reviewer agent loaded a dedicated model from reviewer.md)
- used_reviewer_role: true (explicitly record this)
- reviewer_agent_path: ".cursor/agents/reviewer.md"
- reviewer_model_hint: "<model slug from reviewer.md frontmatter>"
- timestamp: "<ISO 8601 timestamp, e.g. '2026-04-06T14:54:00Z'>"
Note: The actual model used by the reviewer subagent is determined entirely by the
modelfield in.cursor/agents/reviewer.mdfrontmatter. SKILL.md example prompts use a placeholder — the real value comes from the agent file.
Task tool invocation for writer
CRITICAL: The writer subagent must be created via Task tool with subagent_type="generalPurpose". The prompt must start with an agent role declaration block to trigger Cursor's agent auto-matching, which will load .cursor/agents/writer.md and its frontmatter model field. Do not skip this step even if the repair is simple.
Do not let the parent context apply repairs directly. The bounded repair loop requires a dedicated writer subagent. If verdict == "repair" and round < 2, the parent agent must launch a writer subagent via Task tool — it must not apply repairs in its own context.
Use the Task tool with subagent_type="generalPurpose". The prompt must start with:
---
You are operating as the **WRITER AGENT** (grounded-review-writer).
Read and follow .cursor/agents/writer.md now to load your writer configuration.
---
Then the prompt must include:
- The exact
ground_idfor the current grounded unit - Instruction to read
.cursor/agents/writer.mdfirst - The repair actions approved by the reviewer
- Instruction to preserve the draft's structure and substance
Example writer Task invocation:
---
You are operating as the **WRITER AGENT** (grounded-review-writer).
Read and follow .cursor/agents/writer.md now to load your writer configuration.
---
Use the grounded-review skill to repair the current report draft.
Do NOT skip reading writer.md.
ground_id: <the current grounded unit ID>
Repair actions approved by the reviewer (apply ONLY these — do NOT improvise additional changes):
- [list the minimum repair actions from the reviewer diagnosis]
Inputs to read:
- data/reports/<ground_id>/research_report.md (current draft — read before revising)
- data/grounded_notes/<ground_id>/grounded.md
- data/lit_results/<ground_id>/lit.md
- data/report_inputs/<ground_id>/summary.md
- data/review_outputs/<ground_id>/round_<N-1>/review_report.md (previous round diagnosis)
Task: Apply ONLY the specified repair actions. Preserve the draft's
structure and substance. Produce the revised report body and write it to
data/reports/<ground_id>/research_report.md.
IMPORTANT: You are the writer role. Only apply the specified repair actions.
Do NOT self-approve or declare the final verdict. Return the revised report body
for the reviewer to re-check.
After the writer completes, record in your response:
- used_writer_role: true
- writer_agent_path: ".cursor/agents/writer.md"
- writer_model_hint: "<model from writer.md frontmatter, e.g. 'inherit'>"
- repair_actions_applied: [list what was actually applied]
Loop termination
|| Condition | Action |
||-----------|--------|
|| weighted_total >= 90 and no hard-gate failure | Pass — launch writer for light repair, then finalize (no re-review needed) |
|| weighted_total < 90 or any hard-gate failure | Repair — run writer repair, then re-review via bounded loop (max 5 rounds) |
|| Initial + 5 repair rounds exhausted without pass | Finalize with explicit "loop exhausted" status in review_state.json |
Pass verdict: light repair without re-review
Pass means the report meets the quality gate. It does not mean the report is perfect or that reviewer diagnoses should be discarded.
When verdict is pass but the reviewer identified quality improvement suggestions (even if not hard-gate failures):
- The parent agent launches a writer subagent via
Tasktool to apply the reviewer's suggested improvements. - The writer reads the reviewer diagnosis and applies only the suggested improvements.
- No re-review round is triggered — the report is finalized directly after writer completes.
review_state.jsonrecordsused_writer_role: trueandrepair_actions: [suggested improvements applied].
The distinction from the repair loop:
- Repair loop (verdict = repair): writer repairs → reviewer re-checks → repeat. The reviewer gate determines when to stop.
- Light repair after pass (verdict = pass): writer applies suggestions → done. No reviewer re-check because the gate already passed.
This ensures reviewer diagnoses are never discarded, but avoids infinite re-review cycles when the report already meets quality standards.
Repair Enforcement Rule
This is a hard rule. Violations must not be silently accepted.
When a reviewer subagent returns verdict == "repair" and round < 5:
- The parent agent must launch a writer subagent via
Tasktool. - The parent agent must not apply repairs in its own context.
- After the writer completes, the parent agent must launch a reviewer subagent again for re-check.
If used_writer_role == false when verdict == "repair" and round < 2, the bounded repair loop was not executed. This is a contract violation regardless of whether research_report.md was written.
Remedy: Re-run the repair loop by launching a writer subagent, then re-review. Do not finalize with a repair verdict still pending.
In practice: Before finalizing research_report.md, check review_state.json. If verdict == "repair" and used_writer_role == false, re-launch the writer subagent and complete the loop.
Subagent File Contract
When your Cursor environment supports custom subagents, the review stage should use the following files when present:
.cursor/agents/reviewer.md.cursor/agents/writer.md
Expected role split
Reviewer subagent
The reviewer subagent is responsible for:
- reading
grounded.md,lit.md,summary.md, and available supporting artifacts - assigning rubric scores
- checking hard gates
- diagnosing concrete weaknesses
- producing the repair plan
- re-checking the revised draft
- deciding
pass/repair/insufficient
The reviewer should not be the primary author of the final report body except for the diagnosis documents.
If the reviewer subagent is configured as read-only, it should not directly modify repository files. In that case, it should return a structured diagnosis payload to the parent agent, and the parent agent must write:
data/review_outputs/<ground_id>/round_<N>/review_report.mddata/review_outputs/<ground_id>/round_<N>/review_state.json
from the reviewer output.
The reviewer is therefore responsible for the content of the review, while the parent agent is responsible for persisting that content to disk in the appropriate round_<N>/ directory when the reviewer cannot write files directly.
Writer subagent
The writer subagent is responsible for:
- preserving the draft structure where possible
- performing only the approved repair actions
- revising the report into
research_report.md - not inventing new evidence or new citations
- returning the revised report for reviewer re-check
The writer should not self-approve the final report without reviewer confirmation.
The writer may write research_report.md directly because the writer is the editing role.
However, the writer must not overwrite review_report.md or review_state.json with its own independent judgement.
Those files must reflect the reviewer-approved diagnosis and verdict.
Preferred execution order
When reviewer/writer subagents are available, use this order (as detailed in the Execution Flow section above):
- Parent agent launches reviewer subagent via
Tasktool to score and diagnose the current draft - Parent agent creates
round_<N>/directory and persistsreview_report.mdandreview_state.jsonfrom the reviewer output - Parent agent launches writer subagent via
Tasktool to apply the approved repair actions (if verdict isrepairand round < 5) - Parent agent archives the current round: if round > 0, copy
round_<N>/→round_<N-1>/ - Parent agent launches reviewer subagent via
Tasktool to re-check the revised draft - Parent agent creates a new
round_<N+1>/directory and persists the newreview_report.mdandreview_state.json - Repeat steps 3–6 for at most 5 repair rounds
- Only after reviewer passes, update
review_history.json, update symlinks/copies atreview_outputs/<ground_id>/, and finalizeresearch_report.md
The parent agent must not skip the Task-tool subagent calls or collapse the reviewer → writer → reviewer cycle into a single pass.
If the subagents are unavailable, the same base model may simulate both roles, but the role separation and reviewer_independence record must still be preserved.
Reviewer return contract
A reviewer subagent may be configured as read-only. If so, it should return a structured review payload in chat rather than trying to write files directly. That payload must contain enough information for the parent agent to persist both:
data/review_outputs/<ground_id>/round_<N>/review_report.mddata/review_outputs/<ground_id>/round_<N>/review_state.json
The parent agent must not skip file persistence merely because the reviewer role was read-only.
Parent-agent persistence rule
The parent/main agent is always responsible for ensuring that the required output files exist on disk at the end of the review stage. Read-only reviewer configuration does not relax the output-file requirement.
Core Review Responsibilities
The review stage should check and improve the report draft along these dimensions:
1. Topic alignment
Check whether the draft remains tightly aligned with:
- the grounded topic
- the source-derived problem setting
- the actual decision needs of the current project
Downweight or remove weakly related literature that inflates length without improving relevance.
2. Coverage completeness
Check whether the report adequately covers:
- the major grounded questions
- the most important themes from
lit.md - relevant downloaded and parsable papers when such coverage is expected from the research stage
3. Evidence specificity
Check whether major claims are actually supported by:
- the grounded note
- the literature result
- downloaded/opened evidence when available
If not, soften, qualify, or remove them.
4. Analytical depth
Check whether the report explains:
- what the method/problem actually is
- why a design matters
- what the evidence really supports
- what remains uncertain
5. Structure and narrative coherence
Check whether the report is well structured and readable.
You may reorganize or rewrite for clarity, but do not remove substance merely to make it shorter or cleaner.
6. Deliverability
Check whether the report is strong enough to be handed to the output layer.
The result should be:
- structurally clear
- technically grounded
- evidence-aware
- readable
- suitable for rendering/export
7. Removal of search-process noise
Check whether the report draft contains:
- query counts
- hit counts
- opened-link counts
- download counts
- manifest-like inventories
- retrieval execution commentary
These do not belong in the final report body and should be removed.
Review Authority Boundary
This skill may:
- revise wording
- reorganize material
- restore omitted important details
- improve balance across sections
- soften unsupported claims
- clarify evidence strength
- strengthen section transitions
- improve the final report structure
- trigger one bounded targeted evidence-recovery action when the current inputs are clearly insufficient
This skill must not:
- add new external evidence casually or without diagnosis
- run a full new literature survey by default
- fabricate experiments
- invent citations
- move substantive literature analysis out of the report body just to shorten the document
- change the research direction without support from the inputs
- run an unbounded repair loop
Structured Scoring Rubric
Before rewriting the final report, score the current draft across the following dimensions.
Use a 1–5 scale for each dimension, where:
1= very weak / seriously insufficient2= weak3= acceptable but clearly imperfect4= strong5= very strong
Dimensions
topic_alignmentcoverage_completenessevidence_specificityanalytical_depthstructure_and_narrative_coherencedeliverability
Weights
Use the following weights when computing the weighted total (/100):
topic_alignment: 15coverage_completeness: 20evidence_specificity: 20analytical_depth: 20structure_and_narrative_coherence: 15deliverability: 10
Weighted total formula
Compute the final weighted total as:
weighted_total =
(topic_alignment / 5) * 15+ (coverage_completeness / 5) * 20+ (evidence_specificity / 5) * 20+ (analytical_depth / 5) * 20+ (structure_and_narrative_coherence / 5) * 15+ (deliverability / 5) * 10
Rules:
- use exactly this formula
- do not invent any alternative scaling rule
- do not rescale weights
- round only at the final reported total if needed
- keep the unrounded internal calculation consistent with the displayed subtotal values
Required score presentation rule
For each dimension, the review report must explicitly show:
- the raw score (for example
4/5) - the weight
- the weighted subtotal computed as
(score / 5) × weight - an evidence-backed explanation
For example:
topic_alignment: 4/5, weight=15, subtotal=12.0
The final weighted_total must equal the sum of the six displayed subtotals.
If the displayed subtotals and final total are inconsistent, the review is invalid and must be corrected before finalizing review_report.md or review_state.json.
Scoring rule
Every score must be evidence-backed.
Do not assign a score without explicitly naming:
- what in the current draft supports the score
- what in the current draft weakens the score
- which sections or paper analyses are responsible for the weakness
The review report must not contain empty statements such as “depth is somewhat weak” without concrete explanation.
Dimension-specific scoring constraints
Use the following concrete rules to disambiguate score boundaries. These constraints reduce subjective leniency and make scoring consistent across reviewers.
1. topic_alignment
| Score | Condition |
|---|---|
| 5/5 | All grounded project directions are addressed with dedicated coverage; no drift; source context fully preserved |
| 4/5 | All directions covered, but 1-2 minor tangential sections present OR 1 important nuance from source weakened |
| 3/5 | 1 grounded direction missing or substantially weakened; moderate drift detected |
| 2/5 or below | Multiple directions missing or severely weakened; substantial drift |
Constraint: A direction that appears in grounded.md must appear in the report with comparable depth. Downweighting is permitted; silent omission is not. Assigning 4/5 requires naming exactly which section is tangential or which nuance was weakened.
2. coverage_completeness
| Score | Condition |
|---|---|
| 5/5 | All lit.md papers analyzed; all major themes covered; unresolved questions match lit.md's scope |
| 4/5 | All papers present but 1-2 papers receive shallow analysis; OR 1 theme underdeveloped |
| 3/5 | 1-2 papers silently omitted OR 2+ papers shallow; major theme missing or placeholder-level |
| 2/5 or below | Multiple papers omitted; report is clearly thinner than lit.md |
Constraint: Each downloaded PDF must be explicitly acknowledged with its refinement contribution. Assigning 4/5 requires naming exactly which paper(s) are shallow and what is missing. Assigning 3/5 or below requires naming the specific paper(s) omitted or shallow.
3. evidence_specificity
| Score | Condition |
|---|---|
| 5/5 | Every major claim cites specific numbers from lit.md; unvalidated claims are explicitly flagged with evidence boundary |
| 4/5 | Most claims supported by numbers, but 1-2 claims lack specific quantification OR 1 unvalidated claim not flagged |
| 3/5 | Several claims lack quantitative support; OR 2+ unvalidated claims presented without flag |
| 2/5 or below | Most claims unsupported; overclaiming pervasive |
Constraint: The phrase "unvalidated claim" must appear in the review when a grounded note claim is not matched by published literature. Assigning 4/5 requires naming the specific unsupported claim(s) and where they appear. A report that presents a claim like "within 3 centimeters" without flagging it as unvalidated cannot receive 5/5.
4. analytical_depth
| Score | Condition |
|---|---|
| 5/5 | For every major unresolved question: explains WHY it is unresolved, what evidence bears on it, and what would resolve it |
| 4/5 | Most questions analyzed, but 1-2 questions only restated without explaining the causal gap; OR 1 question lacks actionability |
| 3/5 | Analysis stays on the surface; questions mostly restated rather than analyzed; few causal explanations |
| 2/5 or below | No genuine analysis; report is a summary without synthesis |
Constraint: Assigning 4/5 requires naming the specific question(s) that were restated without causal analysis. Simply labeling a question "unresolved" does not count as analysis. The reviewer must identify why existing evidence does not resolve it and what specific evidence would.
5. structure_and_narrative_coherence
| Score | Condition |
|---|
…(truncated)