Bio-Logic: Scientific Reasoning Evaluation
Use structured frameworks to evaluate scientific claims, methodology, and evidence strength.
Instructions
- Identify the task (claim assessment, paper critique, study design review, project interpretation, or hypothesis revision).
- For project work, maintain a hypothesis register with at least 5 distinct active hypotheses until the project is no longer exploratory.
- After each major intermediate result, reflect on what changed, revise hypothesis status, and identify the next discriminating check.
- When findings need context, pair this reasoning with a literature-search skill such as
/polars-dovmed, then revise hypotheses against the literature.
- Apply the relevant checklist below.
- Structure output using the provided format.
Project Hypothesis Loop
Use this loop for omics projects, unexpected results, exploratory analyses, and any request that asks what results mean.
- Initialize: create at least 5 working hypotheses. Include biological mechanisms, technical artifacts, null explanations, sampling/batch effects, and annotation/database artifacts where relevant.
- Build an analysis playbook: search the literature for the inferred organism, virus group, data type, or closest lineage. Summarize typical analyses, comparison baselines, markers/features, plots, and outlier criteria used by scientists in that literature.
- Reflect: for every major intermediate result or QC gate, state the observation, QC status, strongest interpretation, remaining alternatives, and what evidence would separate them.
- Contextualize: compare findings against the playbook and run additional literature searches for central or unexpected findings. Use broad synonym-aware queries and cite DOI/PMCID when available.
- Revise: update each hypothesis as supported, weakened, ruled out, or unresolved. Keep ruled-out hypotheses visible with the evidence that changed their status.
- Replace: if fewer than 5 active hypotheses remain during exploratory work, add plausible replacements or explicitly state why no additional plausible alternatives exist.
Critique Checklist
Use relevant sections based on the review scope. Skip items not applicable to the study type.
## Methodology
- [ ] Design matches research question (causal claim → RCT needed)
- [ ] Sample size justified (power analysis reported)
- [ ] Randomization/blinding implemented where feasible
- [ ] Confounders identified and controlled
- [ ] Measurements validated and reliable
## Statistics
- [ ] Tests appropriate for data type
- [ ] Assumptions checked
- [ ] Multiple comparisons corrected
- [ ] Effect sizes + CIs reported (not just p-values)
- [ ] Missing data handled appropriately
## Interpretation
- [ ] Conclusions match evidence strength
- [ ] Limitations acknowledged
- [ ] Causal claims only from experimental designs
- [ ] No cherry-picking or overgeneralization
## Red Flags
- [ ] P-values clustered just below .05
- [ ] Outcomes differ from registration
- [ ] Correlation presented as causation
- [ ] Subgroups analyzed without preregistration
Claim Assessment
- Identify claim type (causal, associational, descriptive).
- Match evidence to claim type.
- Check logical connection between data and conclusion.
- Ensure confidence matches evidence strength.
Claim strength ladder:
| Language |
Requires |
| "Proves" / "Demonstrates" |
Strong experimental evidence |
| "Suggests" / "Indicates" |
Observational with controlled confounds |
| "Associated with" |
Observational, no causal claim |
| "May" / "Might" |
Preliminary or hypothesis-generating |
Output Format
## Summary
[1-2 sentences: What was studied and main finding]
## Strengths
- [Specific methodological strengths]
## Concerns
### Critical (threaten main conclusions)
- [Issue + why it matters]
### Important (affect interpretation)
- [Issue + why it matters]
### Minor (worth noting)
- [Issue]
## Evidence Rating
[GRADE level: High/Moderate/Low/Very Low with justification]
## Bottom Line
[What can/cannot be concluded from this evidence]
Project Reasoning Output Format
## Current Result
[Observed intermediate/final result and QC status]
## Hypothesis Register
| Rank | Hypothesis | Status | Evidence For | Evidence Against | Next Discriminating Check |
|------|------------|--------|--------------|------------------|---------------------------|
| 1 | [Hypothesis] | supported/weakened/ruled out/unresolved | [Evidence] | [Evidence] | [Test] |
| 2 | [Hypothesis] | ... | ... | ... | ... |
| 3 | [Hypothesis] | ... | ... | ... | ... |
| 4 | [Hypothesis] | ... | ... | ... | ... |
| 5 | [Hypothesis] | ... | ... | ... | ... |
## Literature Context
- [Paper or search result, DOI/PMCID if available]: supports/contradicts/narrows [hypothesis] because [reason].
## Analysis Playbook
- [Literature-derived analysis/comparison/outlier check and why it applies]
## Revision
[What changed since the previous register and why]
## Next Step
[Most informative analysis, control, or experiment]
Quick Reference
| Task |
Action |
| Review a paper |
Use the Critique Checklist below |
| Evaluate a claim |
Use Claim Assessment steps |
| Assess evidence strength |
See references/evidence.md |
| Identify biases |
See references/biases.md |
| Statistical errors |
See references/stats.md |
| Logical fallacies |
See references/fallacies.md |
| Study design review |
See references/design.md |
Input Requirements
- Manuscript or study summary
- Claim statements to evaluate
- Methods and results (as available)
Output
- Structured critique with strengths, concerns, and evidence rating
- Clear recommendation on claim strength and limitations
Quality Gates
Examples
Example 1: Quick Paper Critique
Input: "Review this RCT on vitamin D and depression"
Output:
## Summary
RCT of 200 adults testing vitamin D supplementation for depression over 12 weeks.
## Strengths
- Randomized, double-blind, placebo-controlled
- Validated depression scale (PHQ-9)
- Adequate power analysis
## Concerns
### Critical
- 35% dropout, differential by group (attrition bias)
- ITT analysis not performed
### Important
- Single-site limits generalizability
## Evidence Rating
Moderate (downgraded from high due to attrition bias)
## Bottom Line
Suggestive but not conclusive due to differential attrition.
Example 2: Claim Assessment
Input: "This study proves that coffee prevents Alzheimer's"
Assessment: If based on observational data, this is a correlation→causation error. Appropriate language: "Coffee consumption is associated with lower Alzheimer's risk."
Troubleshooting
Issue: Insufficient methods details
Solution: Request missing design/statistics information before rating evidence.
Issue: Conflicting results across studies
Solution: Report uncertainty and suggest stronger study designs for resolution.
Related Skills
/ai-scientist-evaluator — rubric-scored evaluation of AI scientist outputs
/manuscript-review-council — multi-reviewer critique of peer-reviewed manuscripts
/proposal-review — structured critique of funding proposals
1---2name: bio-logic3description: Evaluate scientific rigor, methods, biases, and evidence quality for claims, papers, and study designs.4---5
6# Bio-Logic: Scientific Reasoning Evaluation
7
8Use structured frameworks to evaluate scientific claims, methodology, and evidence strength.
9
10## Instructions
11
121. Identify the task (claim assessment, paper critique, study design review, project interpretation, or hypothesis revision).
132. For project work, maintain a hypothesis register with at least 5 distinct active hypotheses until the project is no longer exploratory.
143. After each major intermediate result, reflect on what changed, revise hypothesis status, and identify the next discriminating check.
154. When findings need context, pair this reasoning with a literature-search skill such as `/polars-dovmed`, then revise hypotheses against the literature.
165. Apply the relevant checklist below.
176. Structure output using the provided format.
18
19### Project Hypothesis Loop
20
21Use this loop for omics projects, unexpected results, exploratory analyses, and any request that asks what results mean.
22
231. **Initialize**: create at least 5 working hypotheses. Include biological mechanisms, technical artifacts, null explanations, sampling/batch effects, and annotation/database artifacts where relevant.
242. **Build an analysis playbook**: search the literature for the inferred organism, virus group, data type, or closest lineage. Summarize typical analyses, comparison baselines, markers/features, plots, and outlier criteria used by scientists in that literature.
253. **Reflect**: for every major intermediate result or QC gate, state the observation, QC status, strongest interpretation, remaining alternatives, and what evidence would separate them.
264. **Contextualize**: compare findings against the playbook and run additional literature searches for central or unexpected findings. Use broad synonym-aware queries and cite DOI/PMCID when available.
275. **Revise**: update each hypothesis as supported, weakened, ruled out, or unresolved. Keep ruled-out hypotheses visible with the evidence that changed their status.
286. **Replace**: if fewer than 5 active hypotheses remain during exploratory work, add plausible replacements or explicitly state why no additional plausible alternatives exist.
29
30### Critique Checklist
31
32Use relevant sections based on the review scope. Skip items not applicable to the study type.
33
34```
35## Methodology
36- [ ] Design matches research question (causal claim → RCT needed)
37- [ ] Sample size justified (power analysis reported)
38- [ ] Randomization/blinding implemented where feasible
39- [ ] Confounders identified and controlled
40- [ ] Measurements validated and reliable
41
42## Statistics
43- [ ] Tests appropriate for data type
44- [ ] Assumptions checked
45- [ ] Multiple comparisons corrected
46- [ ] Effect sizes + CIs reported (not just p-values)
47- [ ] Missing data handled appropriately
48
49## Interpretation
50- [ ] Conclusions match evidence strength
51- [ ] Limitations acknowledged
52- [ ] Causal claims only from experimental designs
53- [ ] No cherry-picking or overgeneralization
54
55## Red Flags
56- [ ] P-values clustered just below .05
57- [ ] Outcomes differ from registration
58- [ ] Correlation presented as causation
59- [ ] Subgroups analyzed without preregistration
60```
61
62### Claim Assessment
63
641. Identify claim type (causal, associational, descriptive).
652. Match evidence to claim type.
663. Check logical connection between data and conclusion.
674. Ensure confidence matches evidence strength.
68
69**Claim strength ladder**:
70| Language | Requires |
71|----------|----------|
72| "Proves" / "Demonstrates" | Strong experimental evidence |
73| "Suggests" / "Indicates" | Observational with controlled confounds |
74| "Associated with" | Observational, no causal claim |
75| "May" / "Might" | Preliminary or hypothesis-generating |
76
77### Output Format
78
79```markdown
80## Summary
81[1-2 sentences: What was studied and main finding]
82
83## Strengths
84- [Specific methodological strengths]
85
86## Concerns
87### Critical (threaten main conclusions)
88- [Issue + why it matters]
89
90### Important (affect interpretation)
91- [Issue + why it matters]
92
93### Minor (worth noting)
94- [Issue]
95
96## Evidence Rating
97[GRADE level: High/Moderate/Low/Very Low with justification]
98
99## Bottom Line
100[What can/cannot be concluded from this evidence]
101```
102
103### Project Reasoning Output Format
104
105```markdown
106## Current Result
107[Observed intermediate/final result and QC status]
108
109## Hypothesis Register
110| Rank | Hypothesis | Status | Evidence For | Evidence Against | Next Discriminating Check |
111|------|------------|--------|--------------|------------------|---------------------------|
112| 1 | [Hypothesis] | supported/weakened/ruled out/unresolved | [Evidence] | [Evidence] | [Test] |
113| 2 | [Hypothesis] | ... | ... | ... | ... |
114| 3 | [Hypothesis] | ... | ... | ... | ... |
115| 4 | [Hypothesis] | ... | ... | ... | ... |
116| 5 | [Hypothesis] | ... | ... | ... | ... |
117
118## Literature Context
119- [Paper or search result, DOI/PMCID if available]: supports/contradicts/narrows [hypothesis] because [reason].
120
121## Analysis Playbook
122- [Literature-derived analysis/comparison/outlier check and why it applies]
123
124## Revision
125[What changed since the previous register and why]
126
127## Next Step
128[Most informative analysis, control, or experiment]
129```
130
131## Quick Reference
132
133| Task | Action |
134|------|--------|
135| Review a paper | Use the Critique Checklist below |
136| Evaluate a claim | Use Claim Assessment steps |
137| Assess evidence strength | See `references/evidence.md` |
138| Identify biases | See `references/biases.md` |
139| Statistical errors | See `references/stats.md` |
140| Logical fallacies | See `references/fallacies.md` |
141| Study design review | See `references/design.md` |
142
143## Input Requirements
144
145- Manuscript or study summary
146- Claim statements to evaluate
147- Methods and results (as available)
148
149## Output
150
151- Structured critique with strengths, concerns, and evidence rating
152- Clear recommendation on claim strength and limitations
153
154## Quality Gates
155
156- [ ] Claim type matches evidence type
157- [ ] Key methodological risks are identified
158- [ ] Evidence rating is justified with concrete issues
159- [ ] Exploratory project work includes at least 5 active hypotheses or explains why fewer are plausible
160- [ ] Discovery work includes a literature-derived analysis playbook before selecting analyses
161- [ ] Intermediate results update hypothesis status before the next workflow step
162- [ ] Central or unexpected findings are checked against the literature before final interpretation
163
164## Examples
165
166### Example 1: Quick Paper Critique
167
168**Input**: "Review this RCT on vitamin D and depression"
169
170**Output**:
171```
172## Summary
173RCT of 200 adults testing vitamin D supplementation for depression over 12 weeks.
174
175## Strengths
176- Randomized, double-blind, placebo-controlled
177- Validated depression scale (PHQ-9)
178- Adequate power analysis
179
180## Concerns
181### Critical
182- 35% dropout, differential by group (attrition bias)
183- ITT analysis not performed
184
185### Important
186- Single-site limits generalizability
187
188## Evidence Rating
189Moderate (downgraded from high due to attrition bias)
190
191## Bottom Line
192Suggestive but not conclusive due to differential attrition.
193```
194
195### Example 2: Claim Assessment
196
197**Input**: "This study proves that coffee prevents Alzheimer's"
198
199**Assessment**: If based on observational data, this is a correlation→causation error. Appropriate language: "Coffee consumption is associated with lower Alzheimer's risk."
200
201## Troubleshooting
202
203**Issue**: Insufficient methods details
204**Solution**: Request missing design/statistics information before rating evidence.
205
206**Issue**: Conflicting results across studies
207**Solution**: Report uncertainty and suggest stronger study designs for resolution.
208
209## Related Skills
210
211- `/ai-scientist-evaluator` — rubric-scored evaluation of AI scientist outputs
212- `/manuscript-review-council` — multi-reviewer critique of peer-reviewed manuscripts
213- `/proposal-review` — structured critique of funding proposals