Bio-Logic: Scientific Reasoning Evaluation
Use structured frameworks to evaluate scientific claims, methodology, and evidence strength.
Instructions
- Identify the task (claim assessment, paper critique, study design review, project interpretation, or hypothesis revision).
- For project work, maintain a hypothesis register with at least 5 distinct active hypotheses until the project is no longer exploratory.
- After each major intermediate result, reflect on what changed, revise hypothesis status, and identify the next discriminating check.
- When findings need context, pair this reasoning with a literature-search skill such as
/polars-dovmed, then revise hypotheses against the literature.
- Select the study-type profile before applying a checklist. Do not grade a
computational benchmark, phylogenetic analysis, or evolutionary inference
as if it were an intervention trial.
- Structure output using the provided format.
Study-Type Profiles
| Profile |
Main evidence checks |
Rating language |
| Intervention |
allocation, controls, adherence, attrition, estimand, harms |
GRADE when the review question and evidence synthesis support it |
| Observational or quasi-experimental |
identification assumptions, temporality, confounding, negative controls, sensitivity analyses |
risk-of-bias plus causal-confidence statement |
| Computational or machine learning |
data provenance, leakage, splits, baselines, calibration, external validation, reproducibility |
computational evidence confidence |
| Evolutionary or comparative genomics |
orthology, taxon/reference sampling, model fit, support, contamination, topology sensitivity |
phylogenetic/comparative evidence confidence |
| Descriptive or exploratory omics |
sampling, measurement, multiple testing, effect sizes, replication, alternative explanations |
descriptive confidence; avoid causal labels |
Project Hypothesis Loop
Use this loop for omics projects, unexpected results, exploratory analyses, and any request that asks what results mean.
- Initialize: create at least 5 working hypotheses. Include biological mechanisms, technical artifacts, null explanations, sampling/batch effects, and annotation/database artifacts where relevant.
- Build an analysis playbook: search the literature for the inferred organism, virus group, data type, or closest lineage. Summarize typical analyses, comparison baselines, markers/features, plots, and outlier criteria used by scientists in that literature.
- Reflect: for every major intermediate result or QC gate, state the observation, QC status, strongest interpretation, remaining alternatives, and what evidence would separate them.
- Contextualize: compare findings against the playbook and run additional literature searches for central or unexpected findings. Use broad synonym-aware queries and cite DOI/PMCID when available.
- Revise: update each hypothesis as supported, weakened, ruled out, or unresolved. Keep ruled-out hypotheses visible with the evidence that changed their status.
- Replace: if fewer than 5 active hypotheses remain during exploratory work, add plausible replacements or explicitly state why no additional plausible alternatives exist.
Critique Checklist
Use relevant sections based on the review scope. Skip items not applicable to the study type.
## Methodology
- [ ] Design identifies the causal estimand through randomization or a justified quasi-experimental or natural-experiment strategy
- [ ] Sample size justified (power analysis reported)
- [ ] Randomization/blinding implemented where feasible
- [ ] Confounders identified and controlled
- [ ] Measurements validated and reliable
## Statistics
- [ ] Tests appropriate for data type
- [ ] Assumptions checked
- [ ] Multiple comparisons corrected
- [ ] Effect sizes + CIs reported (not just p-values)
- [ ] Missing data handled appropriately
## Interpretation
- [ ] Conclusions match evidence strength
- [ ] Limitations acknowledged
- [ ] Causal claims require an identified causal design: randomized evidence or
a justified quasi-experimental/natural-experiment strategy with its
assumptions and sensitivity checks
- [ ] No cherry-picking or overgeneralization
## Red Flags
- [ ] P-values clustered just below .05
- [ ] Outcomes differ from registration
- [ ] Correlation presented as causation
- [ ] Subgroups analyzed without preregistration
Claim Assessment
- Identify claim type (causal, associational, descriptive).
- Match evidence to claim type.
- Check logical connection between data and conclusion.
- Ensure confidence matches evidence strength.
Claim strength ladder:
| Language |
Requires |
| "Demonstrates" |
Strong evidence under a design that identifies the target claim |
| "Suggests" / "Indicates" |
Observational with controlled confounds |
| "Associated with" |
Observational, no causal claim |
| "May" / "Might" |
Preliminary or hypothesis-generating |
Output Format
## Summary
[1-2 sentences: What was studied and main finding]
## Strengths
- [Specific methodological strengths]
## Concerns
### Critical (threaten main conclusions)
- [Issue + why it matters]
### Important (affect interpretation)
- [Issue + why it matters]
### Minor (worth noting)
- [Issue]
## Evidence Rating
[Framework appropriate to the study profile, rating, and justification. Use
GRADE only when it applies to the review question.]
## Bottom Line
[What can/cannot be concluded from this evidence]
Project Reasoning Output Format
## Current Result
[Observed intermediate/final result and QC status]
## Hypothesis Register
| Rank | Hypothesis | Status | Evidence For | Evidence Against | Next Discriminating Check |
|------|------------|--------|--------------|------------------|---------------------------|
| 1 | [Hypothesis] | supported/weakened/ruled out/unresolved | [Evidence] | [Evidence] | [Test] |
| 2 | [Hypothesis] | ... | ... | ... | ... |
| 3 | [Hypothesis] | ... | ... | ... | ... |
| 4 | [Hypothesis] | ... | ... | ... | ... |
| 5 | [Hypothesis] | ... | ... | ... | ... |
## Literature Context
- [Paper or search result, DOI/PMCID if available]: supports/contradicts/narrows [hypothesis] because [reason].
## Analysis Playbook
- [Literature-derived analysis/comparison/outlier check and why it applies]
## Revision
[What changed since the previous register and why]
## Next Step
[Most informative analysis, control, or experiment]
Quick Reference
| Task |
Action |
| Review a paper |
Use the Critique Checklist below |
| Evaluate a claim |
Use Claim Assessment steps |
| Assess evidence strength |
See references/evidence.md |
| Identify biases |
See references/biases.md |
| Statistical errors |
See references/stats.md |
| Logical fallacies |
See references/fallacies.md |
| Study design review |
See references/design.md |
Input Requirements
- Manuscript or study summary
- Claim statements to evaluate
- Methods and results (as available)
Output
- Structured critique with strengths, concerns, and evidence rating
- Clear recommendation on claim strength and limitations
Quality Gates
Examples
Example 1: Quick Paper Critique
Input: "Review this RCT on vitamin D and depression"
Output:
## Summary
RCT of 200 adults testing vitamin D supplementation for depression over 12 weeks.
## Strengths
- Randomized, double-blind, placebo-controlled
- Validated depression scale (PHQ-9)
- Adequate power analysis
## Concerns
### Critical
- 35% dropout, differential by group (attrition bias)
- ITT analysis not performed
### Important
- Single-site limits generalizability
## Evidence Rating
Moderate (downgraded from high due to attrition bias)
## Bottom Line
Suggestive but not conclusive due to differential attrition.
Example 2: Claim Assessment
Input: "This study proves that coffee prevents Alzheimer's"
Assessment: If based only on an association with no identified causal design, this is a correlation→causation error. Appropriate language: "Coffee consumption is associated with lower Alzheimer's risk." Observational data with a justified quasi-experimental or natural-experiment strategy, its assumptions stated and tested, can support a causal claim; the design is what decides, not whether the data were observational.
Troubleshooting
Issue: Insufficient methods details
Solution: Request missing design/statistics information before rating evidence.
Issue: Conflicting results across studies
Solution: Report uncertainty and suggest stronger study designs for resolution.
Related Skills
/ai-scientist-evaluator — rubric-scored evaluation of AI scientist outputs
/manuscript-review-council — multi-reviewer critique of peer-reviewed manuscripts
/proposal-review — structured critique of funding proposals
1---2name: bio-logic3description: Assess a scientific claim, study design, method, or interpretation against its evidence. Use when testing causal reasoning, finding methodological bias, weighing alternative explanations, or revising hypotheses.4---56# Bio-Logic: Scientific Reasoning Evaluation78Use structured frameworks to evaluate scientific claims, methodology, and evidence strength.910## Instructions11121. Identify the task (claim assessment, paper critique, study design review, project interpretation, or hypothesis revision).132. For project work, maintain a hypothesis register with at least 5 distinct active hypotheses until the project is no longer exploratory.143. After each major intermediate result, reflect on what changed, revise hypothesis status, and identify the next discriminating check.154. When findings need context, pair this reasoning with a literature-search skill such as `/polars-dovmed`, then revise hypotheses against the literature.165. Select the study-type profile before applying a checklist. Do not grade a17 computational benchmark, phylogenetic analysis, or evolutionary inference18 as if it were an intervention trial.196. Structure output using the provided format.2021### Study-Type Profiles2223| Profile | Main evidence checks | Rating language |24|---|---|---|25| Intervention | allocation, controls, adherence, attrition, estimand, harms | GRADE when the review question and evidence synthesis support it |26| Observational or quasi-experimental | identification assumptions, temporality, confounding, negative controls, sensitivity analyses | risk-of-bias plus causal-confidence statement |27| Computational or machine learning | data provenance, leakage, splits, baselines, calibration, external validation, reproducibility | computational evidence confidence |28| Evolutionary or comparative genomics | orthology, taxon/reference sampling, model fit, support, contamination, topology sensitivity | phylogenetic/comparative evidence confidence |29| Descriptive or exploratory omics | sampling, measurement, multiple testing, effect sizes, replication, alternative explanations | descriptive confidence; avoid causal labels |3031### Project Hypothesis Loop3233Use this loop for omics projects, unexpected results, exploratory analyses, and any request that asks what results mean.34351. **Initialize**: create at least 5 working hypotheses. Include biological mechanisms, technical artifacts, null explanations, sampling/batch effects, and annotation/database artifacts where relevant.362. **Build an analysis playbook**: search the literature for the inferred organism, virus group, data type, or closest lineage. Summarize typical analyses, comparison baselines, markers/features, plots, and outlier criteria used by scientists in that literature.373. **Reflect**: for every major intermediate result or QC gate, state the observation, QC status, strongest interpretation, remaining alternatives, and what evidence would separate them.384. **Contextualize**: compare findings against the playbook and run additional literature searches for central or unexpected findings. Use broad synonym-aware queries and cite DOI/PMCID when available.395. **Revise**: update each hypothesis as supported, weakened, ruled out, or unresolved. Keep ruled-out hypotheses visible with the evidence that changed their status.406. **Replace**: if fewer than 5 active hypotheses remain during exploratory work, add plausible replacements or explicitly state why no additional plausible alternatives exist.4142### Critique Checklist4344Use relevant sections based on the review scope. Skip items not applicable to the study type.4546```47## Methodology48- [ ] Design identifies the causal estimand through randomization or a justified quasi-experimental or natural-experiment strategy49- [ ] Sample size justified (power analysis reported)50- [ ] Randomization/blinding implemented where feasible51- [ ] Confounders identified and controlled52- [ ] Measurements validated and reliable5354## Statistics55- [ ] Tests appropriate for data type56- [ ] Assumptions checked57- [ ] Multiple comparisons corrected58- [ ] Effect sizes + CIs reported (not just p-values)59- [ ] Missing data handled appropriately6061## Interpretation62- [ ] Conclusions match evidence strength63- [ ] Limitations acknowledged64- [ ] Causal claims require an identified causal design: randomized evidence or65 a justified quasi-experimental/natural-experiment strategy with its66 assumptions and sensitivity checks67- [ ] No cherry-picking or overgeneralization6869## Red Flags70- [ ] P-values clustered just below .0571- [ ] Outcomes differ from registration72- [ ] Correlation presented as causation73- [ ] Subgroups analyzed without preregistration74```7576### Claim Assessment77781. Identify claim type (causal, associational, descriptive).792. Match evidence to claim type.803. Check logical connection between data and conclusion.814. Ensure confidence matches evidence strength.8283**Claim strength ladder**:84| Language | Requires |85|----------|----------|86| "Demonstrates" | Strong evidence under a design that identifies the target claim |87| "Suggests" / "Indicates" | Observational with controlled confounds |88| "Associated with" | Observational, no causal claim |89| "May" / "Might" | Preliminary or hypothesis-generating |9091### Output Format9293```markdown94## Summary95[1-2 sentences: What was studied and main finding]9697## Strengths98- [Specific methodological strengths]99100## Concerns101### Critical (threaten main conclusions)102- [Issue + why it matters]103104### Important (affect interpretation)105- [Issue + why it matters]106107### Minor (worth noting)108- [Issue]109110## Evidence Rating111[Framework appropriate to the study profile, rating, and justification. Use112GRADE only when it applies to the review question.]113114## Bottom Line115[What can/cannot be concluded from this evidence]116```117118### Project Reasoning Output Format119120```markdown121## Current Result122[Observed intermediate/final result and QC status]123124## Hypothesis Register125| Rank | Hypothesis | Status | Evidence For | Evidence Against | Next Discriminating Check |126|------|------------|--------|--------------|------------------|---------------------------|127| 1 | [Hypothesis] | supported/weakened/ruled out/unresolved | [Evidence] | [Evidence] | [Test] |128| 2 | [Hypothesis] | ... | ... | ... | ... |129| 3 | [Hypothesis] | ... | ... | ... | ... |130| 4 | [Hypothesis] | ... | ... | ... | ... |131| 5 | [Hypothesis] | ... | ... | ... | ... |132133## Literature Context134- [Paper or search result, DOI/PMCID if available]: supports/contradicts/narrows [hypothesis] because [reason].135136## Analysis Playbook137- [Literature-derived analysis/comparison/outlier check and why it applies]138139## Revision140[What changed since the previous register and why]141142## Next Step143[Most informative analysis, control, or experiment]144```145146## Quick Reference147148| Task | Action |149|------|--------|150| Review a paper | Use the Critique Checklist below |151| Evaluate a claim | Use Claim Assessment steps |152| Assess evidence strength | See `references/evidence.md` |153| Identify biases | See `references/biases.md` |154| Statistical errors | See `references/stats.md` |155| Logical fallacies | See `references/fallacies.md` |156| Study design review | See `references/design.md` |157158## Input Requirements159160- Manuscript or study summary161- Claim statements to evaluate162- Methods and results (as available)163164## Output165166- Structured critique with strengths, concerns, and evidence rating167- Clear recommendation on claim strength and limitations168169## Quality Gates170171- [ ] Claim type matches evidence type172- [ ] Key methodological risks are identified173- [ ] Evidence rating is justified with concrete issues174- [ ] The evidence framework matches the study type; computational and175 evolutionary work is not forced into intervention-centric GRADE rules176- [ ] Exploratory project work includes at least 5 active hypotheses or explains why fewer are plausible177- [ ] Discovery work includes a literature-derived analysis playbook before selecting analyses178- [ ] Intermediate results update hypothesis status before the next workflow step179- [ ] Central or unexpected findings are checked against the literature before final interpretation180181## Examples182183### Example 1: Quick Paper Critique184185**Input**: "Review this RCT on vitamin D and depression"186187**Output**:188```189## Summary190RCT of 200 adults testing vitamin D supplementation for depression over 12 weeks.191192## Strengths193- Randomized, double-blind, placebo-controlled194- Validated depression scale (PHQ-9)195- Adequate power analysis196197## Concerns198### Critical199- 35% dropout, differential by group (attrition bias)200- ITT analysis not performed201202### Important203- Single-site limits generalizability204205## Evidence Rating206Moderate (downgraded from high due to attrition bias)207208## Bottom Line209Suggestive but not conclusive due to differential attrition.210```211212### Example 2: Claim Assessment213214**Input**: "This study proves that coffee prevents Alzheimer's"215216**Assessment**: If based only on an association with no identified causal design, this is a correlation→causation error. Appropriate language: "Coffee consumption is associated with lower Alzheimer's risk." Observational data with a justified quasi-experimental or natural-experiment strategy, its assumptions stated and tested, can support a causal claim; the design is what decides, not whether the data were observational.217218## Troubleshooting219220**Issue**: Insufficient methods details221**Solution**: Request missing design/statistics information before rating evidence.222223**Issue**: Conflicting results across studies224**Solution**: Report uncertainty and suggest stronger study designs for resolution.225226## Related Skills227228- `/ai-scientist-evaluator` — rubric-scored evaluation of AI scientist outputs229- `/manuscript-review-council` — multi-reviewer critique of peer-reviewed manuscripts230- `/proposal-review` — structured critique of funding proposals