Scientific Critical Thinking
Overview
Critical thinking is a systematic process for evaluating scientific rigor. Assess methodology, experimental design, statistical validity, biases, confounding, and evidence quality using GRADE and Cochrane ROB frameworks. Apply this skill for critical analysis of scientific claims.
When to Use This Skill
This skill should be used when:
- Evaluating research methodology and experimental design
- Assessing statistical validity and evidence quality
- Identifying biases and confounding in studies
- Reviewing scientific claims and conclusions
- Conducting systematic reviews or meta-analyses
- Applying GRADE or Cochrane risk of bias assessments
- Providing critical analysis of research papers
Core Capabilities
1. Methodology Critique
Evaluate research methodology for rigor, validity, and potential flaws.
Apply when:
- Reviewing research papers
- Assessing experimental designs
- Evaluating study protocols
- Planning new research
Evaluation framework:
Study Design Assessment
- Is the design appropriate for the research question?
- Can the design support causal claims being made?
- Are comparison groups appropriate and adequate?
- Consider whether experimental, quasi-experimental, or observational design is justified
Validity Analysis
- Internal validity: Can we trust the causal inference?
- Check randomization quality
- Evaluate confounding control
- Assess selection bias
- Review attrition/dropout patterns
- External validity: Do results generalize?
- Evaluate sample representativeness
- Consider ecological validity of setting
- Assess whether conditions match target application
- Construct validity: Do measures capture intended constructs?
- Review measurement validation
- Check operational definitions
- Assess whether measures are direct or proxy
- Statistical conclusion validity: Are statistical inferences sound?
- Verify adequate power/sample size
- Check assumption compliance
- Evaluate test appropriateness
Control and Blinding
- Was randomization properly implemented (sequence generation, allocation concealment)?
- Was blinding feasible and implemented (participants, providers, assessors)?
- Are control conditions appropriate (placebo, active control, no treatment)?
- Could performance or detection bias affect results?
Measurement Quality
- Are instruments validated and reliable?
- Are measures objective when possible, or subjective with acknowledged limitations?
- Is outcome assessment standardized?
- Are multiple measures used to triangulate findings?
Reference: See references/scientific_method.md for detailed principles and references/experimental_design.md for comprehensive design checklist.
2. Bias Detection
Identify and evaluate potential sources of bias that could distort findings.
Apply when:
- Reviewing published research
- Designing new studies
- Interpreting conflicting evidence
- Assessing research quality
Systematic bias review spans five families: cognitive biases (confirmation, HARKing, publication bias, cherry-picking), selection biases (sampling, volunteer, attrition, survivorship), measurement biases (observer, recall, social desirability, instrument), analysis biases (p-hacking, outcome switching, selective reporting, subgroup fishing), and confounding (measured, unmeasured, and alternative explanations).
Reference: Full detection-question checklist for every bias in each family moved to references/evaluation_checklists.md (Bias Detection section). See references/common_biases.md for the comprehensive bias taxonomy with mitigation strategies.
3. Statistical Analysis Evaluation
Critically assess statistical methods, interpretation, and reporting.
Apply when:
- Reviewing quantitative research
- Evaluating data-driven claims
- Assessing clinical trial results
- Reviewing meta-analyses
Statistical review checklist covers eight areas: sample size and power, statistical test appropriateness, multiple-comparison correction, p-value interpretation, effect sizes and confidence intervals, missing-data handling, regression/modeling assumptions, and common pitfalls (correlation-as-causation, regression to the mean, base-rate neglect, Texas sharpshooter, Simpson's paradox).
Reference: Full question-by-question checklist for all eight areas moved to references/evaluation_checklists.md (Statistical Analysis Evaluation section). See references/statistical_pitfalls.md for detailed pitfalls and correct practices.
4. Evidence Quality Assessment
Evaluate the strength and quality of evidence systematically.
Apply when:
- Weighing evidence for decisions
- Conducting literature reviews
- Comparing conflicting findings
- Determining confidence in conclusions
Evidence evaluation framework works through five components: the study-design hierarchy (meta-analysis down to expert opinion, with the caveat that level does not guarantee quality), quality within a design type (risk-of-bias tools, rigor, transparency, conflicts of interest), GRADE considerations (start from design type, then downgrade/upgrade), convergence of evidence (what makes it stronger vs. weaker), and contextual factors (plausibility, consistency, temporality, specificity, strength).
Reference: Full framework with the complete GRADE downgrade/upgrade criteria and convergence signals moved to references/evaluation_checklists.md (Evidence Quality Assessment section). See references/evidence_hierarchy.md for the detailed hierarchy, GRADE system, and quality assessment tools.
5. Logical Fallacy Identification
Detect and name logical errors in scientific arguments and claims.
Apply when:
- Evaluating scientific claims
- Reviewing discussion/conclusion sections
- Assessing popular science communication
- Identifying flawed reasoning
Common fallacies in science fall into six categories: causation fallacies (post hoc, correlation = causation, reverse causation, single cause), generalization fallacies (hasty generalization, anecdotal, cherry-picking, ecological), authority/source fallacies (appeal to authority, ad hominem, genetic, appeal to nature), statistical fallacies (base-rate neglect, Texas sharpshooter, multiple comparisons, prosecutor's fallacy), structural fallacies (false dichotomy, moving goalposts, begging the question, straw man), and science-specific fallacies (Galileo gambit, argument from ignorance, nirvana fallacy, unfalsifiability).
Reference: Full fallacy catalog with the plain-language definition of each named fallacy moved to references/evaluation_checklists.md (Logical Fallacy Identification section). See references/logical_fallacies.md for examples and detection strategies.
When identifying fallacies:
- Name the specific fallacy
- Explain why the reasoning is flawed
- Identify what evidence would be needed for valid inference
- Note that fallacious reasoning doesn't prove the conclusion false—just that this argument doesn't support it
Reference: See references/logical_fallacies.md for comprehensive fallacy catalog with examples and detection strategies.
6. Research Design Guidance
Provide constructive guidance for planning rigorous studies.
Apply when:
- Helping design new experiments
- Planning research projects
- Reviewing research proposals
- Improving study protocols
Design process:
Research Question Refinement
- Ensure question is specific, answerable, and falsifiable
- Verify it addresses a gap or contradiction in literature
- Confirm feasibility (resources, ethics, time)
- Define variables operationally
Design Selection
- Match design to question (causal → experimental; associational → observational)
- Consider feasibility and ethical constraints
- Choose between-subjects, within-subjects, or mixed designs
- Plan factorial designs if testing multiple factors
Bias Minimization Strategy
- Implement randomization when possible
- Plan blinding at all feasible levels (participants, providers, assessors)
- Identify and plan to control confounds (randomization, matching, stratification, statistical adjustment)
- Standardize all procedures
- Plan to minimize attrition
Sample Planning
- Conduct a priori power analysis (specify expected effect, desired power, alpha)
- Account for attrition in sample size
- Define clear inclusion/exclusion criteria
- Consider recruitment strategy and feasibility
- Plan for sample representativeness
Measurement Strategy
- Select validated, reliable instruments
- Use objective measures when possible
- Plan multiple measures of key constructs (triangulation)
- Ensure measures are sensitive to expected changes
- Establish inter-rater reliability procedures
Analysis Planning
- Prespecify all hypotheses and analyses
- Designate primary outcome clearly
- Plan statistical tests with assumption checks
- Specify how missing data will be handled
- Plan to report effect sizes and confidence intervals
- Consider multiple comparison corrections
Transparency and Rigor
- Preregister study and analysis plan
- Use reporting guidelines (CONSORT, STROBE, PRISMA)
- Plan to report all outcomes, not just significant ones
- Distinguish confirmatory from exploratory analyses
- Commit to data/code sharing
Reference: See references/experimental_design.md for comprehensive design checklist covering all stages from question to dissemination.
7. Claim Evaluation
Systematically evaluate scientific claims for validity and support.
Apply when:
- Assessing conclusions in papers
- Evaluating media reports of research
- Reviewing abstract or introduction claims
- Checking if data support conclusions
Claim evaluation process:
Identify the Claim
- What exactly is being claimed?
- Is it a causal claim, associational claim, or descriptive claim?
- How strong is the claim (proven, likely, suggested, possible)?
Assess the Evidence
- What evidence is provided?
- Is evidence direct or indirect?
- Is evidence sufficient for the strength of claim?
- Are alternative explanations ruled out?
Check Logical Connection
- Do conclusions follow from the data?
- Are there logical leaps?
- Is correlational data used to support causal claims?
- Are limitations acknowledged?
Evaluate Proportionality
- Is confidence proportional to evidence strength?
- Are hedging words used appropriately?
- Are limitations downplayed?
- Is speculation clearly labeled?
Check for Overgeneralization
- Do claims extend beyond the sample studied?
- Are population restrictions acknowledged?
- Is context-dependence recognized?
- Are caveats about generalization included?
Red Flags
- Causal language from correlational studies
- "Proves" or absolute certainty
- Cherry-picked citations
- Ignoring contradictory evidence
- Dismissing limitations
- Extrapolation beyond data
Provide specific feedback:
- Quote the problematic claim
- Explain what evidence would be needed to support it
- Suggest appropriate hedging language if warranted
- Distinguish between data (what was found) and interpretation (what it means)
Application Guidelines
General Approach
Be Constructive
- Identify strengths as well as weaknesses
- Suggest improvements rather than just criticizing
- Distinguish between fatal flaws and minor limitations
- Recognize that all research has limitations
Be Specific
- Point to specific instances (e.g., "Table 2 shows..." or "In the Methods section...")
- Quote problematic statements
- Provide concrete examples of issues
- Reference specific principles or standards violated
Be Proportionate
- Match criticism severity to issue importance
- Distinguish between major threats to validity and minor concerns
- Consider whether issues affect primary conclusions
- Acknowledge uncertainty in your own assessments
Apply Consistent Standards
- Use same criteria across all studies
- Don't apply stricter standards to findings you dislike
- Acknowledge your own potential biases
- Base judgments on methodology, not results
Consider Context
- Acknowledge practical and ethical constraints
- Consider field-specific norms for effect sizes and methods
- Recognize exploratory vs. confirmatory contexts
- Account for resource limitations in evaluating studies
When Providing Critique
Structure feedback as:
- Summary: Brief overview of what was evaluated
- Strengths: What was done well (important for credibility and learning)
- Concerns: Issues organized by severity
- Critical issues (threaten validity of main conclusions)
- Important issues (affect interpretation but not fatally)
- Minor issues (worth noting but don't change conclusions)
- Specific Recommendations: Actionable suggestions for improvement
- Overall Assessment: Balanced conclusion about evidence quality and what can be concluded
Use precise terminology:
- Name specific biases, fallacies, and methodological issues
- Reference established standards and guidelines
- Cite principles from scientific methodology
- Use technical terms accurately
When Uncertain
- Acknowledge uncertainty: "This could be X or Y; additional information needed is Z"
- Ask clarifying questions: "Was [methodological detail] done? This affects interpretation."
- Provide conditional assessments: "If X was done, then Y follows; if not, then Z is concern"
- Note what additional information would resolve uncertainty
Reference Materials
This skill includes comprehensive reference materials that provide detailed frameworks for critical evaluation:
references/scientific_method.md - Core principles of scientific methodology, the scientific process, critical evaluation criteria, red flags in scientific claims, causal inference standards, peer review, and open science principles
references/common_biases.md - Comprehensive taxonomy of cognitive, experimental, methodological, statistical, and analysis biases with detection and mitigation strategies
references/statistical_pitfalls.md - Common statistical errors and misinterpretations including p-value misunderstandings, multiple comparisons problems, sample size issues, effect size mistakes, correlation/causation confusion, regression pitfalls, and meta-analysis issues
references/evidence_hierarchy.md - Traditional evidence hierarchy, GRADE system, study quality assessment criteria, domain-specific considerations, evidence synthesis principles, and practical decision frameworks
references/logical_fallacies.md - Logical fallacies common in scientific discourse organized by type (causation, generalization, authority, relevance, structure, statistical) with examples and detection strategies
references/experimental_design.md - Comprehensive experimental design checklist covering research questions, hypotheses, study design selection, variables, sampling, blinding, randomization, control groups, procedures, measurement, bias minimization, data management, statistical planning, ethical considerations, validity threats, and reporting standards
references/evaluation_checklists.md - Operational, question-form review checklists for the Bias Detection, Statistical Analysis Evaluation, Evidence Quality Assessment, and Logical Fallacy Identification capabilities (the full item-by-item detail relocated from the Core Capabilities section)
When to consult references:
- Load references into context when detailed frameworks are needed
- Use grep to search references for specific topics:
grep -r "pattern" references/
- References provide depth; SKILL.md provides procedural guidance
- Consult references for comprehensive lists, detailed criteria, and specific examples
Remember
Scientific critical thinking is about:
- Systematic evaluation using established principles
- Constructive critique that improves science
- Proportional confidence to evidence strength
- Transparency about uncertainty and limitations
- Consistent application of standards
- Recognition that all research has limitations
- Balance between skepticism and openness to evidence
Always distinguish between:
- Data (what was observed) and interpretation (what it means)
- Correlation and causation
- Statistical significance and practical importance
- Exploratory and confirmatory findings
- What is known and what is uncertain
- Evidence against a claim and evidence for the null
Goals of critical thinking:
- Identify strengths and weaknesses accurately
- Determine what conclusions are supported
- Recognize limitations and uncertainties
- Suggest improvements for future work
- Advance scientific understanding
1---2name: alterlab-scientific-thinking3description: Evaluate scientific claims and evidence quality using evidence grading frameworks (GRADE, Cochrane Risk of Bias), assessing experimental design validity and identifying biases, confounders, statistical pitfalls, and logical fallacies. Use when judging evidence quality, grading certainty of evidence, spotting design or causal-inference flaws, identifying biases or confounders, naming statistical fallacies, or teaching critical analysis. For writing a formal submittable peer review use alterlab-peer-review; for a multi-reviewer mock panel verdict use alterlab-paper-reviewer; for IRB/consent/conflict-of-interest ethics use alterlab-research-ethics. Part of the AlterLab Academic Skills suite.4license: MIT5---67# Scientific Critical Thinking89## Overview1011Critical thinking is a systematic process for evaluating scientific rigor. Assess methodology, experimental design, statistical validity, biases, confounding, and evidence quality using GRADE and Cochrane ROB frameworks. Apply this skill for critical analysis of scientific claims.1213## When to Use This Skill1415This skill should be used when:16- Evaluating research methodology and experimental design17- Assessing statistical validity and evidence quality18- Identifying biases and confounding in studies19- Reviewing scientific claims and conclusions20- Conducting systematic reviews or meta-analyses21- Applying GRADE or Cochrane risk of bias assessments22- Providing critical analysis of research papers2324## Core Capabilities2526### 1. Methodology Critique2728Evaluate research methodology for rigor, validity, and potential flaws.2930**Apply when:**31- Reviewing research papers32- Assessing experimental designs33- Evaluating study protocols34- Planning new research3536**Evaluation framework:**37381. **Study Design Assessment**39 - Is the design appropriate for the research question?40 - Can the design support causal claims being made?41 - Are comparison groups appropriate and adequate?42 - Consider whether experimental, quasi-experimental, or observational design is justified43442. **Validity Analysis**45 - **Internal validity:** Can we trust the causal inference?46 - Check randomization quality47 - Evaluate confounding control48 - Assess selection bias49 - Review attrition/dropout patterns50 - **External validity:** Do results generalize?51 - Evaluate sample representativeness52 - Consider ecological validity of setting53 - Assess whether conditions match target application54 - **Construct validity:** Do measures capture intended constructs?55 - Review measurement validation56 - Check operational definitions57 - Assess whether measures are direct or proxy58 - **Statistical conclusion validity:** Are statistical inferences sound?59 - Verify adequate power/sample size60 - Check assumption compliance61 - Evaluate test appropriateness62633. **Control and Blinding**64 - Was randomization properly implemented (sequence generation, allocation concealment)?65 - Was blinding feasible and implemented (participants, providers, assessors)?66 - Are control conditions appropriate (placebo, active control, no treatment)?67 - Could performance or detection bias affect results?68694. **Measurement Quality**70 - Are instruments validated and reliable?71 - Are measures objective when possible, or subjective with acknowledged limitations?72 - Is outcome assessment standardized?73 - Are multiple measures used to triangulate findings?7475**Reference:** See `references/scientific_method.md` for detailed principles and `references/experimental_design.md` for comprehensive design checklist.7677### 2. Bias Detection7879Identify and evaluate potential sources of bias that could distort findings.8081**Apply when:**82- Reviewing published research83- Designing new studies84- Interpreting conflicting evidence85- Assessing research quality8687**Systematic bias review** spans five families: cognitive biases (confirmation, HARKing, publication bias, cherry-picking), selection biases (sampling, volunteer, attrition, survivorship), measurement biases (observer, recall, social desirability, instrument), analysis biases (p-hacking, outcome switching, selective reporting, subgroup fishing), and confounding (measured, unmeasured, and alternative explanations).8889**Reference:** Full detection-question checklist for every bias in each family moved to `references/evaluation_checklists.md` (Bias Detection section). See `references/common_biases.md` for the comprehensive bias taxonomy with mitigation strategies.9091### 3. Statistical Analysis Evaluation9293Critically assess statistical methods, interpretation, and reporting.9495**Apply when:**96- Reviewing quantitative research97- Evaluating data-driven claims98- Assessing clinical trial results99- Reviewing meta-analyses100101**Statistical review checklist** covers eight areas: sample size and power, statistical test appropriateness, multiple-comparison correction, p-value interpretation, effect sizes and confidence intervals, missing-data handling, regression/modeling assumptions, and common pitfalls (correlation-as-causation, regression to the mean, base-rate neglect, Texas sharpshooter, Simpson's paradox).102103**Reference:** Full question-by-question checklist for all eight areas moved to `references/evaluation_checklists.md` (Statistical Analysis Evaluation section). See `references/statistical_pitfalls.md` for detailed pitfalls and correct practices.104105### 4. Evidence Quality Assessment106107Evaluate the strength and quality of evidence systematically.108109**Apply when:**110- Weighing evidence for decisions111- Conducting literature reviews112- Comparing conflicting findings113- Determining confidence in conclusions114115**Evidence evaluation framework** works through five components: the study-design hierarchy (meta-analysis down to expert opinion, with the caveat that level does not guarantee quality), quality within a design type (risk-of-bias tools, rigor, transparency, conflicts of interest), GRADE considerations (start from design type, then downgrade/upgrade), convergence of evidence (what makes it stronger vs. weaker), and contextual factors (plausibility, consistency, temporality, specificity, strength).116117**Reference:** Full framework with the complete GRADE downgrade/upgrade criteria and convergence signals moved to `references/evaluation_checklists.md` (Evidence Quality Assessment section). See `references/evidence_hierarchy.md` for the detailed hierarchy, GRADE system, and quality assessment tools.118119### 5. Logical Fallacy Identification120121Detect and name logical errors in scientific arguments and claims.122123**Apply when:**124- Evaluating scientific claims125- Reviewing discussion/conclusion sections126- Assessing popular science communication127- Identifying flawed reasoning128129**Common fallacies in science** fall into six categories: causation fallacies (post hoc, correlation = causation, reverse causation, single cause), generalization fallacies (hasty generalization, anecdotal, cherry-picking, ecological), authority/source fallacies (appeal to authority, ad hominem, genetic, appeal to nature), statistical fallacies (base-rate neglect, Texas sharpshooter, multiple comparisons, prosecutor's fallacy), structural fallacies (false dichotomy, moving goalposts, begging the question, straw man), and science-specific fallacies (Galileo gambit, argument from ignorance, nirvana fallacy, unfalsifiability).130131**Reference:** Full fallacy catalog with the plain-language definition of each named fallacy moved to `references/evaluation_checklists.md` (Logical Fallacy Identification section). See `references/logical_fallacies.md` for examples and detection strategies.132133**When identifying fallacies:**134- Name the specific fallacy135- Explain why the reasoning is flawed136- Identify what evidence would be needed for valid inference137- Note that fallacious reasoning doesn't prove the conclusion false—just that this argument doesn't support it138139**Reference:** See `references/logical_fallacies.md` for comprehensive fallacy catalog with examples and detection strategies.140141### 6. Research Design Guidance142143Provide constructive guidance for planning rigorous studies.144145**Apply when:**146- Helping design new experiments147- Planning research projects148- Reviewing research proposals149- Improving study protocols150151**Design process:**1521531. **Research Question Refinement**154 - Ensure question is specific, answerable, and falsifiable155 - Verify it addresses a gap or contradiction in literature156 - Confirm feasibility (resources, ethics, time)157 - Define variables operationally1581592. **Design Selection**160 - Match design to question (causal → experimental; associational → observational)161 - Consider feasibility and ethical constraints162 - Choose between-subjects, within-subjects, or mixed designs163 - Plan factorial designs if testing multiple factors1641653. **Bias Minimization Strategy**166 - Implement randomization when possible167 - Plan blinding at all feasible levels (participants, providers, assessors)168 - Identify and plan to control confounds (randomization, matching, stratification, statistical adjustment)169 - Standardize all procedures170 - Plan to minimize attrition1711724. **Sample Planning**173 - Conduct a priori power analysis (specify expected effect, desired power, alpha)174 - Account for attrition in sample size175 - Define clear inclusion/exclusion criteria176 - Consider recruitment strategy and feasibility177 - Plan for sample representativeness1781795. **Measurement Strategy**180 - Select validated, reliable instruments181 - Use objective measures when possible182 - Plan multiple measures of key constructs (triangulation)183 - Ensure measures are sensitive to expected changes184 - Establish inter-rater reliability procedures1851866. **Analysis Planning**187 - Prespecify all hypotheses and analyses188 - Designate primary outcome clearly189 - Plan statistical tests with assumption checks190 - Specify how missing data will be handled191 - Plan to report effect sizes and confidence intervals192 - Consider multiple comparison corrections1931947. **Transparency and Rigor**195 - Preregister study and analysis plan196 - Use reporting guidelines (CONSORT, STROBE, PRISMA)197 - Plan to report all outcomes, not just significant ones198 - Distinguish confirmatory from exploratory analyses199 - Commit to data/code sharing200201**Reference:** See `references/experimental_design.md` for comprehensive design checklist covering all stages from question to dissemination.202203### 7. Claim Evaluation204205Systematically evaluate scientific claims for validity and support.206207**Apply when:**208- Assessing conclusions in papers209- Evaluating media reports of research210- Reviewing abstract or introduction claims211- Checking if data support conclusions212213**Claim evaluation process:**2142151. **Identify the Claim**216 - What exactly is being claimed?217 - Is it a causal claim, associational claim, or descriptive claim?218 - How strong is the claim (proven, likely, suggested, possible)?2192202. **Assess the Evidence**221 - What evidence is provided?222 - Is evidence direct or indirect?223 - Is evidence sufficient for the strength of claim?224 - Are alternative explanations ruled out?2252263. **Check Logical Connection**227 - Do conclusions follow from the data?228 - Are there logical leaps?229 - Is correlational data used to support causal claims?230 - Are limitations acknowledged?2312324. **Evaluate Proportionality**233 - Is confidence proportional to evidence strength?234 - Are hedging words used appropriately?235 - Are limitations downplayed?236 - Is speculation clearly labeled?2372385. **Check for Overgeneralization**239 - Do claims extend beyond the sample studied?240 - Are population restrictions acknowledged?241 - Is context-dependence recognized?242 - Are caveats about generalization included?2432446. **Red Flags**245 - Causal language from correlational studies246 - "Proves" or absolute certainty247 - Cherry-picked citations248 - Ignoring contradictory evidence249 - Dismissing limitations250 - Extrapolation beyond data251252**Provide specific feedback:**253- Quote the problematic claim254- Explain what evidence would be needed to support it255- Suggest appropriate hedging language if warranted256- Distinguish between data (what was found) and interpretation (what it means)257258## Application Guidelines259260### General Approach2612621. **Be Constructive**263 - Identify strengths as well as weaknesses264 - Suggest improvements rather than just criticizing265 - Distinguish between fatal flaws and minor limitations266 - Recognize that all research has limitations2672682. **Be Specific**269 - Point to specific instances (e.g., "Table 2 shows..." or "In the Methods section...")270 - Quote problematic statements271 - Provide concrete examples of issues272 - Reference specific principles or standards violated2732743. **Be Proportionate**275 - Match criticism severity to issue importance276 - Distinguish between major threats to validity and minor concerns277 - Consider whether issues affect primary conclusions278 - Acknowledge uncertainty in your own assessments2792804. **Apply Consistent Standards**281 - Use same criteria across all studies282 - Don't apply stricter standards to findings you dislike283 - Acknowledge your own potential biases284 - Base judgments on methodology, not results2852865. **Consider Context**287 - Acknowledge practical and ethical constraints288 - Consider field-specific norms for effect sizes and methods289 - Recognize exploratory vs. confirmatory contexts290 - Account for resource limitations in evaluating studies291292### When Providing Critique293294**Structure feedback as:**2952961. **Summary:** Brief overview of what was evaluated2972. **Strengths:** What was done well (important for credibility and learning)2983. **Concerns:** Issues organized by severity299 - Critical issues (threaten validity of main conclusions)300 - Important issues (affect interpretation but not fatally)301 - Minor issues (worth noting but don't change conclusions)3024. **Specific Recommendations:** Actionable suggestions for improvement3035. **Overall Assessment:** Balanced conclusion about evidence quality and what can be concluded304305**Use precise terminology:**306- Name specific biases, fallacies, and methodological issues307- Reference established standards and guidelines308- Cite principles from scientific methodology309- Use technical terms accurately310311### When Uncertain312313- **Acknowledge uncertainty:** "This could be X or Y; additional information needed is Z"314- **Ask clarifying questions:** "Was [methodological detail] done? This affects interpretation."315- **Provide conditional assessments:** "If X was done, then Y follows; if not, then Z is concern"316- **Note what additional information would resolve uncertainty**317318## Reference Materials319320This skill includes comprehensive reference materials that provide detailed frameworks for critical evaluation:321322- **`references/scientific_method.md`** - Core principles of scientific methodology, the scientific process, critical evaluation criteria, red flags in scientific claims, causal inference standards, peer review, and open science principles323324- **`references/common_biases.md`** - Comprehensive taxonomy of cognitive, experimental, methodological, statistical, and analysis biases with detection and mitigation strategies325326- **`references/statistical_pitfalls.md`** - Common statistical errors and misinterpretations including p-value misunderstandings, multiple comparisons problems, sample size issues, effect size mistakes, correlation/causation confusion, regression pitfalls, and meta-analysis issues327328- **`references/evidence_hierarchy.md`** - Traditional evidence hierarchy, GRADE system, study quality assessment criteria, domain-specific considerations, evidence synthesis principles, and practical decision frameworks329330- **`references/logical_fallacies.md`** - Logical fallacies common in scientific discourse organized by type (causation, generalization, authority, relevance, structure, statistical) with examples and detection strategies331332- **`references/experimental_design.md`** - Comprehensive experimental design checklist covering research questions, hypotheses, study design selection, variables, sampling, blinding, randomization, control groups, procedures, measurement, bias minimization, data management, statistical planning, ethical considerations, validity threats, and reporting standards333334- **`references/evaluation_checklists.md`** - Operational, question-form review checklists for the Bias Detection, Statistical Analysis Evaluation, Evidence Quality Assessment, and Logical Fallacy Identification capabilities (the full item-by-item detail relocated from the Core Capabilities section)335336**When to consult references:**337- Load references into context when detailed frameworks are needed338- Use grep to search references for specific topics: `grep -r "pattern" references/`339- References provide depth; SKILL.md provides procedural guidance340- Consult references for comprehensive lists, detailed criteria, and specific examples341342## Remember343344**Scientific critical thinking is about:**345- Systematic evaluation using established principles346- Constructive critique that improves science347- Proportional confidence to evidence strength348- Transparency about uncertainty and limitations349- Consistent application of standards350- Recognition that all research has limitations351- Balance between skepticism and openness to evidence352353**Always distinguish between:**354- Data (what was observed) and interpretation (what it means)355- Correlation and causation356- Statistical significance and practical importance357- Exploratory and confirmatory findings358- What is known and what is uncertain359- Evidence against a claim and evidence for the null360361**Goals of critical thinking:**3621. Identify strengths and weaknesses accurately3632. Determine what conclusions are supported3643. Recognize limitations and uncertainties3654. Suggest improvements for future work3665. Advance scientific understanding367