Research Methods in Psychology
Psychology's claim to be a science rests on its methods. Unlike philosophy or folk wisdom, psychology makes empirical claims testable through systematic observation and experimentation. This skill covers the experimental method as applied in psychological research, the ethical framework governing human subjects research, the statistical tools used to evaluate evidence, and the ongoing replication crisis that has forced the field to confront its methodological shortcomings.
Agent affinity: kahneman (research design, statistical thinking), james (methodological pragmatism, what constitutes evidence)
Concept IDs: psych-perception-construction, psych-cognitive-biases, psych-social-influence, psych-learning-theory
Research Methods at a Glance
| # |
Domain |
Core Question |
Key Concepts |
| 1 |
Experimental design |
How do we establish causation? |
IV/DV, random assignment, control, confounds |
| 2 |
Non-experimental methods |
How do we study what we cannot manipulate? |
Correlation, observation, case studies, surveys |
| 3 |
Research ethics |
How do we protect participants? |
Informed consent, deception, debriefing, IRB |
| 4 |
Statistics in psychology |
How do we evaluate evidence? |
NHST, effect size, confidence intervals, power |
| 5 |
The replication crisis |
Can we trust published findings? |
Publication bias, p-hacking, preregistration |
Domain 1 -- Experimental Design
The logic of experimentation
An experiment establishes causation by manipulating one variable (the independent variable, IV) while holding all other variables constant, and measuring the effect on another variable (the dependent variable, DV). Random assignment to conditions ensures that participant characteristics are distributed evenly across groups, eliminating systematic confounds.
Key elements
- Independent variable (IV) -- the factor the researcher manipulates. Must have at least two levels (e.g., treatment vs. control).
- Dependent variable (DV) -- the outcome measured. Must be operationally defined (e.g., "anxiety" measured by the Beck Anxiety Inventory score, not a vague assessment).
- Random assignment -- each participant has an equal chance of being assigned to any condition. This is what distinguishes a true experiment from a quasi-experiment.
- Control condition -- a comparison group that does not receive the treatment. May be no-treatment, waitlist, active control (alternative treatment), or placebo.
- Confound -- a variable that varies systematically with the IV, making it impossible to determine which caused the effect on the DV.
Between-subjects vs. within-subjects
| Design |
Each participant |
Advantage |
Disadvantage |
| Between-subjects |
Experiences one condition |
No order effects, no demand characteristics from multiple testing |
Needs more participants, individual differences add noise |
| Within-subjects (repeated measures) |
Experiences all conditions |
Each participant serves as their own control, more statistical power |
Order effects, practice effects, fatigue |
| Mixed |
Between on one IV, within on another |
Combines advantages |
Complexity in analysis and interpretation |
Counterbalancing (varying the order of conditions across participants) controls for order effects in within-subjects designs.
Factorial designs
When two or more IVs are crossed (every level of each IV is combined with every level of every other IV), the design is factorial. A 2x3 factorial has 6 conditions. Factorial designs reveal main effects (the effect of each IV averaging over others) and interactions (the effect of one IV depends on the level of another). Interactions are often more theoretically interesting than main effects.
Validity
- Internal validity -- confidence that the IV caused the change in the DV. Threatened by confounds, selection bias, maturation, history, and attrition.
- External validity -- generalizability to other populations, settings, and times. Threatened by non-representative samples, artificial laboratory settings, and demand characteristics.
- Construct validity -- the degree to which the IV and DV actually measure the constructs of interest.
- Statistical conclusion validity -- appropriate use of statistical tests, adequate sample size, and correct interpretation of results.
Domain 2 -- Non-Experimental Methods
Correlational research
Measures the relationship between two variables without manipulating either. Correlation does not establish causation because of the third-variable problem (an unmeasured variable may cause both) and the directionality problem (does A cause B or B cause A?). Correlation coefficients range from -1 to +1; the sign indicates direction, the magnitude indicates strength.
Observational methods
- Naturalistic observation -- observing behavior in its natural setting without intervention. High ecological validity, low internal validity.
- Structured observation -- observing behavior in a controlled setting with standardized procedures (e.g., Ainsworth's Strange Situation).
- Participant observation -- the researcher joins the group being studied.
Case studies
Intensive investigation of a single individual or small group. Invaluable for rare conditions (H.M., Phineas Gage, Genie) and for generating hypotheses, but cannot establish causation or generalize to populations.
Surveys
Self-report measures administered to large samples. Efficient for gathering data on attitudes, beliefs, and behaviors. Vulnerable to social desirability bias, acquiescence bias, and poorly worded questions.
Domain 3 -- Research Ethics
Historical catalysts
- Nuremberg Code (1947) -- established voluntary consent as essential, in response to Nazi medical experiments.
- Tuskegee Syphilis Study (1932-1972) -- untreated Black men with syphilis were studied for 40 years without informed consent. Led directly to the Belmont Report and modern IRB requirements.
- Milgram (1963) and Zimbardo (1971) -- raised questions about psychological harm from research participation.
APA Ethical Principles
The American Psychological Association's Ethical Principles (2017) include:
- Beneficence and nonmaleficence -- maximize benefits, minimize harm
- Fidelity and responsibility -- establish trust, accept responsibility
- Integrity -- promote accuracy, honesty, truthfulness
- Justice -- fair and equitable access to and benefit from research
- Respect for people's rights and dignity -- protect autonomy, privacy, confidentiality
Key ethical requirements
- Informed consent -- participants must understand the study's purpose, procedures, risks, and their right to withdraw without penalty. Written consent is standard.
- Deception -- permitted only when (a) the study cannot be conducted without it, (b) the scientific value justifies it, and (c) participants are debriefed promptly. Milgram's experiments would require extensive justification today.
- Debriefing -- after participation, explain the study's true purpose and any deception. Address any distress caused.
- Institutional Review Board (IRB) -- independent committee that reviews research proposals for ethical compliance before data collection begins.
- Confidentiality -- participant data must be stored securely and reported in ways that prevent identification.
Vulnerable populations
Research with children, prisoners, cognitively impaired individuals, and other vulnerable populations requires additional safeguards: parental consent plus child assent, independent advocates, and heightened risk-benefit scrutiny.
Domain 4 -- Statistics in Psychology
Null Hypothesis Significance Testing (NHST)
The dominant (and controversial) inferential framework in psychology:
- State a null hypothesis (H0: no effect/no difference) and an alternative hypothesis (H1: there is an effect).
- Collect data and compute a test statistic (t, F, chi-square, etc.).
- Calculate the p-value: the probability of obtaining results as extreme as observed, assuming H0 is true.
- If p < alpha (conventionally .05), reject H0.
What the p-value is and is not
| p-value IS |
p-value IS NOT |
| P(data or more extreme | H0 is true) |
P(H0 is true | data) |
| A measure of data surprise under H0 |
A measure of effect size or practical importance |
| Affected by sample size |
A measure of replication probability |
A p-value of .04 does not mean there is a 96% chance the effect is real. It means that if H0 were true, data this extreme would occur about 4% of the time.
Effect size
Effect size quantifies the magnitude of a finding independently of sample size:
- Cohen's d -- standardized mean difference. Small = 0.2, medium = 0.5, large = 0.8.
- Pearson's r -- correlation coefficient. Small = .10, medium = .30, large = .50.
- Odds ratio -- ratio of odds of an event in two groups. Used in clinical and epidemiological research.
- Eta-squared -- proportion of variance explained. Used in ANOVA.
A statistically significant result with a tiny effect size may be practically meaningless. A non-significant result with a large effect size may reflect inadequate sample size.
Confidence intervals
A 95% confidence interval means: if we repeated the study many times, 95% of the constructed intervals would contain the true parameter value. A confidence interval that does not include zero is equivalent to p < .05 for a two-tailed test. Confidence intervals communicate both the estimate and its precision.
Power analysis
Statistical power is the probability of detecting a true effect (1 - beta, where beta is the Type II error rate). Convention: power >= .80. Power depends on effect size, sample size, and alpha level. Underpowered studies are likely to produce false negatives or, when they do reach significance, inflated effect sizes (the "winner's curse"). Cohen (1962, 1992) documented that psychology studies are chronically underpowered.
Domain 5 -- The Replication Crisis
The problem
The Open Science Collaboration (2015) attempted to replicate 100 published psychology studies. Only 36% of replications reached statistical significance (vs. 97% of originals). Mean effect sizes in replications were half the size of originals. This result, published in Science, catalyzed a reckoning across the field.
Causes
- Publication bias -- journals preferentially publish positive results. Studies that fail to find an effect sit in the "file drawer" (Rosenthal, 1979).
- p-hacking -- exploiting researcher degrees of freedom (excluding outliers, adding covariates, testing multiple DVs, stopping data collection when p < .05) to produce significant results from noise. Simmons, Nelson, & Simonsohn (2011) showed that these practices can produce p < .05 for a false hypothesis with high probability.
- HARKing (Hypothesizing After Results are Known) -- presenting post-hoc findings as if they were predicted a priori. Kerr (1998).
- Underpowered studies -- small samples produce noisy estimates, and only the (inflated) significant ones are published.
- Incentive structure -- academic careers reward novel, significant findings. Replications and null results are not rewarded.
Solutions
- Preregistration -- publicly specifying hypotheses, methods, and analysis plans before data collection. Prevents p-hacking and HARKing.
- Registered Reports -- journals accept or reject studies based on the introduction and method, before results are known. Eliminates publication bias.
- Open data and open materials -- sharing data and stimuli enables verification and reanalysis.
- Large-scale replications -- Many Labs projects (Klein et al., 2014) coordinate replications across multiple laboratories.
- Effect size reporting -- APA Publication Manual (7th ed., 2020) requires effect size reporting alongside significance tests.
- Bayesian statistics -- quantify evidence for H1 vs. H0 via Bayes factors, avoiding the binary significant/non-significant framework entirely.
Cross-References
- kahneman agent: Statistical thinking, cognitive biases in research interpretation, and the psychology of judgment under uncertainty that explains why researchers fall prey to p-hacking and HARKing.
- james agent: Pragmatic epistemology -- what counts as evidence, the relationship between theory and data, and the history of psychological methodology.
- piaget agent: Developmental methodology (longitudinal vs. cross-sectional designs, challenges of studying children).
- cognitive-psychology skill: Experimental paradigms (reaction time, signal detection, priming) used to study cognitive processes.
- social-psychology skill: Methodological controversies specific to social psychology (Milgram ethics, demand characteristics, ecological validity).
- behavioral-neuroscience skill: Neuroimaging methodology (fMRI, EEG, PET) as a complement to behavioral measurement.
References
- Cohen, J. (1992). A power primer. Psychological Bulletin, 112(1), 155-159.
- Kerr, N. L. (1998). HARKing: Hypothesizing after the results are known. Personality and Social Psychology Review, 2(3), 196-217.
- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716.
- Rosenthal, R. (1979). The file drawer problem and tolerance for null results. Psychological Bulletin, 86(3), 638-641.
- Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366.
- American Psychological Association. (2017). Ethical Principles of Psychologists and Code of Conduct. APA.
- American Psychological Association. (2020). Publication Manual (7th ed.). APA.
1---2name: research-methods-psych3description: Research methodology in psychology. Covers experimental design (independent/dependent variables, random assignment, control conditions, between/within-subjects designs), research ethics (informed consent, deception, debriefing, IRB review, APA ethical principles), statistical methods in psychology (null hypothesis significance testing, p-values, effect sizes, confidence intervals, power analysis), and the replication crisis (publication bias, p-hacking, questionable research practices, preregistration, open science). Use when designing psychological research, evaluating study quality, interpreting statistical findings, or discussing methodological rigor in psychology.4---5# Research Methods in Psychology
6
7Psychology's claim to be a science rests on its methods. Unlike philosophy or folk wisdom, psychology makes empirical claims testable through systematic observation and experimentation. This skill covers the experimental method as applied in psychological research, the ethical framework governing human subjects research, the statistical tools used to evaluate evidence, and the ongoing replication crisis that has forced the field to confront its methodological shortcomings.
8
9**Agent affinity:** kahneman (research design, statistical thinking), james (methodological pragmatism, what constitutes evidence)
10
11**Concept IDs:** psych-perception-construction, psych-cognitive-biases, psych-social-influence, psych-learning-theory
12
13## Research Methods at a Glance
14
15| # | Domain | Core Question | Key Concepts |
16|---|---|---|---|
17| 1 | Experimental design | How do we establish causation? | IV/DV, random assignment, control, confounds |
18| 2 | Non-experimental methods | How do we study what we cannot manipulate? | Correlation, observation, case studies, surveys |
19| 3 | Research ethics | How do we protect participants? | Informed consent, deception, debriefing, IRB |
20| 4 | Statistics in psychology | How do we evaluate evidence? | NHST, effect size, confidence intervals, power |
21| 5 | The replication crisis | Can we trust published findings? | Publication bias, p-hacking, preregistration |
22
23## Domain 1 -- Experimental Design
24
25### The logic of experimentation
26
27An experiment establishes causation by manipulating one variable (the independent variable, IV) while holding all other variables constant, and measuring the effect on another variable (the dependent variable, DV). Random assignment to conditions ensures that participant characteristics are distributed evenly across groups, eliminating systematic confounds.
28
29### Key elements
30
31- **Independent variable (IV)** -- the factor the researcher manipulates. Must have at least two levels (e.g., treatment vs. control).
32- **Dependent variable (DV)** -- the outcome measured. Must be operationally defined (e.g., "anxiety" measured by the Beck Anxiety Inventory score, not a vague assessment).
33- **Random assignment** -- each participant has an equal chance of being assigned to any condition. This is what distinguishes a true experiment from a quasi-experiment.
34- **Control condition** -- a comparison group that does not receive the treatment. May be no-treatment, waitlist, active control (alternative treatment), or placebo.
35- **Confound** -- a variable that varies systematically with the IV, making it impossible to determine which caused the effect on the DV.
36
37### Between-subjects vs. within-subjects
38
39| Design | Each participant | Advantage | Disadvantage |
40|---|---|---|---|
41| **Between-subjects** | Experiences one condition | No order effects, no demand characteristics from multiple testing | Needs more participants, individual differences add noise |
42| **Within-subjects (repeated measures)** | Experiences all conditions | Each participant serves as their own control, more statistical power | Order effects, practice effects, fatigue |
43| **Mixed** | Between on one IV, within on another | Combines advantages | Complexity in analysis and interpretation |
44
45Counterbalancing (varying the order of conditions across participants) controls for order effects in within-subjects designs.
46
47### Factorial designs
48
49When two or more IVs are crossed (every level of each IV is combined with every level of every other IV), the design is factorial. A 2x3 factorial has 6 conditions. Factorial designs reveal main effects (the effect of each IV averaging over others) and interactions (the effect of one IV depends on the level of another). Interactions are often more theoretically interesting than main effects.
50
51### Validity
52
53- **Internal validity** -- confidence that the IV caused the change in the DV. Threatened by confounds, selection bias, maturation, history, and attrition.
54- **External validity** -- generalizability to other populations, settings, and times. Threatened by non-representative samples, artificial laboratory settings, and demand characteristics.
55- **Construct validity** -- the degree to which the IV and DV actually measure the constructs of interest.
56- **Statistical conclusion validity** -- appropriate use of statistical tests, adequate sample size, and correct interpretation of results.
57
58## Domain 2 -- Non-Experimental Methods
59
60### Correlational research
61
62Measures the relationship between two variables without manipulating either. Correlation does not establish causation because of the third-variable problem (an unmeasured variable may cause both) and the directionality problem (does A cause B or B cause A?). Correlation coefficients range from -1 to +1; the sign indicates direction, the magnitude indicates strength.
63
64### Observational methods
65
66- **Naturalistic observation** -- observing behavior in its natural setting without intervention. High ecological validity, low internal validity.
67- **Structured observation** -- observing behavior in a controlled setting with standardized procedures (e.g., Ainsworth's Strange Situation).
68- **Participant observation** -- the researcher joins the group being studied.
69
70### Case studies
71
72Intensive investigation of a single individual or small group. Invaluable for rare conditions (H.M., Phineas Gage, Genie) and for generating hypotheses, but cannot establish causation or generalize to populations.
73
74### Surveys
75
76Self-report measures administered to large samples. Efficient for gathering data on attitudes, beliefs, and behaviors. Vulnerable to social desirability bias, acquiescence bias, and poorly worded questions.
77
78## Domain 3 -- Research Ethics
79
80### Historical catalysts
81
82- **Nuremberg Code (1947)** -- established voluntary consent as essential, in response to Nazi medical experiments.
83- **Tuskegee Syphilis Study (1932-1972)** -- untreated Black men with syphilis were studied for 40 years without informed consent. Led directly to the Belmont Report and modern IRB requirements.
84- **Milgram (1963) and Zimbardo (1971)** -- raised questions about psychological harm from research participation.
85
86### APA Ethical Principles
87
88The American Psychological Association's Ethical Principles (2017) include:
89
901. **Beneficence and nonmaleficence** -- maximize benefits, minimize harm
912. **Fidelity and responsibility** -- establish trust, accept responsibility
923. **Integrity** -- promote accuracy, honesty, truthfulness
934. **Justice** -- fair and equitable access to and benefit from research
945. **Respect for people's rights and dignity** -- protect autonomy, privacy, confidentiality
95
96### Key ethical requirements
97
98- **Informed consent** -- participants must understand the study's purpose, procedures, risks, and their right to withdraw without penalty. Written consent is standard.
99- **Deception** -- permitted only when (a) the study cannot be conducted without it, (b) the scientific value justifies it, and (c) participants are debriefed promptly. Milgram's experiments would require extensive justification today.
100- **Debriefing** -- after participation, explain the study's true purpose and any deception. Address any distress caused.
101- **Institutional Review Board (IRB)** -- independent committee that reviews research proposals for ethical compliance before data collection begins.
102- **Confidentiality** -- participant data must be stored securely and reported in ways that prevent identification.
103
104### Vulnerable populations
105
106Research with children, prisoners, cognitively impaired individuals, and other vulnerable populations requires additional safeguards: parental consent plus child assent, independent advocates, and heightened risk-benefit scrutiny.
107
108## Domain 4 -- Statistics in Psychology
109
110### Null Hypothesis Significance Testing (NHST)
111
112The dominant (and controversial) inferential framework in psychology:
113
1141. State a null hypothesis (H0: no effect/no difference) and an alternative hypothesis (H1: there is an effect).
1152. Collect data and compute a test statistic (t, F, chi-square, etc.).
1163. Calculate the p-value: the probability of obtaining results as extreme as observed, assuming H0 is true.
1174. If p < alpha (conventionally .05), reject H0.
118
119### What the p-value is and is not
120
121| p-value IS | p-value IS NOT |
122|---|---|
123| P(data or more extreme \| H0 is true) | P(H0 is true \| data) |
124| A measure of data surprise under H0 | A measure of effect size or practical importance |
125| Affected by sample size | A measure of replication probability |
126
127A p-value of .04 does not mean there is a 96% chance the effect is real. It means that if H0 were true, data this extreme would occur about 4% of the time.
128
129### Effect size
130
131Effect size quantifies the magnitude of a finding independently of sample size:
132
133- **Cohen's d** -- standardized mean difference. Small = 0.2, medium = 0.5, large = 0.8.
134- **Pearson's r** -- correlation coefficient. Small = .10, medium = .30, large = .50.
135- **Odds ratio** -- ratio of odds of an event in two groups. Used in clinical and epidemiological research.
136- **Eta-squared** -- proportion of variance explained. Used in ANOVA.
137
138A statistically significant result with a tiny effect size may be practically meaningless. A non-significant result with a large effect size may reflect inadequate sample size.
139
140### Confidence intervals
141
142A 95% confidence interval means: if we repeated the study many times, 95% of the constructed intervals would contain the true parameter value. A confidence interval that does not include zero is equivalent to p < .05 for a two-tailed test. Confidence intervals communicate both the estimate and its precision.
143
144### Power analysis
145
146Statistical power is the probability of detecting a true effect (1 - beta, where beta is the Type II error rate). Convention: power >= .80. Power depends on effect size, sample size, and alpha level. Underpowered studies are likely to produce false negatives or, when they do reach significance, inflated effect sizes (the "winner's curse"). Cohen (1962, 1992) documented that psychology studies are chronically underpowered.
147
148## Domain 5 -- The Replication Crisis
149
150### The problem
151
152The Open Science Collaboration (2015) attempted to replicate 100 published psychology studies. Only 36% of replications reached statistical significance (vs. 97% of originals). Mean effect sizes in replications were half the size of originals. This result, published in *Science*, catalyzed a reckoning across the field.
153
154### Causes
155
156- **Publication bias** -- journals preferentially publish positive results. Studies that fail to find an effect sit in the "file drawer" (Rosenthal, 1979).
157- **p-hacking** -- exploiting researcher degrees of freedom (excluding outliers, adding covariates, testing multiple DVs, stopping data collection when p < .05) to produce significant results from noise. Simmons, Nelson, & Simonsohn (2011) showed that these practices can produce p < .05 for a false hypothesis with high probability.
158- **HARKing** (Hypothesizing After Results are Known) -- presenting post-hoc findings as if they were predicted a priori. Kerr (1998).
159- **Underpowered studies** -- small samples produce noisy estimates, and only the (inflated) significant ones are published.
160- **Incentive structure** -- academic careers reward novel, significant findings. Replications and null results are not rewarded.
161
162### Solutions
163
164- **Preregistration** -- publicly specifying hypotheses, methods, and analysis plans before data collection. Prevents p-hacking and HARKing.
165- **Registered Reports** -- journals accept or reject studies based on the introduction and method, before results are known. Eliminates publication bias.
166- **Open data and open materials** -- sharing data and stimuli enables verification and reanalysis.
167- **Large-scale replications** -- Many Labs projects (Klein et al., 2014) coordinate replications across multiple laboratories.
168- **Effect size reporting** -- APA Publication Manual (7th ed., 2020) requires effect size reporting alongside significance tests.
169- **Bayesian statistics** -- quantify evidence for H1 vs. H0 via Bayes factors, avoiding the binary significant/non-significant framework entirely.
170
171## Cross-References
172
173- **kahneman agent:** Statistical thinking, cognitive biases in research interpretation, and the psychology of judgment under uncertainty that explains why researchers fall prey to p-hacking and HARKing.
174- **james agent:** Pragmatic epistemology -- what counts as evidence, the relationship between theory and data, and the history of psychological methodology.
175- **piaget agent:** Developmental methodology (longitudinal vs. cross-sectional designs, challenges of studying children).
176- **cognitive-psychology skill:** Experimental paradigms (reaction time, signal detection, priming) used to study cognitive processes.
177- **social-psychology skill:** Methodological controversies specific to social psychology (Milgram ethics, demand characteristics, ecological validity).
178- **behavioral-neuroscience skill:** Neuroimaging methodology (fMRI, EEG, PET) as a complement to behavioral measurement.
179
180## References
181
182- Cohen, J. (1992). A power primer. *Psychological Bulletin*, 112(1), 155-159.
183- Kerr, N. L. (1998). HARKing: Hypothesizing after the results are known. *Personality and Social Psychology Review*, 2(3), 196-217.
184- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. *Science*, 349(6251), aac4716.
185- Rosenthal, R. (1979). The file drawer problem and tolerance for null results. *Psychological Bulletin*, 86(3), 638-641.
186- Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. *Psychological Science*, 22(11), 1359-1366.
187- American Psychological Association. (2017). *Ethical Principles of Psychologists and Code of Conduct*. APA.
188- American Psychological Association. (2020). *Publication Manual* (7th ed.). APA.