Fairness and bias analysis
Two distinct ideas that are routinely conflated. Keep them separate.
Fairness is a social concept, not a single statistic
"Fairness" has no single agreed definition (statistical, psychometric, or social). Recognized
meanings include:
- Equal group outcomes (e.g., equal passing rates). The Standards reject this as the
definition of fairness: group outcome differences alone do not indicate bias — though they
should trigger heightened scrutiny for possible bias.
- Equitable treatment of all examinees (testing conditions, access to practice materials,
feedback, retest opportunities, reasonable accommodation, mode of administration).
- Comparable access to the construct — accessible testing so all candidates can show their
standing without being advantaged/disadvantaged by construct-irrelevant characteristics (age,
race, ethnicity, gender, SES, cultural/linguistic background, disability).
- Lack of bias.
There is broad agreement that equitable treatment, access, bias, and scrutiny when subgroup
differences appear are important — but no agreement that "fairness" can be uniquely defined in
terms of any one of them.
Bias is a technical concept — two forms
Bias = systematic error that differentially affects the performance of different subgroups.
- Predictive bias — slope and/or intercept of the predictor→criterion regression differs across
groups (a predictor–criterion relationship issue).
- Measurement bias — construct-irrelevant variance producing systematically higher/lower
scores for a subgroup (a score issue, for predictors or criteria).
Crucial point on consequences
A subgroup-mean difference (adverse impact) is a negative consequence, but it is evidence against
validity only if it traces to a measurement property of the procedure (i.e., bias). If the group
difference on the procedure mirrors a real difference in the work-relevant outcome (i.e., no
predictive bias), the consequence is a policy issue for the user, not a validity defect.
Predictive bias / differential prediction
Test via moderated multiple regression (MMR): regress the criterion on the predictor, subgroup
membership, and their interaction. Slope and/or intercept differences signal predictive bias.
MMR is preferred over comparing separate subgroup correlation coefficients.
- Frame the question as "is the subgroup's performance underpredicted?" — only
underprediction signals bias against that group. Simply knowing slopes/intercepts differ
doesn't answer it. (In U.S. cognitive-ability research, slope differences are rare; when intercept
differences occur they typically take the form of overprediction of minority performance —
Schmidt, Pearlman, & Hunter, 1980; and corrected analyses, e.g., Berry & Zhao, 2015, still find
little underprediction.)
- Consider effect sizes as well as statistical significance (Nye & Sackett, 2017; Dahlke &
Sackett, 2017).
Technical cautions (predictive bias)
- Analyze predictors as operationally used (e.g., test the composite when selection uses a
composite, not each test separately).
- A confident, unbiased criterion is a prerequisite.
- Statistical power is a chronic problem — small total/subgroup samples, unequal subgroup sizes,
range restriction, and predictor unreliability all reduce power to detect slope/intercept
differences.
- Check the homogeneity-of-error-variance assumption; use alternative tests when it's violated.
- Use an unbiased estimate of the intercept difference and operational validity parameters
(not observed parameters).
Predictive bias and mean differences can exist independently; analyze predictive bias when
there's compelling reason to question whether predictor and criterion relate comparably across
subgroups and appropriate data exist. Where relevant research exists, generalized evidence can
inform the question.
Measurement bias
Construct-irrelevant variance raising/lowering scores for a subgroup — hard to detect because it
requires comparing an observed score to a true score. Approaches:
- Item sensitivity review — diverse reviewers examine items (and instructions to candidates and
scorers) for language/content that could carry differing meaning across subgroups or be
demeaning/offensive. Value depends on content; use is a matter of professional judgment.
- Differential item functioning (DIF) — identifies items on which members of different subgroups
with the same total score (or same IRT true score) perform differently. Notes:
- Needs large samples for stable results.
- Domains where DIF is common have rarely shown sizable, replicable DIF (Sackett et al., 2001);
for cognitive tests it's common to find roughly equal numbers of items favoring each subgroup,
netting to little test-level bias.
- DIF is not a routine/expected part of selection development; explore it when appropriate data
exist. Especially useful in cross-cultural / linguistically different testing.
Pitfalls
- Equating adverse impact with bias, or "no bias" with "fair."
- Testing each component instead of the operational composite.
- Running underpowered bias analyses and reading a null as "no bias."
- Comparing subgroup correlations instead of using MMR.
- Treating any slope/intercept difference as bias without asking about underprediction direction.
Checklist
See also
criterion-related-validation · selection-decisions-and-scoring (composites & subgroup tradeoffs)
· candidate-accommodations (equitable treatment/access) · internal-structure-validation ·
technical-validation-report
Source: Principles (5th ed., 2018), "Fairness and Bias."
1---2name: fairness-and-bias-analysis3description: Use when evaluating fairness and bias of a selection procedure — distinguishing the several meanings of "fairness," testing for predictive bias (differential prediction via moderated regression), and examining measurement bias (DIF, item sensitivity review). Covers what subgroup-mean differences do and don't imply, when bias analyses are warranted, and the statistical pitfalls. Triggers: "adverse impact vs bias", "differential prediction", "predictive bias", "measurement bias", "DIF analysis", "is the test fair / biased", "subgroup differences", "item sensitivity review".4license: MIT5---67# Fairness and bias analysis89Two distinct ideas that are routinely conflated. Keep them separate.1011## Fairness is a social concept, not a single statistic1213"Fairness" has **no single agreed definition** (statistical, psychometric, or social). Recognized14meanings include:15- **Equal group outcomes** (e.g., equal passing rates). The *Standards* **reject** this as the16 definition of fairness: group outcome differences alone **do not indicate bias** — though they17 should **trigger heightened scrutiny** for possible bias.18- **Equitable treatment** of all examinees (testing conditions, access to practice materials,19 feedback, retest opportunities, reasonable accommodation, mode of administration).20- **Comparable access to the construct** — accessible testing so all candidates can show their21 standing without being advantaged/disadvantaged by construct-irrelevant characteristics (age,22 race, ethnicity, gender, SES, cultural/linguistic background, disability).23- **Lack of bias.**2425There is broad agreement that equitable treatment, access, bias, and scrutiny when subgroup26differences appear are important — but **no agreement that "fairness" can be uniquely defined** in27terms of any one of them.2829## Bias is a technical concept — two forms3031**Bias** = systematic error that differentially affects the performance of different subgroups.3233- **Predictive bias** — slope and/or intercept of the predictor→criterion regression differs across34 groups (a *predictor–criterion* relationship issue).35- **Measurement bias** — construct-irrelevant variance producing systematically higher/lower36 **scores** for a subgroup (a *score* issue, for predictors or criteria).3738### Crucial point on consequences39A subgroup-mean difference (adverse impact) is a **negative consequence**, but it is evidence against40validity **only if it traces to a measurement property** of the procedure (i.e., bias). If the group41difference on the procedure mirrors a real difference in the work-relevant outcome (i.e., **no42predictive bias**), the consequence is a **policy issue** for the user, not a validity defect.4344## Predictive bias / differential prediction4546Test via **moderated multiple regression (MMR)**: regress the criterion on the predictor, subgroup47membership, and their **interaction**. Slope and/or intercept differences signal predictive bias.48MMR is preferred over comparing separate subgroup correlation coefficients.4950- Frame the question as **"is the subgroup's performance *underpredicted*?"** — only51 **underprediction** signals bias against that group. Simply knowing slopes/intercepts differ52 doesn't answer it. (In U.S. cognitive-ability research, slope differences are rare; when intercept53 differences occur they typically take the form of **overprediction** of minority performance —54 Schmidt, Pearlman, & Hunter, 1980; and corrected analyses, e.g., Berry & Zhao, 2015, still find55 little underprediction.)56- Consider **effect sizes** as well as statistical significance (Nye & Sackett, 2017; Dahlke &57 Sackett, 2017).5859### Technical cautions (predictive bias)601. Analyze predictors **as operationally used** (e.g., test the **composite** when selection uses a61 composite, not each test separately).622. A **confident, unbiased criterion** is a prerequisite.633. **Statistical power is a chronic problem** — small total/subgroup samples, unequal subgroup sizes,64 range restriction, and predictor unreliability all reduce power to detect slope/intercept65 differences.664. Check the **homogeneity-of-error-variance** assumption; use alternative tests when it's violated.675. Use an **unbiased estimate** of the intercept difference and operational validity parameters68 (not observed parameters).6970Predictive bias and mean differences can exist **independently**; analyze predictive bias when71there's compelling reason to question whether predictor and criterion relate comparably across72subgroups *and* appropriate data exist. Where relevant research exists, generalized evidence can73inform the question.7475## Measurement bias7677Construct-irrelevant variance raising/lowering scores for a subgroup — hard to detect because it78requires comparing an observed score to a **true** score. Approaches:7980- **Item sensitivity review** — diverse reviewers examine items (and instructions to candidates and81 scorers) for language/content that could carry **differing meaning** across subgroups or be82 **demeaning/offensive**. Value depends on content; use is a matter of professional judgment.83- **Differential item functioning (DIF)** — identifies items on which members of different subgroups84 with the **same total score (or same IRT true score)** perform differently. Notes:85 - Needs **large samples** for stable results.86 - Domains where DIF is common have **rarely** shown sizable, replicable DIF (Sackett et al., 2001);87 for cognitive tests it's common to find roughly equal numbers of items favoring each subgroup,88 netting to little test-level bias.89 - DIF is **not a routine/expected** part of selection development; explore it when appropriate data90 exist. Especially useful in **cross-cultural / linguistically different** testing.9192## Pitfalls9394- Equating adverse impact with bias, or "no bias" with "fair."95- Testing each component instead of the operational composite.96- Running underpowered bias analyses and reading a null as "no bias."97- Comparing subgroup correlations instead of using MMR.98- Treating any slope/intercept difference as bias without asking about *under*prediction direction.99100## Checklist101102- [ ] "Fairness" meaning(s) at issue named explicitly103- [ ] Subgroup differences treated as a scrutiny trigger, not a verdict104- [ ] Predictive bias tested with MMR on the **operational** predictor/composite105- [ ] Underprediction direction (not mere difference) interpreted; effect sizes reported106- [ ] Power, range restriction, unreliability, error-variance homogeneity addressed107- [ ] Unbiased parameter estimates used108- [ ] Measurement bias considered (item sensitivity review and/or DIF) where data/justification exist109- [ ] Equitable treatment and access to the construct addressed (see accommodations skill)110111## See also112113`criterion-related-validation` · `selection-decisions-and-scoring` (composites & subgroup tradeoffs)114· `candidate-accommodations` (equitable treatment/access) · `internal-structure-validation` ·115`technical-validation-report`116117*Source: Principles (5th ed., 2018), "Fairness and Bias."*