AI predictors: theoretical basis (Concern 1)
Many technologically enhanced assessments use a wide variety of data scraped from applications,
resumes, social media, emails, or the Internet, then run through hundreds of candidate ML
algorithms. The substantive nature of the included variables and their linkages to job
requirements are often unknown.
The problem
- No substantive rationale. Past-employer data scraped from resumes might show that employment at
Employers A, B, C predicts future performance while D, E, F do not — with no substantive post hoc
explanation, even when all six are in the same business. Theology coursework might "predict" sales
success. Voice or facial characteristics may have no obvious theory linking them to KSAOs or
performance — justification is "inferred at best and unknown at worst."
- Proxy variables. Atheoretical predictors readily encode protected characteristics. In credit
scoring, ZIP code is a known proxy for race; AI using millions of correlations can base
decisions on such hidden relationships. A predictor can "work" statistically while being a
construct-irrelevant proxy.
- Interpretability ≠ relevance. Even when relationships are interpretable, they may have little
practical or conceptual relevance to the work performed (Braun & Kuljanin, 2015). Big data is
often massive, messy, and missing.
The debate (represent both sides honestly)
I-O psychology has long debated whether predictors need a theoretical basis:
- "Theory matters." The traditional basis for including a measure is the extent to which it
reflects a KSAO necessary to perform the job, as determined by a job analysis. The Standards
and Principles embed this in the very definition of validity: "the degree to which accumulated
evidence and theory support specific interpretations of scores… entailed by the proposed uses"
(Principles, p. 96; Standards, p. 225; emphasis added).
- "Prediction is enough." Others hold that if scores correlate with a relevant criterion
(performance, engagement, turnover), they're useful predictors and the rationale is merely "nice to
know."
The article's resolution: if the only purpose is mechanical prediction, studying constructs and
jobs is "merely a response to regulatory requirements." But if the purpose looks beyond simple
prediction, understanding the predictive relationship yields improved measures, broader coverage
of the performance domain, greater generalizability, and assurance that the system is sensible with
respect to recruiting, training, diverse applicant pools, and change over time. Systematic research
also surfaces additional variables and data sources that may predict, mediate, or explain work
behavior (Rotolo & Church, 2015).
How adverse impact changes the calculus
- Without adverse impact, even a job-irrelevant predictor would not legally require the "job
related / business necessity" defense (Title VII) and is not unlawful per se. From this view,
theory can look like an avoidable intellectual exercise.
- With adverse impact, job relatedness — which rests on relevance/theory established via job
analysis — becomes a legal requirement (see
ai-selection-legal-landscape,
ai-job-analysis-and-relevancy).
So the theoretical-basis question is partly scientific (do we want to understand prediction?)
and partly contingent on adverse impact (do we legally have to?). The deeper issue: is selection
research propelled by science, prioritizing understanding applicants' suitability through the lens
of job requirements — or is it an atheoretical, purely empirical activity to maximize predicted
outcomes?
Questions to ask (from the article)
- Are theoretical justifications necessary in employment testing?
- Is a technologically enhanced measure that predicts organizational outcomes sufficient, or does
one need to understand why that prediction occurs?
- Do theoretical justifications improve practice in employment testing?
- Do the considerations about theoretical justification change when there is adverse impact versus
when there is not?
Pitfalls
- Accepting "it predicts, so who cares why" when the use goes beyond one-shot mechanical prediction.
- Missing proxy variables that encode protected characteristics through hidden correlations.
- Assuming interpretability implies job relevance.
- Treating a serendipitous, unreplicated correlation as a justified predictor.
Checklist
See also
ai-job-analysis-and-relevancy · ai-selection-legal-landscape · ai-validity-evidence ·
work-analysis · criterion-related-validation
(predictor choice, rationale) · ai-input-data-and-design-audit
Source: Tippins, Oswald & McPhail (2021), Concern: "Lack of a Theoretical Basis for Predictors."
1---2name: ai-predictor-theoretical-basis3description: Use when an AI/ML selection tool uses predictors with no clear theoretical or job-analytic rationale — scraped data (resumes, social media, emails, the Internet), voice/facial features, or opaque big-data correlations. Covers the debate over whether predictors need a theoretical basis, proxy-variable risk (e.g., ZIP code for race), and how the presence or absence of adverse impact changes the analysis. Maps to Concern 1 of Tippins, Oswald & McPhail (2021). Triggers: "atheoretical predictors", "scraped data hiring", "why does this variable predict", "proxy variables in AI hiring", "is a correlation enough", "predictor with no rationale".4license: MIT5---67# AI predictors: theoretical basis (Concern 1)89Many technologically enhanced assessments use a wide variety of data **scraped from applications,10resumes, social media, emails, or the Internet**, then run through hundreds of candidate ML11algorithms. The **substantive nature** of the included variables and their **linkages to job12requirements** are often **unknown**.1314## The problem1516- **No substantive rationale.** Past-employer data scraped from resumes might show that employment at17 Employers A, B, C predicts future performance while D, E, F do not — with no substantive *post hoc*18 explanation, even when all six are in the same business. Theology coursework might "predict" sales19 success. Voice or facial characteristics may have **no obvious theory** linking them to KSAOs or20 performance — justification is "inferred at best and unknown at worst."21- **Proxy variables.** Atheoretical predictors readily encode protected characteristics. In credit22 scoring, **ZIP code is a known proxy for race**; AI using millions of correlations can base23 decisions on such hidden relationships. A predictor can "work" statistically while being a24 construct-irrelevant proxy.25- **Interpretability ≠ relevance.** Even when relationships are interpretable, they may have **little26 practical or conceptual relevance** to the work performed (Braun & Kuljanin, 2015). Big data is27 often massive, messy, and missing.2829## The debate (represent both sides honestly)3031I-O psychology has long debated whether predictors need a theoretical basis:32- **"Theory matters."** The traditional basis for including a measure is the extent to which it33 reflects a **KSAO necessary to perform the job, as determined by a job analysis.** The *Standards*34 and *Principles* embed this in the very definition of validity: "the degree to which accumulated35 evidence **and theory** support specific interpretations of scores… entailed by the proposed uses"36 (*Principles*, p. 96; *Standards*, p. 225; emphasis added).37- **"Prediction is enough."** Others hold that if scores correlate with a relevant criterion38 (performance, engagement, turnover), they're useful predictors and the rationale is merely "nice to39 know."4041The article's resolution: if the **only** purpose is mechanical prediction, studying constructs and42jobs is "merely a response to regulatory requirements." But if the purpose **looks beyond simple43prediction**, understanding the predictive relationship yields **improved measures, broader coverage44of the performance domain, greater generalizability, and assurance** that the system is sensible with45respect to recruiting, training, diverse applicant pools, and **change over time.** Systematic research46also surfaces **additional variables and data sources** that may predict, mediate, or explain work47behavior (Rotolo & Church, 2015).4849## How adverse impact changes the calculus5051- **Without adverse impact**, even a job-irrelevant predictor would not legally require the "job52 related / business necessity" defense (Title VII) and is not unlawful *per se*. From this view,53 theory can look like an avoidable intellectual exercise.54- **With adverse impact**, job relatedness — which rests on relevance/theory established via job55 analysis — becomes a legal requirement (see `ai-selection-legal-landscape`,56 `ai-job-analysis-and-relevancy`).5758So the theoretical-basis question is partly **scientific** (do we want to *understand* prediction?)59and partly **contingent on adverse impact** (do we legally *have to*?). The deeper issue: is selection60research **propelled by science**, prioritizing understanding applicants' suitability through the lens61of **job requirements** — or is it an **atheoretical, purely empirical** activity to maximize predicted62outcomes?6364## Questions to ask (from the article)6566- Are theoretical justifications necessary in employment testing?67- Is a technologically enhanced measure that predicts organizational outcomes **sufficient**, or does68 one need to understand **why** that prediction occurs?69- Do theoretical justifications **improve practice** in employment testing?70- Do the considerations about theoretical justification **change when there is adverse impact versus71 when there is not?**7273## Pitfalls7475- Accepting "it predicts, so who cares why" when the use goes beyond one-shot mechanical prediction.76- Missing proxy variables that encode protected characteristics through hidden correlations.77- Assuming interpretability implies job relevance.78- Treating a serendipitous, unreplicated correlation as a justified predictor.7980## Checklist8182- [ ] Each predictor's substantive link to a job-relevant KSAO examined (or its absence noted)83- [ ] Proxy-variable risk for protected characteristics assessed84- [ ] Purpose clarified: mechanical prediction only, or understanding/coverage/generalizability?85- [ ] Adverse-impact status checked and its legal implications applied86- [ ] Atheoretical/serendipitous relationships flagged for replication before use8788## See also8990`ai-job-analysis-and-relevancy` · `ai-selection-legal-landscape` · `ai-validity-evidence` ·91`work-analysis` · `criterion-related-validation`92(predictor choice, rationale) · `ai-input-data-and-design-audit`9394*Source: Tippins, Oswald & McPhail (2021), Concern: "Lack of a Theoretical Basis for Predictors."*