How I-O Psychology Defines Fairness in Selection
Fairness in hiring has multiple definitions, and they do not always agree. Here is how the field distinguishes procedural fairness from outcome fairness and why the difference matters.
“Fairness” is one of those words that sounds simple until you try to define it precisely. In selection, it has several distinct technical definitions, and they do not always point in the same direction. An assessment can be fair by one definition and unfair by another.
Understanding how the field defines fairness is essential for building selection systems that are both legally defensible and actually equitable.
Procedural fairness
Procedural fairness is about how the process is run. Is every candidate asked the same questions? Are all candidates scored on the same criteria? Does the process give candidates an equal opportunity to demonstrate their qualifications?
Structured interviews with standardized questions and scoring rubrics are procedurally fair in this sense. The process is the same for everyone. This is the baseline.
Procedural fairness is necessary but not sufficient. A process can be procedurally identical for all candidates and still produce systematically different outcomes across demographic groups.
Outcome fairness and adverse impact
Outcome fairness focuses on whether the selection procedure produces equal outcomes across groups. Adverse impact (measured by the 4/5ths rule) is the primary quantitative indicator: does the selection rate differ significantly between protected groups?
A process that is procedurally fair can still produce adverse impact. If a cognitive test is administered identically to all candidates but Group A passes at 80% and Group B passes at 50%, the test has adverse impact even though the procedure was equal.
Measurement equivalence
A third definition: does the assessment measure the same construct in the same way across groups? This is called measurement equivalence or measurement invariance.
If a structured interview is measuring “communication competency” for Group A but inadvertently measuring “familiarity with dominant professional norms” for Group B, the scores are not comparable across groups even if the process looks identical.
Testing for measurement equivalence requires statistical analysis (confirmatory factor analysis, differential item functioning, or item response theory methods). It is more technically demanding than adverse impact monitoring but is increasingly required in rigorous validation contexts.
Differential item functioning
Within a test or interview, individual items can behave differently across groups even when the overall assessment appears equivalent. Differential item functioning (DIF) analysis identifies items that are harder or easier for one group than another after controlling for overall ability level.
An item shows DIF if, say, male and female candidates at the same underlying ability level perform differently on that specific item. DIF does not automatically mean the item is biased: there may be a legitimate reason for the difference related to the construct. But it is a signal that the item warrants closer review.
Why “no adverse impact” doesn’t mean “fair”
Eliminating adverse impact by removing valid predictors makes the process less predictive, which reduces the quality of hiring decisions across all candidates. This is not a better outcome.
The goal is not zero adverse impact at any cost. The goal is to:
- Use the most valid selection methods available
- Monitor for adverse impact at each stage
- Investigate the source of any observed disparity
- Reduce adverse impact where you can do so without sacrificing validity
- Ensure procedural consistency and measurement equivalence
This is a harder target than simply running a process that looks equal. But it is what “fair” actually requires.
Talent Systems AI
See structured, scored interviewing running on your roles.
Validated competency rubrics. Adverse impact monitoring. Full audit trail from day one.
Book a Demo