Back to Insights

Criterion Validity: What It Is and Why Most Employers Can't Demonstrate It

Criterion validity is the gold standard for proving a hiring assessment works. Here is what it actually requires and what to do when a formal study isn't feasible.

When someone says a hiring assessment is “validated,” they usually mean it has criterion-related validity evidence. This is the type of validity people find most intuitive: does a high score on the assessment predict high performance on the job?

The answer requires a study. Most employers have never done one.

What criterion validity actually involves

Criterion-related validity is the statistical relationship between a predictor (the assessment score) and a criterion (some measure of job performance). It is expressed as a correlation coefficient, where 0 means no relationship and 1 means perfect prediction.

No selection assessment reaches 1. The question is whether the correlation is large enough to be useful and statistically significant given your sample.

A validity coefficient of .30 to .40 is considered acceptable. A coefficient above .50 is strong. Structured behavioral interviews typically fall in the .50 to .60 range in well-designed studies.

Predictive validity: You administer the assessment to applicants before hire, then collect performance data after they have been on the job for a defined period. The correlation between their pre-hire scores and later performance is the validity coefficient.

This is the methodologically cleaner approach because it mirrors the actual use case. The practical problem is that it takes time. You need to hire enough people and wait long enough to measure performance before you have results.

Concurrent validity: You administer the assessment to current employees and correlate their scores with current performance ratings. You get results faster, but concurrent samples differ from applicant samples in ways that can bias the results. High-performing employees who might have scored low have already left. New hires who might have scored high are not yet in the sample.

On sample size: The conventional minimum for a criterion validity study is around 150 participants. With smaller samples, correlations are unstable: a coefficient of .40 in a sample of 50 has a confidence interval wide enough to include zero. Most employers cannot reach the required sample size for a single role, which is why synthetic validity and transportability studies exist.

The performance criterion problem

A validity study is only as good as the performance measure it uses. If you are correlating interview scores with supervisor ratings, and those ratings reflect how much the supervisor likes the employee rather than actual performance, your validity coefficient is measuring something other than what you intended.

Good performance criteria are:

  • Objective where possible (sales figures, error rates, completion times)
  • Collected consistently across raters
  • Job-relevant: they measure what the job actually requires, not general impressions
  • Collected at an appropriate time interval (too early, and the employee is still ramping up; too late, and performance has been shaped by factors unrelated to hire-time capability)

Most organizations have weak performance measurement. This is one reason why validity studies are hard to conduct well even when sample size is not a barrier.

What to do when a formal study is not feasible

For most employers, especially smaller ones, a full criterion validity study for each role is not practical. The alternatives:

Transportability: Evidence that a validity study conducted elsewhere (for a similar role, in a similar context) applies to your situation. UGESP permits this, with documentation.

Synthetic validity: Decompose the job into components, assemble validity evidence for each component from published research, and aggregate the estimates. Meta-analytic databases like those compiled by Schmidt and Hunter provide validity coefficients for many assessment types across job families.

Content validity: Demonstrate that the assessment samples the content of the job directly. This is the standard for work samples and job knowledge tests and does not require a criterion study.

.56 Mean validity coefficient for structured interviews in the Schmidt & Hunter (1998) meta-analysis, making them one of the highest-validating practical selection tools

The point is not that every employer must conduct an original study. The point is that employers need some form of validity evidence for procedures that produce adverse impact. Choosing well-validated assessment types and documenting the connection to job requirements is the minimum standard.

Talent Systems AI

See structured, scored interviewing running on your roles.

Validated competency rubrics. Adverse impact monitoring. Full audit trail from day one.

Book a Demo