Cognitive and aptitude testing
Cognitive ability testing, with the adverse impact left in
A cognitive test is the most studied instrument in hiring, and the one with the largest subgroup difference of anything you can buy. Both numbers come out of the same 2022 table. This page prints it.
Two numbers decide whether a cognitive ability test belongs in your hiring process, and the same 2022 paper reports both of them. The first is that a cognitive test predicts job performance at about .31, not the .51 the category quoted for twenty years. The second is that it carries a subgroup difference of about .79 standard deviations, the largest of any common selection method. Most pages selling these tests print the first number and omit the second.
This page is written for the person deciding whether to buy one, not for a candidate preparing to sit one. It cites the primary rule or the primary study wherever one was reachable, because summaries of this material have drifted a long way from the source.
What the evidence says now
The reference point is Sackett, Zhang, Berry and Lievens, “Revisiting meta-analytic estimates of validity in personnel selection”, Journal of Applied Psychology (2022). It corrected a systematic overcorrection for range restriction that had inflated the older estimates, and the corrections moved most methods down by .10 to .20.
Their Table 3 pairs each method’s revised validity with its Black-White standardised mean difference, which is why it is the only table worth reproducing on a page like this one.
| Selection method | Schmidt and Hunter 1998 | Revised validity | Black-White d |
|---|---|---|---|
| Structured interview | .51 | .42 | .23 |
| Job knowledge test | .48 | .40 | .54 |
| Empirically keyed biodata | .35 | .38 | .33 |
| Work sample test | .54 | .33 | .67 |
| Cognitive ability test | .51 | .31 | .79 |
| Integrity test | .41 | .31 | .10 |
| Assessment centre | .37 | .29 | .52 |
| Unstructured interview | .38 | .19 | .32 |
| Conscientiousness, general | .31 | .19 | -.07 |
Read the first column against the second and the marketing claim collapses on its own: cognitive ability is no longer the focal predictor. The authors say so directly, proposing that structured interviews now be treated as the benchmark against which other methods are judged. They also point out what the four methods above cognitive ability have in common. Structured interviews, job knowledge tests, empirically keyed biodata and work samples are all job-specific measures. The general constructs start below them.
Two cautions on the d column, both from the paper’s own note. The cognitive ability figure of .79 comes from Roth and colleagues (2011) and reflects applicant samples. The job knowledge figure of .54 comes from Roth and colleagues (2003) and reflects incumbents, because no meta-analysis of applicant data exists for it; the authors expect the applicant figure to be at least that large and possibly larger. And the residual variability matters as much as the mean: the lower bound of the 80 percent credibility interval for cognitive ability is .13, so there are real settings where the relationship is close to nothing.
What a difference of .79 does to your selection rates
Effect sizes are abstract until they become pass rates. The table below is normal-curve arithmetic on a single assumption set: two groups with the same spread of scores whose means differ by .79 standard deviations, one cut score applied to everyone, no other criteria. It is an illustration of the shape of the problem rather than a prediction about your pipeline.
| If the cut passes this share of the higher-scoring group | The other group’s pass rate is | Impact ratio | Four-fifths rule |
|---|---|---|---|
| 90% | 68.9% | .77 | fails |
| 70% | 39.5% | .57 | fails |
| 50% | 21.5% | .43 | fails |
| 30% | 9.4% | .31 | fails |
| 10% | 1.9% | .19 | fails |
The four-fifths rule at 29 CFR 1607.4(D) treats a selection rate below 80 percent of the highest group’s rate as evidence of adverse impact. On these assumptions a cognitive cut score reaches an .80 ratio only when it is so low that it screens out roughly 8 percent of the higher-scoring group, which is to say when it is barely selecting at all. Selectivity and impact ratio move against each other, and the effect is mechanical.
Run the same arithmetic on the other methods at a median cut and the ordering is unmistakable: a work sample test lands near .50, a job knowledge test near .59, a structured interview near .82. None of them are free. The general measure is simply the most expensive on this axis.
A word on what the four-fifths rule is not. Its own provision says smaller differences may still constitute adverse impact where they are significant in statistical and practical terms, and larger differences may not where the numbers are too small to be significant. The EEOC calls it a rule of thumb. Clearing it is not clearance.
You also cannot compute any of this without applicant flow data. Sections 1607.4(A) and (B) expect employers to keep records of impact by sex and by race or ethnic group, and an impact review you are unable to perform is not a control.
The two corrections that are not available to you
This is the part that makes cognitive testing structurally different from other purchases, and it is why the decision cannot be deferred until results come in.
You may not adjust the scores. Section 106 of the Civil Rights Act of 1991, codified at 42 U.S.C. 2000e-2(l), makes it unlawful “to adjust the scores of, use different cutoff scores for, or otherwise alter the results of, employment related tests on the basis of race, color, religion, sex, or national origin.” The practice this outlawed, within-group score conversion, was once standard. Banding schemes that achieve the same thing by another route live in the same neighbourhood and are a question for counsel, not for a vendor.
You may not simply discard the results either. In Ricci v. DeStefano, 557 U.S. 557 (2009), New Haven threw out firefighter promotion exam results after seeing that white candidates had scored higher. The Court held that doing so violated Title VII: an employer needs a strong basis in evidence that it would face disparate-impact liability before taking race-conscious action to avoid that liability. Fear of a lawsuit is not enough.
Put those together and the sequence is forced. The validation work, the cut score, the role of the score in the decision and the applicant flow records all have to be settled before the first candidate takes the test, because afterwards the two obvious remedies are closed.
Why the validation bar is higher than it is for a skills test
A practical test can rest on content validity: the test visibly samples the work, so the job-relatedness argument is in the task. A cognitive test cannot use that argument at all, and the Guidelines name it explicitly.
Section 1607.14(C)(1) states that a selection procedure based upon inferences about mental processes cannot be supported solely or primarily on the basis of content validity, and lists the constructs it means: “intelligence, aptitude, personality, commonsense, judgment, leadership, and spatial ability.” Aptitude and intelligence sit in that list beside personality. An aptitude test therefore needs criterion-related evidence, meaning scores that empirically relate to outcomes in the job, or construct validity evidence, which the Guidelines themselves describe as an extensive and arduous research effort.
This is not a new position. Griggs v. Duke Power Co., 401 U.S. 424 (1971), the case that created disparate-impact doctrine, was about exactly this instrument: Duke Power required the Wonderlic Personnel Test, “which purports to measure general intelligence,” alongside a mechanical comprehension test, adopted without meaningful study of their relationship to job performance. The line the Court drew is still the clearest statement of the standard: “any tests used must measure the person for the job and not the person in the abstract.”
So when a vendor quotes a validity coefficient, the questions are which job, which outcome measure, and which population it was established in. A number from someone else’s population is not evidence about yours.
Three more places this lands
- Accommodation, under the ADA. A cognitive test is usually not a medical examination, unlike the clinical inventory at issue in Karraker v. Rent-A-Center, which is covered on the personality testing page. The live issue for ability testing is the timer. The ADA requires tests to be administered so that results reflect the skill being measured rather than an applicant’s sensory, manual or speaking impairment, and a speeded test is precisely where that obligation shows up. Decide the extended-time process before someone asks for it.
- Age. Smith v. City of Jackson, 544 U.S. 228 (2005), confirmed that disparate-impact claims exist under the ADEA, narrowed by a defence for reasonable factors other than age. The same selection-rate arithmetic applies to age bands, and it is worth running on a timed test.
- New York City. Local Law 144 attaches an annual independent bias audit, a published summary and ten business days of candidate notice to an automated employment decision tool used on a role in the city. Whether a scored assessment is in scope turns on whether the score is relied on solely, weighted more heavily than any other criterion, or used to override other conclusions. How you use the score determines the obligation.
The three words buyers use interchangeably
| Term | What it measures | What it inherits |
|---|---|---|
| Cognitive ability, GMA | Maximum performance on reasoning, numeric and verbal problems, usually timed | Validity of .31 and d of .79; the whole of this page |
| Aptitude | The same thing, sometimes narrowed to one domain such as numerical or mechanical | Named in 1607.14(C)(1) alongside intelligence; treated identically |
| Behavioural assessment | Typical conduct and preference rather than maximum performance | The personality family: lower validity, different ADA questions |
Vendors mix these labels freely, and the search results mix them worse. What matters is that the legal treatment follows the construct being measured, so a “behavioural” label on a trait questionnaire does not move it out of the personality analysis, and an “aptitude” label on a reasoning test does not move it out of this one.
Where SharpAssessment stands
The library is practical and job-knowledge tests, listed in full in the test library. A general cognitive score is not in it.
That ordering is a position on defensibility, not a queue. Section 1607.14(C)(1) puts aptitude and intelligence in the same category as personality, so neither can lean on content validity, and the moment either produces adverse impact you need criterion evidence gathered on roles and populations like yours. A cognitive module is worth building behind that evidence and not ahead of it, and the same table this page opens with is the reason the ordering costs nothing in predictive power: job knowledge tests at .40 and work samples at .33 already sit at or above the general measure at .31, with materially smaller subgroup differences.
If you need an audited cognitive instrument with decades of normative data for a role today, buy it from a vendor who has one. The established psychometric houses spent a long time accumulating that data and it is the real asset in this category. What is not worth buying from anyone is a reasoning quiz with a percentile attached and no answer to the question of which job, which outcome and which population.
For the adjacent decision on trait instruments, the personality test page covers the same ground for the Big Five and the ADA line, and the longer research background is in personality tests for hiring. For how the practical tests are put together and priced against the category, start with skills assessment software.
This page summarises business risk from primary sources and is not legal advice. Employment law is jurisdiction-specific and the rules covering automated hiring tools are changing quickly. Before running a cognitive or aptitude testing programme, take advice from a US employment lawyer.