Personality testing for hiring, with the legal part left in
Most pages selling personality assessments skip the rules that decide whether you can defend the decision afterwards. This one starts there, with citations you can check.
Personality testing in hiring attracts confident opinions in both directions, and the defensible position is narrower than either camp finds comfortable. These instruments measure something real. What they measure is only weakly connected to job performance. And using one as an early filter, removing candidates before anyone has looked at whether they can do the work, spends your weakest signal first and is the hardest version to defend afterwards.
The legal and statistical claims below cite the primary rule or the primary study wherever one was reachable, and say so explicitly where only secondary sources were available. That distinction matters here, because secondary summaries of this material have drifted a long way from the text.
What the evidence actually says
Conscientiousness is the trait that holds up best across job types. The current reference point for how well it predicts job performance is the 2022 re-analysis by Sackett, Zhang, Berry and Lievens in the Journal of Applied Psychology, which corrected a systematic overcorrection in the older meta-analytic tradition. Their estimate is about .19 in general, and about .25 when items are framed around behaviour at work rather than behaviour in general.
For scale, the same paper revised the long-quoted figure for cognitive ability tests down from .51 to .31. The Big Five traits other than conscientiousness land between roughly .05 and .10 in their general forms, and between about .12 and .23 in work-framed forms.
Two things follow. First, a .19 correlation is a real relationship, and Barrick and Mount’s original 1991 meta-analysis reported a similar magnitude at about .22, so this is not a case of a finding collapsing. Second, the 80 percent credibility interval around that conscientiousness estimate reaches down to roughly .02, which means there are job and measurement contexts where the true relationship is close to nothing. A number that varies that much between settings is not a property the instrument carries with it into your company. It is a claim that has to be established where you intend to use it.
That is also the professional standard rather than an opinion. SIOP’s Principles for the Validation and Use of Personnel Selection Procedures define validity as the degree to which evidence supports a specific interpretation of scores for a proposed use. Validity belongs to a use, not to a product.
If you take one thing from this page: be suspicious of any vendor that states a validity figure without naming the job, the outcome measured, and the population it was established in.
The four rules that decide whether you can defend it
1. Adverse impact, and what the four-fifths rule really says
If a selection procedure passes one group at a materially lower rate than another, the EEOC’s Uniform Guidelines on Employee Selection Procedures put the burden on the employer to show the procedure is job-related and consistent with business necessity.
The four-fifths rule, at 29 CFR 1607.4(D), is the familiar shorthand: a selection rate below four-fifths, or 80 percent, of the highest group’s rate will generally be regarded by federal enforcement agencies as evidence of adverse impact. The words that get dropped matter more than the number. The same provision says smaller differences may still constitute adverse impact where they are significant in statistical and practical terms, and that larger differences may not, where they rest on small numbers that are not statistically significant. The EEOC has described it in its own guidance as a rule of thumb, and courts have held that passing it does not settle the question.
So four-fifths is a screening heuristic, not a safe harbour. Treating an impact ratio above 80 percent as legal clearance is a common overstatement in this category.
There is a practical consequence that gets overlooked. You cannot compute an impact ratio without applicant flow data by sex and by race or ethnicity. Sections 1607.4(A) and (B) expect employers to maintain exactly that, and section 1607.15(A)(1) offers a simplified version for employers with 100 or fewer people. If you are not collecting it, you are not in a position to know whether your test has a problem. The Guidelines also let enforcement agencies infer adverse impact from missing records where a group is underused in the job category, so the gap is not neutral.
2. Personality is a construct, so content validity will not carry it
This is the most useful thing on this page and the least discussed.
A practical skills test can use a content validity strategy, but only to the extent that it representatively samples important parts of the job. A bookkeeping exercise for a bookkeeping role is a strong candidate for that argument, though the argument still has to be made rather than assumed. Personality does not work that way at all, and the Guidelines say so directly. Section 1607.14(C)(1) states that a selection procedure based on inferences about mental processes cannot be supported solely or primarily by content validity, and it names the constructs it means, personality among them, alongside intelligence, aptitude, judgement and leadership.
The consequence is concrete. A personality instrument needs criterion-related evidence, meaning scores empirically relate to job outcomes, or construct validity evidence, which the Guidelines themselves describe as an extensive and arduous research effort. A panel of experts confirming the questions look job-relevant is content validity wearing a different hat, and it is the wrong instrument for the job.
3. The ADA line: when a test becomes a medical examination
The ADA bars medical examinations and disability-related inquiries before a job offer, at 42 U.S.C. 12112(d)(2). It also requires that where an applicant has a disability affecting sensory, manual or speaking skills, tests be administered so the results reflect what the test means to measure rather than that impairment, and it constrains qualification standards that screen out people with disabilities.
Whether a personality test crosses into medical examination territory is a fact-specific question, not a category rule. The EEOC’s enforcement guidance on pre-employment inquiries and medical examinations sets out a multi-factor test, including whether the test is designed to reveal an impairment, whether a health professional interprets it, and whether it is normally administered in a clinical setting. The guidance is explicit in both directions: a psychological test is medical if it yields evidence that would lead to identifying a mental disorder, and a test designed and used to measure only things such as honesty, tastes and habits is not.
Karraker v. Rent-A-Center, 411 F.3d 831 (7th Circuit, 2005), is the decision people cite. Rent-A-Center required a promotion battery that included 502 questions from the MMPI, an instrument scored on clinical scales covering traits such as depression, paranoia and mania. The court held that because the MMPI is designed at least in part to reveal mental illness, and its use had the effect of harming the prospects of people with mental disabilities, it was best categorised as a medical examination, and its use violated the ADA. Notably, the court reached that conclusion even though no psychologist interpreted the results.
Read the scope carefully, because both popular readings are wrong. The court did not hold that personality tests are medical examinations. It applied a multi-factor test to one instrument built on clinical scales, and the framework it adopted expressly contemplates personality tests that are not medical. It is also binding precedent only in the Seventh Circuit, though it is cited well beyond it.
The workable takeaway: an instrument that reports on psychopathology is a different legal object from one that reports on work-relevant dispositions, and you should know which one you have bought. Accommodation obligations apply during testing either way.
4. New York City: an annual bias audit and ten business days
If the role is in New York City, Local Law 144 may apply, enforced by the Department of Consumer and Worker Protection since 5 July 2023. Where it applies, it requires three things: a bias audit by an independent auditor conducted no more than a year before use, a public summary of that audit on the employer’s website before the tool is used, and notice to the candidate at least ten business days ahead, covering the qualifications assessed and the option to request an alternative process or an accommodation.
The audit arithmetic borrows directly from the Uniform Guidelines: selection rates and impact ratios by sex, by race or ethnicity, and by the intersection of the two. Penalties run up to 500 dollars for a first violation and for each additional violation the same day, then 500 to 1,500 dollars for each subsequent violation, with each day of non-compliant use and each failure to give notice counted separately.
Whether your assessment is in scope is the part almost every summary gets wrong. A scored assessment clearly meets the first half of the definition of an automated employment decision tool, since it is a computational process issuing a simplified output. But the city’s final rules narrow the phrase “substantially assist or replace discretionary decision making” to three specific patterns: relying solely on the score, weighting it more heavily than any other criterion, or using it to override conclusions drawn from other factors. A score genuinely used as one non-dominant input among several human-reviewed factors can fall outside that definition.
That is a real distinction rather than a loophole, and it cuts both ways. It means how you use the tool determines your obligations, so the usage decision and the compliance decision are the same decision. Assume coverage unless you can show the score is not doing the deciding.
Where else this bites, as of July 2026
These rules move quickly, so treat the dates as of this page’s date and verify before relying on them.
- Illinois, and this one rests on secondary sources because the state’s own legislative text was unreachable when we checked: HB 3773 amended the Human Rights Act effective 1 January 2026, making discriminatory use of AI in employment decisions a civil rights violation and requiring notice when AI is used, and the Artificial Intelligence Video Interview Act has required notice, explanation and consent for AI analysis of interview videos since 2020. The implementing rules for the newer amendment were unsettled, and we could not confirm their current status.
- Colorado is the trap. The 2024 law most articles still describe, SB 24-205, was repealed and replaced rather than merely delayed. The operative statute is now SB 26-189, signed 14 May 2026 and effective 1 January 2027.
- The EU AI Act classifies employment and worker-management uses as high risk. The date commonly quoted for those obligations, 2 August 2026, has moved: the high-risk obligations for employment now apply from 2 December 2027 following the simplification package that took effect in July 2026. General transparency duties still land in August 2026. Separately, emotion recognition in the workplace is a prohibited practice, not merely a high-risk one, which matters for any tool inferring emotional state from face or voice.
- Maryland requires applicant consent for facial recognition during interviews. That reaches a product doing facial analysis, not a questionnaire.
Which frameworks survive the selection question
| Framework | Built for | Use in selection | Main caution |
|---|---|---|---|
| Big Five, work-framed | Trait measurement | The form the validity evidence above actually measures | Validity is modest, and still needs evidence for your own roles |
| Big Five, generic | Trait measurement | Measurably weaker than work-framed versions | Conscientiousness at about .19 here against about .25 work-framed |
| DISC | Communication style | Mostly used for team work | Not built as a selection predictor |
| MBTI | Self-understanding | Built for development, not for selection | Its design purpose is insight, not prediction |
| Clinical inventories, MMPI type | Psychological diagnosis | Avoid for pre-offer hiring | ADA medical examination exposure, as in Karraker |
The pattern worth noticing: the frameworks with the widest name recognition are largely the ones designed for something other than hiring.
How to use a personality test without creating a problem
- Never as the first gate. Screening on personality before you have looked at whether someone can do the work discards candidates on your weakest signal. If you use it, use it after a work sample.
- Never as a knockout. A trait score predicts a tendency, not an outcome. A candidate scoring low on extraversion is not disqualified from sales.
- As a structured interview prompt. The most defensible use is turning a profile into better questions rather than into a ranking.
- Design the impact check in from the start. Collect applicant flow data from day one. An adverse impact review you cannot perform is not a control.
- Know what your instrument measures. If it reports on anything resembling psychopathology, treat the ADA question as live before you use it pre-offer.
- Ask the validation question in writing. Which job, which outcome, which population, and what kind of validity evidence. A vendor who cannot answer is telling you something.
Where the work produces checkable output, test the output. A candidate reconciling a messy ledger tells you more about a bookkeeping hire than a conscientiousness percentile, and it is far easier to explain to a candidate who asks why they were turned down. Personality measurement earns its place where the work resists direct sampling, or as context for a conversation.
Where SharpAssessment stands
We do not sell a personality test, and that is a deliberate position rather than a gap in the roadmap.
Section 1607.14(C)(1) is the reason. A trait instrument cannot lean on content validity at all, so the moment it produces adverse impact you need criterion validity: real outcome data, gathered on a population like yours. Vendors who have that spent decades collecting it. Shipping a Big Five questionnaire assembled from public-domain items and calling it a hiring instrument skips exactly the part that makes it defensible.
What we sell instead is the kind of test where job-relatedness is visible in the task: practical and job-knowledge work samples, listed in the test library. If a psychometric score has to carry real weight in a decision you might have to defend, the established vendors with long-accumulated normative data are the right buy, and our TestGorilla alternatives page covers that market. The longer explainer on personality tests for hiring covers the research background.
None of this makes personality testing off limits. It means the compliance work is part of the cost of using it, and a vendor who cannot discuss that work is quoting you an incomplete price.
This page summarises business risk from primary sources and is not legal advice. Employment law is jurisdiction-specific and the AI-related rules above are changing quickly. Before running any assessment programme, particularly a psychometric one, take advice from a US employment lawyer.