SharpAssessmentGet started

Personality testing for hiring, with the legal part left in

Most pages selling personality assessments skip the rules that decide whether you can defend the decision afterwards. This one starts there, with citations you can check.

Personality testing in hiring attracts confident opinions in both directions, and the defensible position is narrower than either camp finds comfortable. These instruments measure something real. What they measure is only weakly connected to job performance. And using one as an early filter, removing candidates before anyone has looked at whether they can do the work, spends your weakest signal first and is the hardest version to defend afterwards.

The legal and statistical claims below cite the primary rule or the primary study wherever one was reachable, and say so explicitly where only secondary sources were available. That distinction matters here, because secondary summaries of this material have drifted a long way from the text.

What the evidence actually says

Conscientiousness is the trait that holds up best across job types. The current reference point for how well it predicts job performance is the 2022 re-analysis by Sackett, Zhang, Berry and Lievens in the Journal of Applied Psychology, which corrected a systematic overcorrection in the older meta-analytic tradition. Their estimate is about .19 in general, and about .25 when items are framed around behaviour at work rather than behaviour in general.

For scale, the same paper revised the long-quoted figure for cognitive ability tests down from .51 to .31. The Big Five traits other than conscientiousness land between roughly .05 and .10 in their general forms, and between about .12 and .23 in work-framed forms.

Two things follow. First, a .19 correlation is a real relationship, and Barrick and Mount’s original 1991 meta-analysis reported a similar magnitude at about .22, so this is not a case of a finding collapsing. Second, the 80 percent credibility interval around that conscientiousness estimate reaches down to roughly .02, which means there are job and measurement contexts where the true relationship is close to nothing. A number that varies that much between settings is not a property the instrument carries with it into your company. It is a claim that has to be established where you intend to use it.

That is also the professional standard rather than an opinion. SIOP’s Principles for the Validation and Use of Personnel Selection Procedures define validity as the degree to which evidence supports a specific interpretation of scores for a proposed use. Validity belongs to a use, not to a product.

If you take one thing from this page: be suspicious of any vendor that states a validity figure without naming the job, the outcome measured, and the population it was established in.

The four rules that decide whether you can defend it

1. Adverse impact, and what the four-fifths rule really says

If a selection procedure passes one group at a materially lower rate than another, the EEOC’s Uniform Guidelines on Employee Selection Procedures put the burden on the employer to show the procedure is job-related and consistent with business necessity.

The four-fifths rule, at 29 CFR 1607.4(D), is the familiar shorthand: a selection rate below four-fifths, or 80 percent, of the highest group’s rate will generally be regarded by federal enforcement agencies as evidence of adverse impact. The words that get dropped matter more than the number. The same provision says smaller differences may still constitute adverse impact where they are significant in statistical and practical terms, and that larger differences may not, where they rest on small numbers that are not statistically significant. The EEOC has described it in its own guidance as a rule of thumb, and courts have held that passing it does not settle the question.

So four-fifths is a screening heuristic, not a safe harbour. Treating an impact ratio above 80 percent as legal clearance is a common overstatement in this category.

There is a practical consequence that gets overlooked. You cannot compute an impact ratio without applicant flow data by sex and by race or ethnicity. Sections 1607.4(A) and (B) expect employers to maintain exactly that, and section 1607.15(A)(1) offers a simplified version for employers with 100 or fewer people. If you are not collecting it, you are not in a position to know whether your test has a problem. The Guidelines also let enforcement agencies infer adverse impact from missing records where a group is underused in the job category, so the gap is not neutral.

2. Personality is a construct, so content validity will not carry it

This is the most useful thing on this page and the least discussed.

A practical skills test can use a content validity strategy, but only to the extent that it representatively samples important parts of the job. A bookkeeping exercise for a bookkeeping role is a strong candidate for that argument, though the argument still has to be made rather than assumed. Personality does not work that way at all, and the Guidelines say so directly. Section 1607.14(C)(1) states that a selection procedure based on inferences about mental processes cannot be supported solely or primarily by content validity, and it names the constructs it means, personality among them, alongside intelligence, aptitude, judgement and leadership.

The consequence is concrete. A personality instrument needs criterion-related evidence, meaning scores empirically relate to job outcomes, or construct validity evidence, which the Guidelines themselves describe as an extensive and arduous research effort. A panel of experts confirming the questions look job-relevant is content validity wearing a different hat, and it is the wrong instrument for the job.

3. The ADA line: when a test becomes a medical examination

The ADA bars medical examinations and disability-related inquiries before a job offer, at 42 U.S.C. 12112(d)(2). It also requires that where an applicant has a disability affecting sensory, manual or speaking skills, tests be administered so the results reflect what the test means to measure rather than that impairment, and it constrains qualification standards that screen out people with disabilities.

Whether a personality test crosses into medical examination territory is a fact-specific question, not a category rule. The EEOC’s enforcement guidance on pre-employment inquiries and medical examinations sets out a multi-factor test, including whether the test is designed to reveal an impairment, whether a health professional interprets it, and whether it is normally administered in a clinical setting. The guidance is explicit in both directions: a psychological test is medical if it yields evidence that would lead to identifying a mental disorder, and a test designed and used to measure only things such as honesty, tastes and habits is not.

Karraker v. Rent-A-Center, 411 F.3d 831 (7th Circuit, 2005), is the decision people cite. Rent-A-Center required a promotion battery that included 502 questions from the MMPI, an instrument scored on clinical scales covering traits such as depression, paranoia and mania. The court held that because the MMPI is designed at least in part to reveal mental illness, and its use had the effect of harming the prospects of people with mental disabilities, it was best categorised as a medical examination, and its use violated the ADA. Notably, the court reached that conclusion even though no psychologist interpreted the results.

Read the scope carefully, because both popular readings are wrong. The court did not hold that personality tests are medical examinations. It applied a multi-factor test to one instrument built on clinical scales, and the framework it adopted expressly contemplates personality tests that are not medical. It is also binding precedent only in the Seventh Circuit, though it is cited well beyond it.

The workable takeaway: an instrument that reports on psychopathology is a different legal object from one that reports on work-relevant dispositions, and you should know which one you have bought. Accommodation obligations apply during testing either way.

4. New York City: an annual bias audit and ten business days

If the role is in New York City, Local Law 144 may apply, enforced by the Department of Consumer and Worker Protection since 5 July 2023. Where it applies, it requires three things: a bias audit by an independent auditor conducted no more than a year before use, a public summary of that audit on the employer’s website before the tool is used, and notice to the candidate at least ten business days ahead, covering the qualifications assessed and the option to request an alternative process or an accommodation.

The audit arithmetic borrows directly from the Uniform Guidelines: selection rates and impact ratios by sex, by race or ethnicity, and by the intersection of the two. Penalties run up to 500 dollars for a first violation and for each additional violation the same day, then 500 to 1,500 dollars for each subsequent violation, with each day of non-compliant use and each failure to give notice counted separately.

Whether your assessment is in scope is the part almost every summary gets wrong. A scored assessment clearly meets the first half of the definition of an automated employment decision tool, since it is a computational process issuing a simplified output. But the city’s final rules narrow the phrase “substantially assist or replace discretionary decision making” to three specific patterns: relying solely on the score, weighting it more heavily than any other criterion, or using it to override conclusions drawn from other factors. A score genuinely used as one non-dominant input among several human-reviewed factors can fall outside that definition.

That is a real distinction rather than a loophole, and it cuts both ways. It means how you use the tool determines your obligations, so the usage decision and the compliance decision are the same decision. Assume coverage unless you can show the score is not doing the deciding.

Where else this bites, as of July 2026

These rules move quickly, so treat the dates as of this page’s date and verify before relying on them.

  • Illinois, and this one rests on secondary sources because the state’s own legislative text was unreachable when we checked: HB 3773 amended the Human Rights Act effective 1 January 2026, making discriminatory use of AI in employment decisions a civil rights violation and requiring notice when AI is used, and the Artificial Intelligence Video Interview Act has required notice, explanation and consent for AI analysis of interview videos since 2020. The implementing rules for the newer amendment were unsettled, and we could not confirm their current status.
  • Colorado is the trap. The 2024 law most articles still describe, SB 24-205, was repealed and replaced rather than merely delayed. The operative statute is now SB 26-189, signed 14 May 2026 and effective 1 January 2027.
  • The EU AI Act classifies employment and worker-management uses as high risk. The date commonly quoted for those obligations, 2 August 2026, has moved: the high-risk obligations for employment now apply from 2 December 2027 following the simplification package that took effect in July 2026. General transparency duties still land in August 2026. Separately, emotion recognition in the workplace is a prohibited practice, not merely a high-risk one, which matters for any tool inferring emotional state from face or voice.
  • Maryland requires applicant consent for facial recognition during interviews. That reaches a product doing facial analysis, not a questionnaire.

Which frameworks survive the selection question

FrameworkBuilt forUse in selectionMain caution
Big Five, work-framedTrait measurementThe form the validity evidence above actually measuresValidity is modest, and still needs evidence for your own roles
Big Five, genericTrait measurementMeasurably weaker than work-framed versionsConscientiousness at about .19 here against about .25 work-framed
DISCCommunication styleMostly used for team workNot built as a selection predictor
MBTISelf-understandingBuilt for development, not for selectionIts design purpose is insight, not prediction
Clinical inventories, MMPI typePsychological diagnosisAvoid for pre-offer hiringADA medical examination exposure, as in Karraker

The pattern worth noticing: the frameworks with the widest name recognition are largely the ones designed for something other than hiring.

How to use a personality test without creating a problem

  • Never as the first gate. Screening on personality before you have looked at whether someone can do the work discards candidates on your weakest signal. If you use it, use it after a work sample.
  • Never as a knockout. A trait score predicts a tendency, not an outcome. A candidate scoring low on extraversion is not disqualified from sales.
  • As a structured interview prompt. The most defensible use is turning a profile into better questions rather than into a ranking.
  • Design the impact check in from the start. Collect applicant flow data from day one. An adverse impact review you cannot perform is not a control.
  • Know what your instrument measures. If it reports on anything resembling psychopathology, treat the ADA question as live before you use it pre-offer.
  • Ask the validation question in writing. Which job, which outcome, which population, and what kind of validity evidence. A vendor who cannot answer is telling you something.

Where the work produces checkable output, test the output. A candidate reconciling a messy ledger tells you more about a bookkeeping hire than a conscientiousness percentile, and it is far easier to explain to a candidate who asks why they were turned down. Personality measurement earns its place where the work resists direct sampling, or as context for a conversation.

Where SharpAssessment stands

We do not sell a personality test, and that is a deliberate position rather than a gap in the roadmap.

Section 1607.14(C)(1) is the reason. A trait instrument cannot lean on content validity at all, so the moment it produces adverse impact you need criterion validity: real outcome data, gathered on a population like yours. Vendors who have that spent decades collecting it. Shipping a Big Five questionnaire assembled from public-domain items and calling it a hiring instrument skips exactly the part that makes it defensible.

What we sell instead is the kind of test where job-relatedness is visible in the task: practical and job-knowledge work samples, listed in the test library. If a psychometric score has to carry real weight in a decision you might have to defend, the established vendors with long-accumulated normative data are the right buy, and our TestGorilla alternatives page covers that market. The longer explainer on personality tests for hiring covers the research background.

None of this makes personality testing off limits. It means the compliance work is part of the cost of using it, and a vendor who cannot discuss that work is quoting you an incomplete price.

This page summarises business risk from primary sources and is not legal advice. Employment law is jurisdiction-specific and the AI-related rules above are changing quickly. Before running any assessment programme, particularly a psychometric one, take advice from a US employment lawyer.

Frequently asked questions

Are personality tests legal for hiring in the United States?
There is no federal law banning them. What the law does is attach conditions. If a test screens out a protected group at a materially lower rate, the EEOC Uniform Guidelines expect the employer to show the test is job-related and consistent with business necessity. If the instrument is built to surface a mental impairment, the ADA restricts using it before a job offer. If it is used in New York City in a way that substantially assists or replaces human judgement, Local Law 144 adds an annual bias audit and advance notice to candidates. Legal to use, conditional to defend. This is a summary of business risk and not legal advice.
Do personality tests predict job performance?
Weakly, and less than the older literature claimed. The current meta-analytic reference point (Sackett, Zhang, Berry and Lievens, Journal of Applied Psychology, 2022) puts conscientiousness at roughly .19 for predicting job performance, rising to about .25 when the questions are framed explicitly around behaviour at work. The other four Big Five traits sit lower. That is a real signal and a small one. It is not a basis for ranking candidates on its own, and any vendor quoting a validity number should be asked which job, which outcome measure, and which population it was established in.
Which personality framework is most defensible for selection?
The Big Five, also called the five factor model, is the framework the meta-analytic validity evidence actually covers, and work-framed versions of it measure better than generic ones. MBTI was built for self-understanding and development rather than for selection. DISC is mostly used for team communication. Clinical inventories such as the MMPI are a different category entirely and carry the ADA exposure that the Karraker case turned on. Whichever you pick, the validity question attaches to your use of it rather than to the brand.
What does New York City Local Law 144 actually require?
For an automated employment decision tool used on a role in New York City, it requires a bias audit by an independent auditor conducted within the previous year, publication of a summary of that audit on the employer website before the tool is used, and notice to the candidate at least ten business days in advance, including the qualifications the tool assesses and an opportunity to request an alternative process. Penalties run up to 500 dollars for a first violation and for any further violation the same day, then 500 to 1,500 dollars for each subsequent violation. Each day of non-compliant use counts separately, and so does each failure to give a required notice. Whether a particular scored assessment is in scope depends on how the score is used, which is the part most summaries get wrong.
Can I validate a personality test the way I would validate a skills test?
No, and this is the most common technical mistake in the category. A skills test can rest on content validity, meaning the test visibly samples the work itself. The Uniform Guidelines name personality explicitly as a construct and state that a content validity strategy is not appropriate for procedures that purport to measure traits or constructs. A personality instrument needs criterion-related or construct validity evidence instead, which is slower and more expensive to produce. Subject matter experts agreeing that the questions look relevant is not validation.
Can candidates fake a personality test?
To a degree, yes, and they do. Meta-analytic comparisons of real job applicants against people with nothing at stake find applicants score measurably higher on the traits they think are wanted, with the effect strongest on conscientiousness and emotional stability. We are working from secondary summaries of that literature rather than the underlying papers, and published estimates of the size of the effect vary a lot by study design, so treat the direction as solid and the magnitude as indicative. Applicants can manage the impression they give whatever the question format. That is a reason to treat a profile as one input into a conversation rather than as a score to sort by.