Pre-employment assessment, measured against what you do now
Every guide to this subject lists the test types and concludes that testing is good. This one starts with the number nobody quotes: how well the resume screen and the unstructured interview you are running instead actually predict anything.
First, which of the three markets you are in
“Pre-employment testing” is used by three separate industries, and searching for it returns all of them mixed together.
| What people mean | What it answers | When it happens |
|---|---|---|
| Assessment | Can this person do the work | Before the hiring decision, on a shortlist |
| Background and reference screening | Is what they told you true | Usually after a conditional offer |
| Drug screening and physicals | A medical or safety question | After a conditional offer, for ADA reasons |
Only the first is on this page. The other two are different vendors, different budgets and a different rulebook: medical and disability-related enquiries are restricted before an offer under the ADA, which is why they sit where they sit. Our breakdown of who sells what is in pre-employment testing companies.
What an assessment is worth, against what you do now
The useful question is not whether assessment works. It is what you are comparing it against.
Most small teams decide between candidates on two things: a resume, which mainly conveys years of experience, and an interview that was not structured because nobody had time to structure it. Both have been measured, and this is the comparison the vendor pages leave out.
The figures below come from Sackett, Zhang, Berry and Lievens (2022) in the Journal of Applied Psychology, the paper that corrected a systematic overcorrection for range restriction which had inflated this entire literature for decades. It is why the familiar line that cognitive ability is the best predictor at .51 is now wrong.
| Selection method | Correlation with job performance |
|---|---|
| Structured interview | .42 |
| Job knowledge test | .40 * |
| Empirically keyed biodata | .38 |
| Work sample test | .33 * |
| Cognitive ability test | .31 |
| Integrity test | .31 |
| Assessment centre | .29 * |
| Situational judgement test | .26 |
| Conscientiousness questionnaire | .19 |
| Unstructured interview | .19 * |
| Years of job experience | .07 * |
Three things in that table should change how a small team spends its hiring effort.
Time served is not evidence. At .07, years of experience is the weakest thing on the list, and it is the single factor most resume screens are built around. Sorting a pile of applications by experience is close to sorting it at random.
The cheapest upgrade available is not a test. Structured interviewing is the top row, and it costs nothing except writing the questions down in advance and scoring them the same way for everyone. Buying an assessment and then keeping an unstructured interview afterwards means paying for a .33 signal and diluting it with a .19 one.
Tests that sample the job beat tests that measure a trait. Job knowledge at .40 and work samples at .33 sit above cognitive ability at .31 and well above conscientiousness at .19. That ordering is the reversal in the 2022 paper, and it has not reached most of the pages selling assessments. It also happens to line up with which tests are cheapest to defend, which is unusual and worth taking advantage of.
How not to misread those numbers
A correlation of .40 does not mean the test is right 40 percent of the time. It means that across many hires, higher scorers tended to perform better, with a lot of individual exceptions. At .40 you will still hire people who tested well and worked out badly.
These are averages across jobs and studies, not a property the instrument carries into your company. The structured interview figure at the top has an 80 percent credibility interval running from roughly .18 to .66, which is a very wide band. And validities do not add: running a work sample and a cognitive test together does not give you .33 plus .31, because the two overlap.
The honest reading is directional. Sample the work, structure the conversation, and stop treating time served as evidence.
Choosing between the practical test types, and matching them to a specific role, is a longer job that has its own page: skills assessment software and pre-employment skills tests covers what each type settles, which tests fit which roles, and how to read the scores when they come back.
Where the assessment goes in your process
Position matters more than buyers expect, because the same test produces a different result depending on what it is filtering.
- Application. Knockout questions only: right to work, the licence, the shift pattern. Those are facts, not assessments. Running a scored test on everyone who clicks apply mostly buys a large bill and a pile of unopened invitations.
- First screen, after a quick resume pass. This is where a work sample or job knowledge test earns its money. It replaces the phone screen that was going to be unstructured anyway, using the strongest evidence available at that point.
- Structured interview, informed by the results. Take the two weakest areas of the assessment into the conversation as questions. This is where an assessment becomes a hiring decision rather than a ranking.
- Reference and background checks, once an offer decision is close. Verification steps, not predictors.
The common mistake is putting a personality or cognitive questionnaire at step one because it is the cheapest thing to administer at volume. That spends the weakest signal on the largest population, and it is the version that is hardest to explain when a rejected candidate asks why.
The conditions that decide whether it holds up
Almost none of the pages ranking for this search mention any of this, and it is short.
Job-relatedness is the whole test. The EEOC’s guidance on employment tests and selection procedures and the underlying Uniform Guidelines ban nothing by name. They say that where a procedure screens out a protected group at a materially lower rate, you must show it is job-related and consistent with business necessity.
Content validity only covers half the category. A test that visibly samples the work can rest on it. Section 1607.14(C)(1) states that a procedure measuring a construct, and it names personality, intelligence, aptitude and judgement, cannot be supported primarily by content validity. Those need criterion-related evidence gathered on a population like yours.
You cannot detect a problem you are not measuring. Knowing whether your assessment has adverse impact requires applicant flow data by sex and by race or ethnicity. Without it you are not in a position to know.
Accommodation applies during testing. A timed test administered without adjustment can end up measuring an impairment rather than the skill.
Some jurisdictions add notice and audit duties. New York City’s Local Law 144 requires an annual independent bias audit, a published summary and ten business days’ notice for automated employment decision tools used on roles there. Illinois, Colorado and the EU each have rules covering the AI end of this category. The detail, with citations, is on the personality testing page, because that is where the exposure concentrates.
This is a summary of business risk and not legal advice.
When not to run an assessment
Skip it when the shortlist is three people you already know can do the work. An assessment is a comparison instrument. With a small, well-qualified pool and a role you understand, it adds a week and a drop-off risk without adding information.
Three more cases where testing is the wrong call:
- When you cannot say what the test is for. “It seemed like good practice” is how a hiring process acquires a step nobody can defend and nobody can remove.
- When the market is tighter than your funnel. For a role where three qualified people apply a month, an unpaid forty minute task at the front of the process costs more candidates than it screens.
- When the score will not change the decision. If you already know who you are hiring, running everyone through an assessment to document it is theatre, and a rejected candidate’s lawyer can read it that way too.
Why the billing unit decides the price
Published pricing in this category is unreliable, and the reason is that vendors do not charge for the same thing. Seats, active jobs, candidates, credits and annual licences all appear, so one product can be several times cheaper or more expensive per candidate than another depending on the shape of your hiring rather than on the sticker price.
The pattern that catches small teams is the annual contract, which prices as though you hire continuously. Hire eight people in March and nobody until November, and you have funded a year of software to use it for one month. What fourteen vendors actually publish is written up in talent assessment tools.
Where SharpAssessment fits
We sell the two rows of that table a small team can defend without a validation budget: practical work samples and job knowledge tests, scored consistently and reported side by side. One link goes out, every candidate gets the same task and the same scoring, and you keep a record of how the decision was made. Every test in the library states what it measures, how it is scored and where it stops being useful.
We do not sell a personality or cognitive instrument, and that is a position rather than a gap. Those need criterion validity and real normative data, which takes decades to collect, and shipping a questionnaire without them is exactly what leaves a hiring team unable to explain a rejection.
Pricing is by the volume of candidates you assess. No seats, no per-job fee, no annual contract, because a team that hires in bursts should not have to fund eleven quiet months.
If you have a talent assessment function and a validation budget, the established psychometric vendors are the right buy, and the comparison page names them.