Job knowledge test
Customer service assessment test
Judgement in realistic support situations: what to do first, what to escalate, what to promise.
- Time to complete
- About 15 minutes
- Format
- Six scenarios, ranked responses, one written justification
- Test family
- Job knowledge
- Included in
- All plans
What this test can see, and what an interview cannot
Testing customer service is harder than testing data entry, because there is no keystroke count to grade. What separates a good support hire from an adequate one is judgement under mild pressure: which of four waiting customers to answer first, when a policy exception is worth making, and when to stop improvising and escalate. None of that shows on a CV, and an interview only asks a candidate to describe a situation rather than sit in one.
This test puts them in the situation. Each item arrives the way it would at work, as a chat message or a ticket in a queue, with responses a real support team might plausibly choose between. None of them is a joke option, which is the usual giveaway in a weak scenario test. The candidate ranks the responses, then writes a short justification for the one they put first. The ranking carries the score; the justification is what you read when two candidates land on the same number.
The scenarios a candidate works through
A refund outside policy
A customer asks for a refund eleven days into a ten-day window, and has been paying you for three years.
What it separates Whether policy is a floor to reason from or a script to read out, and whether they know what they can decide without asking.
Four things at once
A broken checkout, an angry public post, a routine password reset and a large account asking something only engineering can answer.
What it separates Whether they prioritise by who is blocked and how many, or by whoever is loudest.
An angry opener
The first message contains an insult and a threat to go public, over something that is genuinely your fault.
What it separates Whether the apology comes before the explanation, and whether it is followed by a specific next step or a vague reassurance.
A colleague's wrong promise
The customer is quoting back something a teammate told them that was not true.
What it separates How they correct the record without either blaming the colleague to the customer or confirming the error to keep the peace.
An answer nobody has yet
The question needs engineering, engineering has not replied, and the customer wants a date.
What it separates Whether they invent a timeline, go quiet, or commit to an interim check-in they can actually keep.
The third contact about the same thing
A returning ticket with two previous conversations attached.
What it separates Whether they read the history before replying. A surprising number of candidates do not.
How an answer to a conflict scenario is scored
Responses are not marked right or wrong. Each option carries a weight running from the best available action down to the one that does the most damage, and the weights were set by experienced support staff rather than derived from the wording of the question. A candidate who picks a defensible second-best option scores close to the top. One who picks the option that ends the conversation at the customer's expense scores near zero, even when it is technically within policy.
On the conflict scenarios in particular, the refund and the angry opener, the ranking exposes three things and the score is built from them:
- Order of operations. Acknowledging the problem before explaining the constraint scores higher than the same two moves in the other order. Most weak answers contain the right content arranged the wrong way round.
- Size of the commitment. Promising a resolution the candidate cannot deliver is penalised at about the same weight as refusing outright. Both end the same way, one just ends later.
- Where the conversation lands. An answer that leaves the customer with a named next step and a time scores above one that resolves the emotion and nothing else.
The written justification is scored separately, by your team, against a rubric shown beside the answer. That part stays human on purpose: it is the only place in the test where you hear how the candidate sounds with nobody handing them options. Have two reviewers score the first batch independently before you trust one reviewer's numbers.
What separates a strong result from a weak one
| Reading the report for | A strong candidate | A weak one |
|---|---|---|
| De-escalation | Names the specific problem in the first line, then explains. | Opens with the policy, or with an apology so general it could be pasted into any ticket. |
| Exceptions | Makes the exception and says why it is defensible. | Either refuses on principle or grants whatever is asked for. |
| Escalation | Escalates the two items that need it and handles the rest. | Escalates almost nothing, or escalates almost everything. |
| Prioritisation | Ranks by who is blocked and how many. | Ranks by tone of voice. |
| Commitments | Gives a next step and a time that depends only on them. | Gives a resolution time that depends on somebody who has not answered yet. |
| Written justification | Two sentences with the reasoning visible. | Restates the chosen option in different words. |
Roles it suits
- Customer support, email and chat
- The closest fit, and the written justification doubles as a writing sample. If the role is mostly email, pair it with the business writing test, which scores tone on a full reply rather than on two sentences.
- Call centre and phone support
- The judgement transfers; nothing spoken is measured. Pace, handling an interruption and being intelligible on a bad line are not in here. Use it to cut a long list down, then spend fifteen minutes on the phone with what is left.
- Retail and front desk
- Weight the queue and de-escalation scenarios higher and the ticket-history one lower. Retail conflict happens in front of other customers and cannot be parked for an hour, so how someone opens matters more than it does in a ticket queue.
- Patient coordinator and healthcare front office
- The judgement transfers, but no scenario here touches a regulated situation. Add your own items on privacy and on what staff may say to a family member before you lean on the score.
How long it takes, and where it goes in the process
Fifteen minutes, and it is built to hold to that: six scenarios at roughly ninety seconds of reading and ranking each, plus a few minutes on the written justification. The per-scenario timer is generous enough that a careful reader is not punished for being careful, which would be an odd thing to select against in a support hire.
It sits after the application and before a phone screen. At fifteen minutes you can send it to a whole shortlist rather than to the three people you already liked, which is where most of its value shows up: the candidates it moves up are usually the ones a CV screen would have dropped, because judgement in a queue leaves almost no trace on a CV.
What it will not tell you
Answer keys encode one company's policy. The weights need reviewing against your own escalation rules before use.
It does not measure product knowledge, deliberately, because a candidate who has never seen your product should not be marked down for it. It also says nothing about the twentieth ticket of a shift, and shift-end judgement is what a good deal of support attrition is actually about.
And a candidate who has worked in support before recognises the shape of these items, so familiarity lifts scores a little. That matters when a career switcher is being compared against somebody with three years on a helpdesk, and it is the argument for reading the written justifications rather than sorting on the score.
Every test has a boundary like this. A test used outside it stops predicting anything, which is the usual reason a hiring team loses faith in testing altogether.
Frequently asked questions
What is a customer service assessment test?
How do you test customer service skills before hiring?
Is this a customer service aptitude test or a personality test?
Does it work for call centre hiring?
Are customer service simulations better than a scenario test?
Tests that pair with this one
Business writing test
Clarity, tone and correctness when writing a short work email or customer reply from a brief.
English proficiency test
Reading comprehension, grammar and written expression at the level ordinary office work needs.
Attention to detail test
Ability to spot discrepancies between documents and to follow multi step written instructions exactly.