Skip to content
← THE JOURNAL

Verify Candidate Skills in the AI Era: Designing Assessments People Can’t Fake

AI · SEPTEMBER 2026 · 12 MIN READ

By The Humavera Team

verify candidate skills ai era

A recruiter told me last spring that candidates were pasting her assessment questions into ChatGPT, and asked what she could do to stop it. The honest answer was not the one she wanted: nothing, and that is not the useful question anyway. The useful question is what her assessment was measuring that a model can now do in nine seconds.

If you want to verify candidate skills in the AI era, the first move is to stop trying to keep AI out of the room. The assessments that survive are the ones where AI assistance is either irrelevant to the answer or plainly visible in it. Surveillance buys you about a year. Design buys you the decade.

This is a real cost, not a hypothetical one. In Robert Half research released in March 2026 — fielded November 2025 among more than 2,000 U.S. hiring managers — 65% said the surge in applications had made verifying candidate skills harder. What follows is four design moves you can apply to an assessment you already have, and what each one costs you in candidate time and reviewer time, because they all cost something and pieces that pretend otherwise are not written by people who have run them.

Where this sits in your funnel — what happens to the pile before anyone reaches an assessment — is a separate problem [pending: T-01].

What broke, precisely

“Assessments are broken” is too broad to act on. Some item types broke completely and others did not move at all, and the boundary is learnable.

Broke: knowledge-recall items. Anything with a correct answer that exists in public text. Definitional questions. “Explain the difference between X and Y.” Written-response items where the score tracks how well the answer is composed. Take-home exercises that reward polish and structure, which is most take-homes, because polish is exactly what a model adds.

Did not break: items requiring the candidate’s own context — their team, their last escalation, their actual decision. Items with genuine trade-offs and no clean right answer. Anything where the reasoning is observed live rather than inferred from a finished artefact.

The pattern underneath: if the value of the answer lives in the artefact, it broke. If the value lives in the process that produced it, it did not.

Watching this happen across assessment tooling was uncomfortable, because the items that stopped discriminating first were the well-built ones. A clean, unambiguous, carefully-worded knowledge item is precisely the item a model handles perfectly. Ambiguity, context-dependence and messiness — the things assessment design has spent decades trying to remove — turned out to be the properties that survived.

Which means a lot of what got graded as good assessment practice was measuring articulacy plus recall, and we could not see it while articulacy was expensive.

The proctoring answer, and why I don’t buy it

The SERP for skills assessment and AI cheating is dominated by lockdown browsers, webcam monitoring and keystroke analysis. Let me be direct: for a mid-market company, proctoring is the wrong purchase, and I would say that even if it were free.

It detects the wrong thing. Proctoring establishes that a person sat alone in a room. It does not establish that they can do the job, which is the question. You have spent budget and candidate goodwill to verify a proxy for a proxy.

It penalises the wrong people. Webcam requirements land hardest on candidates in shared housing, on unreliable connections, in different time zones, with caring responsibilities, and on neurodivergent candidates whose natural behaviour — looking away while thinking, moving, reading aloud — trips exactly the flags these systems raise. Your false positives are not distributed randomly. They concentrate on people your process already handles badly.

And it is an arms race with a structural loser. Every countermeasure has a workaround within months, and you are paying a subscription to stay one version behind. A second device pointed at a screen defeats most of it, and always will.

I am arguing against the category on its merits, not against any particular vendor. Some of these products are well built. They are well built for a question you should stop asking.

The retreat to in-person testing is more defensible, and worth naming honestly rather than dismissing. It genuinely works. It also costs you every candidate who cannot travel on a Tuesday, which in practice means parents, people currently employed, and anyone outside your city. That is a real trade, and for some roles it is worth making. Just make it deliberately, knowing you are buying verification with pool size, rather than sliding into it because the alternative felt hard.

Four design moves

Each of these works on an existing assessment. None requires a platform.

Ask for judgement, not answers

Build the item around a situation with a genuine trade-off and no clean right answer. An ops analyst gets a shipment that will miss a contractual deadline; they can expedite at a cost that blows the quarter’s freight budget, partial-ship and split the invoice, or tell the customer today. All three are defensible. Ask which they would do and why, in five sentences.

A model will produce a competent answer to that. That is fine — it is not what you are scoring. You are scoring which stakeholder they chose to disappoint and whether their reasoning survives a follow-up question, and a model cannot produce this candidate’s judgement about which risk they are willing to own.

Cost: harder to score. You need a rubric with three or four levels and at least two people calibrated against it, or you have converted an objective item into an opinion.

Anchor it in the candidate’s own experience

Behavioural questions requiring a specific personal example, with follow-ups that go two layers past the prepared answer. Not “tell me about a time you handled conflict” — that has been answerable by a model for years, and by a decent candidate with prep for decades. Instead: “Walk me through the last time you had to tell a manager something they did not want to hear. What did you say first? What did they say back? What did you do the next day?”

The Willo Hiring Trends Report 2026 found live behavioural interviews with real examples were the most trusted talent indicator among respondents, cited by 68%. Worth flagging the sample: just over 100 hiring professionals worldwide, collected by an interview-software vendor with a commercial interest in that conclusion. Directional, not definitive. But it matches what practitioners describe, which is that the follow-up is where the interview actually happens.

Cost: interviewer time, and it collapses without follow-ups. An untrained interviewer who accepts the first polished answer gets nothing from this method at all.

Make it live, not take-home

A short working session — twenty to thirty minutes — where you watch the thinking rather than reading its output. An HR coordinator gets a messy leave-balance dispute with three contradictory records and talks through how they would resolve it. A finance associate is handed a reconciliation that does not tie and reasons out loud about where to look first.

You are not testing whether they arrive at the answer. You are watching what they check first, what they ask you, and what they do when the obvious approach fails. That is unfakeable in a way no artefact is.

Cost: scheduling, which is genuinely the biggest constraint for a small team, and raised stakes for anxious candidates. Mitigate the second by sending the brief and the format twenty-four hours in advance. Surprise adds nothing but noise, and the candidates it disadvantages are not the ones you are trying to screen out.

Let them use AI, and assess the use

The strongest move and the least used. If the job permits AI — and for most roles you are hiring into now, it does — the assessment should too. Then you are measuring the thing you are actually hiring for.

Give the task, permit the tools, and ask for the working: what they prompted, what came back, what they changed and why. A candidate who accepts a plausible-sounding output uncritically is showing you something important. So is one who catches that the model invented a policy that does not exist.

The demand side supports this. Resume Genius’s 2026 Hiring Insights report, surveying 1,000 U.S. hiring managers, found 60% want to test, discuss, or see proof of a candidate’s AI abilities rather than take resume claims at face value, and that they respond better to AI skills shown in context — in an interview or a task — than emphasised in application materials.

Cost: you have to define what good AI use looks like in that role before you can score it, and most teams have not done that. Budget an hour with someone who does the job well. It is the highest-value hour in this article.

Structured assessment tooling is the category that supports running these consistently across candidates — Humavera’s assessments capability among others — but the design work above is yours regardless of what you run it on, and no product does it for you. One note in passing: scored assessments used in hiring carry regulatory obligations in some jurisdictions, particularly around automated scoring and record-keeping. That is its own subject [pending: T-03].

The follow-up conversation is your verification layer

If you take one thing from this: ten minutes of “walk me through why you did it that way” verifies more than any detector, any proctoring subscription, and any AI-detection score.

It works because explanation is much harder to fake than production. Someone who did the work can tell you what they tried first and abandoned. Someone who did not will describe the finished answer in the passive voice, stay at the level of generality the artefact already established, and get vaguer under specifics rather than sharper.

What to ask: what did you try first that did not work; what would change your answer; what did you leave out and why; what would you do differently with another hour. Thin answers restate the output. Real ones contain a discarded path.

One rule, and it is not optional: the same questions for everyone, in the same order, scored on the same rubric. The moment you probe harder on candidates you are unsure about, you have built a bias machine and given it an audit trail. This is authentic skills verification only if it is applied identically.

What it costs, and what to cut

For an HR function of three to eight people, realistically, per role: an hour defining good AI use, an hour writing two items, two hours training interviewers on the rubric, then about forty minutes of assessment and conversation per shortlisted candidate.

Against fifteen candidates that is roughly fourteen hours. You do not have fourteen spare hours, so here is what to cut. Drop the take-home entirely — it is the item type that broke hardest and it costs candidates more than everything else combined. Drop the second screening call, which is where the follow-up conversation now lives. Cut your shortlist from fifteen to eight, which the earlier stages should be doing anyway. That gets you to about eight hours, most of it one-time setup you reuse on every subsequent hire for that role.

Do not try all four moves at once. Take one existing assessment item, ask whether a model answers it well, and if it does, rewrite that single item to ask for judgement or the candidate’s own context instead. Run it on the next role. That is a Tuesday afternoon, and it will tell you more about your process than a proctoring trial ever will.

FAQ

How do you stop candidates using ChatGPT on an assessment? You don’t, and chasing it wastes the effort. Any at-home assessment can be defeated by a second device, and detection tools produce false positives that cluster on non-native English speakers and neurodivergent candidates. Redesign the item instead: if a model answers it well, it was measuring recall or composition. Ask for judgement, personal context, or live reasoning.

Should you let candidates use AI during a hiring assessment? Yes, if the job allows it — which for most roles it now does. Permitting the tools and asking candidates to show their working turns AI use from a threat into the thing you are measuring. Resume Genius found 60% of 1,000 hiring managers surveyed want to test or see proof of AI abilities rather than accept resume claims.

Are work sample tests unpaid labour? They cross that line when they exceed roughly thirty minutes or produce output you could actually use. A short, clearly scoped exercise with a time cap that generates nothing of commercial value to you is defensible. Anything resembling real deliverable work should be paid at a fair rate. State which one yours is when you send the invitation.

Is proctoring software worth it for a mid-size company? Rarely. It verifies that someone sat alone in a room, not that they can do the job, and its false positives land hardest on candidates with poor connections, shared housing, disabilities or neurodivergence. It is also an arms race you fund by subscription. Redesigning two assessment items costs less and addresses the actual question.

What kind of assessment questions still work in 2026? Questions requiring the candidate’s own context, genuine trade-offs with no clean right answer, and live reasoning you observe rather than infer from a finished document. Knowledge recall, definitional questions and polished written responses have stopped discriminating. If the value sits in the artefact, it broke; if it sits in the process that produced it, it held.

One essay a week. No product spam.

Newsletter delivery is not wired up in this theme — connect your email provider to enable it.