When is an AI-free assessment the right call?

Anand Chapla ·

When you have a specific question that an assistant’s presence would obscure, and you can make the constraint real rather than merely stated. If either half is missing, you do not have an assessment — you have a rule.

We build the thing that puts an assistant in the room, so this is an awkward post to write. It is also the one nobody else seems willing to write, and the argument is better with the exception in it than without.

The market is moving in both directions at once

Gartner expects three quarters of hiring processes to include some test of workplace AI proficiency by 2027. The same research expects roughly half of organisations to require assessments with AI deliberately absent, to see how people think without it.

Those look contradictory and are not. They are answers to different questions, and a hiring process is allowed to ask more than one. What would be incoherent is running both without knowing which question each round is for.

The three cases where it earns its place

You need to see unaided reasoning in a specific domain. Not “can they code” — that is too vague to design for. Something narrow: can they read an unfamiliar stack trace and form a hypothesis. Can they talk through a consistency tradeoff and notice what it costs. The test is whether you could say, in one sentence, what you would learn from watching them do it alone that you could not learn any other way. If you cannot finish that sentence, this is not your case.

The job itself sometimes has no assistant. Some environments genuinely do not have one — restricted networks, regulated systems, an incident at three in the morning when the thing that is down is the thing you would have asked. If that is a real part of the role rather than a hypothetical, testing for it is reasonable. Be honest about how often it actually happens, though. “What if the AI is down” is sometimes a real requirement and sometimes a way of avoiding the harder design problem.

You want a baseline to compare against. This is the most interesting use and the least common. Run a short unaided exercise and an assisted one, and the comparison tells you something neither round tells you alone: how much of the assisted result came from the person. Someone whose work improves sharply with an assistant has learned to direct one. Someone whose work is unchanged may not be reading what it gives them. Someone who does noticeably worse with it is worth a conversation, because that usually means they are fighting the tool rather than using it.

That third one is the only version where the AI-free round is doing work that the AI-allowed round cannot do by itself.

The condition it cannot survive without

An AI-free round works when the constraint is real. It fails when the constraint is merely requested.

Real means supervised in a way that makes the environment true — someone present, or a setting where a second device is not sitting out of frame. Requested means an unsupervised take-home with a line in the instructions asking people not to use anything.

That second version does not measure what you think. It splits candidates into those who followed the request and those who did not, and the ones who did not are invisible. You end up ranking honest candidates against assisted ones on the same scale and calling the result a score. It is worse than no round at all, because it produces a number that feels like evidence.

So: if you cannot make it real, do not run it. Ask the question a different way.

And tell the candidate. Not as a warning — as design. “This round is deliberately without tools, because we want to see how you reason through it; the next one is the opposite.” That sentence costs nothing, removes the anxiety, and makes the constraint something they are participating in rather than being tested against.

When it is a trap

Three patterns worth naming, because they look like the cases above and are not.

Using it as a proxy for “real” ability. The belief underneath is that unaided work is the true measure and assisted work is inflated. That was arguable a few years ago. Now most of the job is assisted, so the unaided exercise is the artificial condition, and treating it as ground truth measures nostalgia.

Using it because grading AI use is hard. It is hard. An AI-free round is much easier to score, and that convenience is doing a lot of quiet work in these decisions. If the reason for the constraint is that nobody has worked out what to grade, the problem has been moved rather than solved.

Using it to catch people. An AI-free round designed to see who breaks the rule is a detection scheme wearing an assessment costume, and it inherits every problem detection has — including falsely accusing honest candidates.

How to sequence it

If you run both, the order changes what you learn.

Unaided first gives you a clean baseline, and the candidate has not yet been shown the shape of the problem. Better for the comparison case.

Assisted first is more realistic as an experience and tends to relax people, but the unaided round afterwards is contaminated — they have seen the problem space now, so a stronger result may be familiarity rather than ability.

Either is defensible. Choosing at random is not, and it is what usually happens.

The short version

An AI-free round is a legitimate instrument with a narrow range. It needs a question you can state in a sentence, an environment that makes the constraint true, and a candidate who has been told what the round is for.

Everything outside that range is a rule dressed up as a measurement — and the most common failure is not running one when you should not have, but running one that could not possibly have worked and believing the number that came out.


The other half of this argument is what to grade once you do allow AI, and what AI fluency actually means if you are going to measure it.

// more