Can you detect AI use in a coding interview?

Anand Chapla ·

Not reliably, no. But the more useful answer is that detection is the wrong thing to build, because a detector that worked perfectly would still not tell you what you need to know.

Two separate problems, and the second one survives even if you solve the first.

The detectors are not accurate enough to act on

Start with the practical position. The tools built to help candidates through an interview undetected are funded products with paying users, actively maintained by people whose whole job is staying ahead of detection. Anything you deploy is on the losing side of that race by design: they get to test against your detector, and you do not get to test against their next release.

Then there is the plain physical problem. A second device sitting beside the laptop produces no signal on the machine you are watching, and no amount of screen analysis will find it.

So the realistic ceiling is a system that is right much of the time and wrong some of the time, with no way to tell which case you are looking at. That sounds tolerable until you work out what it does at volume.

The arithmetic almost nobody runs

Suppose a detector is right 95% of the time. That is a generous assumption for this kind of tool, and it still produces a result most teams would not accept if they wrote it down.

Run two hundred candidates through it. Say forty of them used an assistant in a way your policy forbids. The detector finds thirty-eight of them, and misses two. It also flags eight of the hundred and sixty people who did nothing wrong.

(Illustrative arithmetic, not measured figures — the point is the shape, not the numbers.)

So of forty-six flagged candidates, eight are innocent. Roughly one in six accusations is false, and there is no property of an individual flag that tells you which ones. You have built a system whose confident outputs you cannot act on confidently.

Now look at what happens to those eight people. They are rejected, and they are not told why, because explaining the flag means explaining the detector. They cannot appeal something they were never shown. They walk away believing they failed on merit, and you record a result that is simply false — and you do it without ever finding out.

The two errors are not symmetric. Missing two people who used an assistant costs you two interviews that told you less than they should have. Wrongly accusing eight costs eight people a job and costs you eight engineers who will tell their friends.

Even a perfect detector answers the wrong question

Now assume the technical problem away. Suppose you had a flawless oracle that told you, with certainty, whether an assistant was involved.

What have you learned?

You have learned whether someone followed a rule. That is a real fact about a person, and it is not the fact the job description asked about. It tells you nothing about whether they can take an ambiguous problem, direct a tool at it, notice when the answer is wrong, and fix it.

And there is a quieter problem underneath. A rule that only the honest follow does not measure honesty either — it measures honesty plus the belief that the rule is enforced. Change the enforcement and the population changes. You are measuring your own process, not the candidate.

Meanwhile every one of those candidates will use an assistant on their first day, because that is what the job now involves. An interview built to exclude the main tool of the work is measuring a version of the job that no longer exists.

What replaces it

The move is not a better detector. It is to make the question irrelevant.

Allow the assistant, say so plainly in the invitation, and put the whole session on the record — what was asked, what came back, what was kept, what was thrown away. There is nothing to detect when nothing is hidden, and the arms race stops because there is no longer a prize.

That alone is not enough, though, and this is where teams get stuck. Allowing AI and then scoring the finished code makes the round easier and tells you less than before, because a competent assistant will get almost anyone to something that works. You have to move what you score onto the decisions: the context they gave it, whether they questioned the first answer, what they changed, and whether they can explain why. Those four, in detail.

The pleasant side effect is that the anxiety goes out of the room. A candidate who is not worried about being falsely accused does better work, and you see more of how they actually think.

The honest exception

None of this means AI has to be present in every round.

There are genuine reasons to run an exercise without it — checking whether someone can reason through a problem unaided is a real question, and Gartner expects roughly half of organisations to want some form of AI-free assessment. That is a legitimate design choice.

But it is a design choice, not a detection problem. An AI-free round works when it is supervised in a way that makes the constraint real and when the candidate is told plainly that it is the point of that round. It does not work as an unenforceable rule on an unsupervised take-home, which is where most of this started.

The difference matters: one is a stated constraint everybody understands, the other is a trap.

The question to ask instead

“How do we detect AI use?” has no good answer. The question underneath it usually does.

If it is how do we know this work is theirs? — record the session and read it. Authorship is visible in the prompts and the edits in a way it never was in a finished file.

If it is how do we compare candidates fairly? — give them all the same tools and write down what good looks like in advance.

If it is how do we know they can work without it? — say that is what this round is for, and design it honestly.

Each of those is answerable. The detection question is not, and the effort spent on it is effort not spent on the interview.


How the codesolara scorecard is built — criteria with authored bands, probes that run the submission, and a verification pass over every citation. The argument this follows from is whether to allow AI at all.

// more