Should candidates use AI in a technical interview?

Anand Chapla ·

Yes. Allow it, tell candidates it is allowed, and grade how they use it.

The reason is not generosity. It is that the alternative no longer works. A ban on AI in a technical interview is a rule you cannot enforce, and an unenforceable rule does not produce a fair process — it produces a process where the honest candidate is at a disadvantage.

The ban is already failing

Take-home tests were the first to go. There is no way to know what ran on a candidate’s machine over four hours, and pretending otherwise was always a convention rather than a control.

Live interviews held out longer, on the theory that you can watch someone’s face. That has stopped being true. There are assistants now that listen to a spoken interview and suggest answers while the candidate is talking, designed not to appear on a shared screen. A conversation can no longer tell you whose answer you are hearing.

So a ban leaves you with two groups. The candidates who follow it, and the candidates who do not and are not caught. You are not measuring engineering ability. You are measuring willingness to follow a rule that the process cannot check, which is a real trait, but not the one the job description asked for.

The company that makes the model looked at this and said no

In January 2026 Anthropic published a write-up of how it designs technical evaluations. One of its engineering take-homes explicitly allows candidates to use AI tools, as they would on the job.

What happened next is the interesting part. Claude Opus 4 produced a more optimised solution than almost every human applicant inside the four-hour limit. So they rebuilt the test. Then Claude Opus 4.5 matched the best human score at the two-hour mark — and that best human had themselves made heavy use of Claude with steering. So they rebuilt it again.

At that point some colleagues suggested the obvious fix: ban AI. The engineer who designed the test declined. His reasoning, in his own words, was that beyond the enforcement problems, he wanted people to be able to distinguish themselves in a setting with AI — the setting they would actually be working in.

It is worth being precise here, because this story gets flattened in both directions. Anthropic’s general guidance to candidates is that take-homes are done without Claude unless we indicate otherwise, and that live interviews have no AI assistance on the same condition. This take-home is the “otherwise”. It is an exception, not a company-wide policy. But it is an exception made deliberately, by people who know exactly what the model can do, after watching it beat their applicants twice.

What you are actually measuring now

The objection to allowing AI is that it removes the signal. It does remove one: whether the candidate can produce working code for a self-contained problem. That signal is gone, and it is not coming back. Everyone’s code works now.

What replaces it is more specific and, for most roles, more useful. When a candidate works with an assistant in front of you, you find out:

  • What they gave it. A prompt that names the symptom and not the surface sends the assistant through four files before it finds the right one. A prompt that carries the constraint, the file and the failure mode does not.
  • Whether they took the first answer. This is the single cheapest thing to observe and the most predictive. Plenty of people accept whatever comes back.
  • What they changed. The edits a candidate makes to a generated block are a direct readout of what they understood in it.
  • Whether they can say why. Not whether they can narrate a rationale afterwards — whether the decision and the explanation match.

None of those four are new ideas. A talent acquisition lead listed them almost verbatim in a public post this month, arriving at them from the opposite direction: he was looking for a way to stop trying to catch people. His phrasing was that he is not trying to catch someone using AI, he is trying to see how they work with it.

The part that is genuinely harder

Allowing AI is the easy half. The hard half is that you now have to grade something less tidy than a diff.

A test that produces a pass or fail can be scored by anyone. “Used the assistant well” cannot — not consistently, not across two interviewers, not in a way that survives someone asking why three weeks later. Two people watching the same session will write the same sentence on the scorecard and put different numbers next to it, and neither of them is being careless. There is just nothing underneath the words to check.

That is a real problem and it is worth naming rather than waving through. It is also the reason this is not simply a matter of changing a policy line in your hiring handbook. If you allow AI and keep scoring the finished code, you have made the interview easier and learned less. If you allow AI and grade the work, you need the session itself to be something a colleague can open and read.

What to do on Monday

If you want to try this without rebuilding your process:

  1. Pick one round. Say in writing, before it starts, that AI is allowed and that using it is expected.
  2. Give a task inside an existing codebase rather than a blank file. The choice of which file to touch is most of the judgement, and a blank file removes it.
  3. Write down, before the interview, the four things above as your scoring guide.
  4. Keep the record. Whatever the candidate typed and whatever came back is the evidence, and it is the only part that a hiring manager who was not in the room can check.

Step four is the one most teams skip, and it is the one that decides whether the round produces a defensible decision or a stronger feeling.

Do not detect the assistant. Grade how it is used.

If the detection question is the one you are actually stuck on, it has its own answer — including what a detector that is right 95% of the time does to the honest candidates it flags. And if you want an unaided round anyway, there is a way to do that honestly.


This is how the codesolara workspace is built: the candidate gets a real editor, terminal and file tree with an assistant beside them, and what they asked it, kept and threw away becomes a scorecard that cites the moment behind every score.

// more