How do you assess an engineer who uses AI?
Grade four things: the context they gave the assistant, whether they accepted its first answer, what they changed, and whether they can explain why. Then make every score point at the moment in the session that produced it.
That is the whole answer. The rest of this is what each one looks like when you are actually watching, and the failure mode that makes all four useless if you skip it.
Why the finished code stopped working as a signal
The old interview graded an artefact. The candidate produced code; the code worked or it did not; that was the score.
That worked while producing the artefact was the hard part. It is not any more. A competent assistant will get a self-contained problem to a working state for almost anyone, which means a working solution now separates nobody from nobody. If you allow AI into the room and keep scoring the diff, you have made the interview easier and learned less than you did before.
The signal moved. It is no longer in what was produced. It is in what was refused.
The four
1. What context did they give it?
Watch the first prompt. The difference between candidates shows up before the assistant has said anything.
A weak prompt names the symptom: the tests are failing. The assistant has nowhere to start, so it goes exploring — three or four files before it reaches the one that matters. A strong prompt carries the constraint and the surface: this scheduler drops the DST offset when it normalises to UTC, here is the failing case, do not change the public signature.
This is not prompt-craft as a skill in itself. It is a readout of whether the candidate has understood the problem before reaching for help, which is the same thing you were trying to learn from a whiteboard, observable directly for once.
2. Did they take the first answer?
This is the cheapest thing on the list to observe and the most predictive.
A large number of people accept whatever comes back. It looks reasonable, it is formatted confidently, and reading it properly is slower than running it. The candidates worth hiring do something visible here: they read it, and they pause.
Design for this. A task where the obvious answer is subtly wrong somewhere separates the two groups in about ninety seconds. Accepting everything without reading it is exactly the behaviour that walks into it, and the candidate who catches it has shown you something no self-report ever will.
3. What did they change?
The edits a candidate makes to a generated block are a direct readout of what they actually understood in it.
Deleting a defensive branch that could not fire, renaming a variable that shadowed something, pulling out an abstraction that only had one caller, reverting the whole thing and asking again with a better prompt — each of these is a decision, and each is legible. The candidate who changes nothing has either received something perfect or read none of it, and the transcript tells you which.
4. Can they explain why?
The trap in this one is timing.
If you ask why did you build it this way? after the work is finished, you are grading the story they tell about the work, not the work. People are good at constructing rationales after the fact, and a confident narrator will outscore a better engineer.
The defensible version asks during, not after, and checks the answer against the record. If someone says they rejected the assistant’s first approach because it would have broken on empty input, that claim is either in the session or it is not.
The failure mode that makes all four useless
You can grade all four and still end up with nothing you can defend, because of how the result gets written down.
“Showed good judgement” is a feeling with a number next to it. Two interviewers will write that same sentence about the same candidate and score it 3 and 5. Neither is lying. There is simply nothing underneath the words for the other one to check.
Here is the same judgement with the session behind it:
The assistant rewrote a date function and quietly dropped the daylight-saving offset. The candidate caught it, reverted it, and asked for a failing test before trying again.
src/schedule.ts, lines 118 to 134. Transcript at 00:41:12. Test run, three failing.
(Illustrative — not a real candidate.)
Same sentence on the scorecard. The difference is that now you can open it.
This matters in two rooms. In the hiring meeting, nobody has to take your word for it — they can go and read the moment. And when you reject someone, you can tell them which moment the number came from, rather than “the panel felt”.
So the rule the four things hang off is this: every score names the moment that produced it. A score that cannot point at anything is an opinion with a number attached.
That rule has a second half, and it is the one almost nobody gets to: a citation can be wrong. Something has to re-read it and check — and withdraw the score when the evidence does not hold, rather than quietly lowering it.
What this asks of your process
Two things, and they are both awkward.
You have to keep the record. A ten-minute exercise on paper leaves you with your memory of it. An AI-assisted session leaves every prompt and every edit, which is the first time this kind of judgement has had evidence underneath it at all. If you throw the transcript away at the end of the round, you are back to feelings.
You have to check the citations. A score that points at a timestamp is only better than a feeling if the timestamp says what the score claims. That means a second pass over the evidence, not just the first pass that produced it — otherwise you have made your opinions look rigorous, which is worse than leaving them obviously subjective.
Neither of those is free. Both are cheaper than hiring the wrong person, and considerably cheaper than being unable to explain why you did not hire the right one.
Open the last scorecard you wrote. Behind each number, is there something a colleague could go and read?
How the codesolara scorecard is built — dimensions you weight, criteria with authored bands, probes that exercise the submission, and a verification pass over every citation. The other half of this argument is whether to allow AI at all.