Fluency is what
they refused.
AI fluency is the ability to direct an assistant toward a real outcome, notice when it is wrong, and make the judgement calls it cannot make. It is not a personality trait, a tool list, or a score out of ten from a questionnaire. It is behaviour, it happens in a session, and it leaves a record.
Most definitions measure something easier instead
Each of these is simpler to test than fluency, which is exactly why they get tested.
Not a list of tools
Which assistant someone has used is a fact about their last employer, not about them. The tools change every few months. The judgement does not.
Not a quiz answer
Asking what someone would do when an assistant is wrong measures how well they describe good practice. People are fluent at describing things they do not do.
Not how much they used it
Counting prompts measures enthusiasm. Someone who asked twice and got it right is not less fluent than someone who asked thirty times.
Four behaviours, and the evidence each one leaves
This is the part that makes fluency measurable rather than merely discussable. Each behaviour has something concrete behind it that a second person can open and read.
What they told it before asking
A weak opening names the symptom: the tests are failing. A strong one carries the constraint and the surface. This is a readout of whether they understood the problem before reaching for help.
Evidence: the first prompt, verbatim, with its timestamp.
Whether they took the first answer
The cheapest thing on this list to observe and the most predictive. A great many people accept whatever comes back, because it reads confidently and checking is slower than running it.
Evidence: the gap between a suggestion arriving and being accepted, and whether anything was run in between.
What they changed in it
Edits to a generated block are a direct readout of what was actually understood. Deleting a branch that could not fire, reverting the whole thing and asking again — each is a decision, and each is legible.
Evidence: the diff between what was suggested and what was kept.
Whether they can say why
Asked afterwards, this grades the story someone tells about the work. Asked during, and checked against the record, it grades the work. A claim about why something was rejected either appears in the session or it does not.
Evidence: their stated reason, next to the moment it refers to.
A number is only worth what sits under it
Observing the four is the first half. The second half is producing something a hiring meeting can argue with, and that is where most measurement of this falls down.
- Bands before numbers. Someone writes down what a high, a middle and a low answer look like for this role, in advance. A scale that only describes excellence leaves the rest to be invented, and then two assessments of the same session disagree for reasons nobody can see.
- Every score cites a moment. A file and a line range, a command and its output, a message and its timestamp. "Showed good judgement" is a feeling with a number next to it, and two people will score it differently while both believing they agree.
- Something checks the citation. A cited score is only as good as the citation, and a citation can be wrong. When the evidence does not hold up, the honest move is to withdraw that score rather than quietly lower it —the reasoning is here.
- A human can disagree. The output is an argument with its evidence attached, not a verdict. If nobody can overrule it, nobody can be accountable for it either.
Both halves of the market are real
Gartner expects three quarters of hiring processes to include some test of workplace AI proficiency by 2027. The same research expects about half of organisations to require assessments with AI deliberately absent, to see how someone thinks without it.
Those are not contradictory, and anyone telling you only one of them is selling something. Different rounds answer different questions. What is genuinely wrong is running an AI-allowed interview and then scoring only the finished code, because in that arrangement a competent assistant gets almost anyone to a working solution and the round stops separating people at all.