// ai fluency

Fluency is what
they refused.

AI fluency is the ability to direct an assistant toward a real outcome, notice when it is wrong, and make the judgement calls it cannot make. It is not a personality trait, a tool list, or a score out of ten from a questionnaire. It is behaviour, it happens in a session, and it leaves a record.

// three things it is not

Most definitions measure something easier instead

Each of these is simpler to test than fluency, which is exactly why they get tested.

01 —

Not a list of tools

Which assistant someone has used is a fact about their last employer, not about them. The tools change every few months. The judgement does not.

02 —

Not a quiz answer

Asking what someone would do when an assistant is wrong measures how well they describe good practice. People are fluent at describing things they do not do.

03 —

Not how much they used it

Counting prompts measures enthusiasm. Someone who asked twice and got it right is not less fluent than someone who asked thirty times.

// what to actually look at

Four behaviours, and the evidence each one leaves

This is the part that makes fluency measurable rather than merely discussable. Each behaviour has something concrete behind it that a second person can open and read.

01 — context

What they told it before asking

A weak opening names the symptom: the tests are failing. A strong one carries the constraint and the surface. This is a readout of whether they understood the problem before reaching for help.

Evidence: the first prompt, verbatim, with its timestamp.

02 — scepticism

Whether they took the first answer

The cheapest thing on this list to observe and the most predictive. A great many people accept whatever comes back, because it reads confidently and checking is slower than running it.

Evidence: the gap between a suggestion arriving and being accepted, and whether anything was run in between.

03 — editing

What they changed in it

Edits to a generated block are a direct readout of what was actually understood. Deleting a branch that could not fire, reverting the whole thing and asking again — each is a decision, and each is legible.

Evidence: the diff between what was suggested and what was kept.

04 — account

Whether they can say why

Asked afterwards, this grades the story someone tells about the work. Asked during, and checked against the record, it grades the work. A claim about why something was rejected either appears in the session or it does not.

Evidence: their stated reason, next to the moment it refers to.

// turning behaviour into a score

A number is only worth what sits under it

Observing the four is the first half. The second half is producing something a hiring meeting can argue with, and that is where most measurement of this falls down.

  • Bands before numbers. Someone writes down what a high, a middle and a low answer look like for this role, in advance. A scale that only describes excellence leaves the rest to be invented, and then two assessments of the same session disagree for reasons nobody can see.
  • Every score cites a moment. A file and a line range, a command and its output, a message and its timestamp. "Showed good judgement" is a feeling with a number next to it, and two people will score it differently while both believing they agree.
  • Something checks the citation. A cited score is only as good as the citation, and a citation can be wrong. When the evidence does not hold up, the honest move is to withdraw that score rather than quietly lower it —the reasoning is here.
  • A human can disagree. The output is an argument with its evidence attached, not a verdict. If nobody can overrule it, nobody can be accountable for it either.
// where this is going

Both halves of the market are real

Gartner expects three quarters of hiring processes to include some test of workplace AI proficiency by 2027. The same research expects about half of organisations to require assessments with AI deliberately absent, to see how someone thinks without it.

Those are not contradictory, and anyone telling you only one of them is selling something. Different rounds answer different questions. What is genuinely wrong is running an AI-allowed interview and then scoring only the finished code, because in that arrangement a competent assistant gets almost anyone to a working solution and the round stops separating people at all.

see how the scorecard is built →read the longer argument
// questions

What people ask about this

Is AI fluency the same thing as prompt engineering?
No. Prompt wording is the smallest part of it. What separates people is whether they understood the problem before reaching for help, whether they read what came back, and whether they can say why they kept or rejected it. A well-worded prompt from someone who cannot evaluate the answer is not fluency.
Can you measure AI fluency with a test or a quiz?
A quiz measures what someone says they would do. Fluency is what they actually did when the assistant was confidently wrong. Those come apart, and the second one is the thing you are hiring for, so it has to be observed rather than self-reported.
Do candidates who use the assistant more score higher?
No, and a score that worked that way would be measuring enthusiasm. Volume of use is not a signal in either direction. What matters is what was asked, what was kept, and what was thrown away.
Is this just measuring whether someone is a good developer?
It overlaps, and that is fine — the two were never separable. Directing a tool well requires knowing what correct looks like. The difference is that this is observable directly, in the session, instead of being inferred from whether the final code passed.
What about roles that are not engineering?
The same four behaviours apply wherever someone works with an assistant on real output, which increasingly means QA, data, support and operations. The evidence differs — a document history rather than a diff — but the question of what they asked for and what they changed does not.
Should every interview allow AI?
Not necessarily. There are things worth testing without it, and Gartner expects about half of organisations to require some AI-free assessment. The mistake is running an AI-free process and believing it is measuring how someone will actually work, or running an AI-allowed one and still scoring only the finished code.