Who checks the evidence behind an AI score?

Anand Chapla ·

A second pass has to re-read each citation and confirm it says what the score claimed. When it does not hold, the right move is to withdraw that score — not to mark it down.

That is the whole argument. The rest of this is why marking it down is the wrong answer, which is less obvious than it first sounds, and what has to be true for withdrawal to work.

A citation is not automatically evidence

A score that points at the moment behind it is better than a score that points at nothing. That much is settled, and it is where most of the thinking in this area currently stops.

But look at what a citation actually is. The thing writing it is a language model, and a language model is very good at producing text that has the shape of a reference. src/schedule.ts, lines 118 to 134. Transcript at 00:41:12. Three failing tests.

Those are sentences. Each one is either true, or it is a sentence that looks true. Nothing about the format tells you which.

This is not a reason to distrust the approach. It is a reason to finish it. A grader that cites its evidence has made a claim that can be checked, which is a large improvement over a grader that offers a paragraph of praise. Having made checking possible, you then have to actually check.

The same problem exists with people, and has for much longer. An interviewer who writes “handled the ambiguity well” has cited nothing, so there was never anything to verify. The difference now is that the evidence exists — the prompts, the edits, the test runs, the timestamps — so “I remember it that way” stopped being the best available answer.

What makes a citation checkable

Before anything can be verified, the citation has to name something a second reader can open:

  • a probe and what it printed
  • a file and a line range
  • a commit
  • a specific tool call
  • a message in the transcript, with its timestamp

“The candidate showed good judgment” names nothing, so there is nothing to confirm or contradict.

That constraint does useful work before any verification happens. A grader required to cite something resolvable cannot fall back on the comfortable sentence, because the comfortable sentence has no referent and fails the format. Half the value of verification is the discipline it forces on the thing being verified.

The second pass

The check is a separate run, not the first run reviewing its own work. That detail is load-bearing. A model asked whether it was right about something tends to say yes; it is agreeing with text in its own context rather than going back to the source.

So the second run starts from the evidence, takes each criterion’s answer and the citations attached to it, and returns one of two verdicts: confirmed, or unsupported.

Unsupported does not mean the candidate did badly. It means the reason given for the score could not be found in the session.

Why an unsupported citation must not lower the score

Here is the part that took the longest to get right. There are two obvious things to do with a criterion whose evidence did not check out, and both are wrong.

Score it low. Now an invented citation costs the candidate marks. They are being penalised for a mistake the grader made. They may well have done the thing the criterion asks about — nobody knows, because the only account of it turned out to be unreliable. This is the worse of the two errors, because it does harm to a person who has no way of seeing it happen.

Leave it scored. Now an invented citation earns marks. The number carries a claim that has already failed a check, and the card still looks clean.

A fabricated justification must not move the number in either direction. So the criterion is withdrawn — marked not assessed, with the reason recorded — and the number is built from what is left.

For that to be honest rather than cosmetic, one more thing has to be true: a withdrawn criterion leaves both sides of the arithmetic. It comes out of the total and out of what the total is measured against. If it only came out of the top, “not assessed” would be zero with a politer name, and the candidate would be penalised exactly as before, less visibly.

When the checker itself does not run

There is a third state, and it needs saying out loud because it is the one that is easiest to hide.

Sometimes the verifier does not complete. Two designs are tempting and both are worse than the obvious one.

Treat it as required, and one verifier failure withholds the result of a grading run that was completely fine. That trades a small honesty problem for a large availability problem.

Treat its absence as silence, and an unverified score is indistinguishable from a verified one. Every card looks equally trustworthy, including the ones nobody checked. This is the failure worth designing against hardest, because it is invisible by construction.

So the third state is announced. The card carries a note saying that the verifier did not complete and no citation behind this score was checked, along with why.

An unchecked score is not a scandal. An unchecked score that looks checked is.

What this costs

Two honest costs, because a claim about rigour that hides its price is not a claim about rigour.

It is a second run, so it is slower and it costs more per candidate. That is straightforwardly fine for a decision about whether to hire someone, and straightforwardly wrong as a filter on four hundred applications. If what you need is a cheap first cut, this is the wrong tool and it will feel like it.

It will sometimes withdraw a score that was actually correct, because the citation was sloppy rather than false — a line range off by a few lines, a timestamp that drifted. That direction of error is the one to accept. A score quietly removed costs you a little information. A score wrongly kept costs a candidate a job and costs you the ability to explain why.

The question worth asking

If you are choosing how to run this — with a vendor, or with your own process — the useful question is not whether the score cites its evidence. That is becoming common, and it is the easy half.

Ask what happens when the citation is wrong. Whether anything checks. Whether you are told when nothing did. And whether a score whose evidence did not hold up is removed, or just quietly lowered onto the candidate’s head.


How the codesolara scorecard is built — dimensions you weight, criteria with authored bands, probes that run the submission, and the verification pass described here. The argument this follows from is what to grade once the finished code stops being the signal.

// more