A score you can
defend in the room.
Two hours after an interview, someone will ask you why. Not "what did they get" — why. Every number on a Codesolara scorecard is attached to the moment in the session that produced it, and every one of those attachments has been checked.
Reverted the assistant's timezone change after noticing it dropped the DST offset, then asked for a failing test before re-attempting.
Early prompts named the symptom but not the surface; the assistant explored three files before reaching the scheduler. Later prompts were specific.
The scheduler was corrected; the feed still computes its own window and disagrees with it after a DST transition.
Withdrawn by the verifier: the cited probe exercised the scheduler, not the feed, so it cannot support this claim. Left out of the score rather than marked down.
Illustrative — not a real candidate.
Four things, in order
What you decided to measure
We ship two by default — AI fluency and code quality — and you can rename them, reweight them, drop one or add your own. Weight is a five-step importance scale, not a percentage to balance: the console shows you the share each one ends up with.
What separates one score from another
Each dimension is a set of authored criteria, and each criterion states what high, mid and low look like — all three, always. A rubric that describes only excellence leaves the grader to invent the rest of the scale, and then two runs of the same submission disagree for reasons nobody can see.
Exercising the thing they built
Probes actually run the submission and report what happened. A probe that fails is a mark against the work; a probe that errored — the app wouldn't start, a dependency was missing — is a grading failure and is reported as one, never quietly folded into a low score.
Checking the checker
A second pass takes each criterion's citations and confirms they say what the first pass claimed. A citation has to name something checkable — a probe, a file and line range, a commit, a specific tool call, a transcript message. "The candidate showed good judgment" is not evidence, and the verifier has nothing to point at when it is offered one.Why a failed check withdraws the score rather than lowering it →
An evaluation we can't stand behind doesn't get a recommendation
The most dangerous scorecard is a confident one built on a run that half-worked. So the failure modes are visible rather than smoothed over.
- If a dimension the rubric asked for could not be graded, the composite still reports — but the hiring recommendation is withheld, not guessed.
- A criterion whose probes all errored reports that it could not be assessed. It does not score low, because those are not the same claim.
- Probes the grader never reached are shown as not run, so you can see the coverage rather than assume it.
- Where a surface the probe needs simply doesn't exist, that's recorded as evidence for a human to weigh — on an open-ended brief, leaving something out is frequently the right call.
You can override anything
A scorecard is an input to your decision, never the decision. Any score can be overridden by a reviewer with a reason attached — and overrides are an append-only record, so the original machine score and every change to it stay visible to whoever reads the card next.
Nobody is ever auto-rejected. Integrity signals on a scorecard are flags for a person to read, not penalties, and never act on their own.
"She's borderline on the composite, but look at 14:22 — the assistant handed her a fix that passed the suite and she reverted it anyway, because she'd read it. That's the person I want on the on-call rota."
Related: what AI fluency actually means · how five assessment platforms handle AI · what to grade instead of the finished code