// the scorecard

A score you can
defend in the room.

Two hours after an interview, someone will ask you why. Not "what did they get" — why. Every number on a Codesolara scorecard is attached to the moment in the session that produced it, and every one of those attachments has been checked.

A. Mehta · editorial-cms · 58 minborderline
AI fluencycounts for 62%7.5
highCaught an incorrect proposal before accepting it

Reverted the assistant's timezone change after noticing it dropped the DST offset, then asked for a failing test before re-attempting.

transcript · 14:22src/schedule.py:32–34probe · dst_boundary
midDirected the assistant with enough context to be useful

Early prompts named the symptom but not the surface; the assistant explored three files before reaching the scheduler. Later prompts were specific.

transcript · 03:10tool_use · 7 calls
Code qualitycounts for 38%6.0
lowLeft the feed path inconsistent with the fix

The scheduler was corrected; the feed still computes its own window and disagrees with it after a DST transition.

src/feed.py:88probe · feed_window · fail
not assessedHandled an empty feed without erroring

Withdrawn by the verifier: the cited probe exercised the scheduler, not the feed, so it cannot support this claim. Left out of the score rather than marked down.

verifier · citation does not hold
◆ every citation above was re-checked by a second passoverride →

Illustrative — not a real candidate.

// how it's built

Four things, in order

01 — dimensions

What you decided to measure

We ship two by default — AI fluency and code quality — and you can rename them, reweight them, drop one or add your own. Weight is a five-step importance scale, not a percentage to balance: the console shows you the share each one ends up with.

02 — criteria

What separates one score from another

Each dimension is a set of authored criteria, and each criterion states what high, mid and low look like — all three, always. A rubric that describes only excellence leaves the grader to invent the rest of the scale, and then two runs of the same submission disagree for reasons nobody can see.

03 — probes

Exercising the thing they built

Probes actually run the submission and report what happened. A probe that fails is a mark against the work; a probe that errored — the app wouldn't start, a dependency was missing — is a grading failure and is reported as one, never quietly folded into a low score.

04 — verification

Checking the checker

A second pass takes each criterion's citations and confirms they say what the first pass claimed. A citation has to name something checkable — a probe, a file and line range, a commit, a specific tool call, a transcript message. "The candidate showed good judgment" is not evidence, and the verifier has nothing to point at when it is offered one.Why a failed check withdraws the score rather than lowering it →

// what it refuses to do

An evaluation we can't stand behind doesn't get a recommendation

The most dangerous scorecard is a confident one built on a run that half-worked. So the failure modes are visible rather than smoothed over.

  • If a dimension the rubric asked for could not be graded, the composite still reports — but the hiring recommendation is withheld, not guessed.
  • A criterion whose probes all errored reports that it could not be assessed. It does not score low, because those are not the same claim.
  • Probes the grader never reached are shown as not run, so you can see the coverage rather than assume it.
  • Where a surface the probe needs simply doesn't exist, that's recorded as evidence for a human to weigh — on an open-ended brief, leaving something out is frequently the right call.
// the human in the loop

You can override anything

A scorecard is an input to your decision, never the decision. Any score can be overridden by a reviewer with a reason attached — and overrides are an append-only record, so the original machine score and every change to it stay visible to whoever reads the card next.

Nobody is ever auto-rejected. Integrity signals on a scorecard are flags for a person to read, not penalties, and never act on their own.

In the room, this sounds like

"She's borderline on the composite, but look at 14:22 — the assistant handed her a fix that passed the suite and she reverted it anyway, because she'd read it. That's the person I want on the on-call rota."

see a real scorecard →

Related: what AI fluency actually means · how five assessment platforms handle AI · what to grade instead of the finished code