The method

No black box. Just evidence.

The long form of what the Storm page summarises: every stage the engine takes, what we built underneath it, and the questions people ask when they are trying to find the hole.

The guarantees

Three things every score guarantees.

  • Cited

    Every yes or no points at an exact quote. No quote, no finding.

  • Deterministic

    Same evidence, same prompt, same pinned model, same result. A score you cannot reproduce is an opinion with a number on it.

  • You decide

    The engine scores. A person decides. Nothing is auto rejected, auto advanced, or hidden from a reviewer.

End to end

Six stages, from what good means to a score you can replay.

Every step the engine takes, and what we do at each one to make the result trustworthy.

  1. Stage 01 · Define

    Turn the role into testable questions

    You define what matters: competencies, weights, importance. The engine turns each criterion into a small set of yes or no observation questions a reviewer could ask while watching the work.

    They are the operational definition of good for that role: versioned, reviewed, and bound to it. You can edit, add or veto any question.

    You can veto any question
  2. Stage 02 · Capture

    Capture what they actually did, not what they said they would do

    As the operator works, every artifact is captured: voice transcript, code edits, spreadsheet formulas, slide content, every line of chat. It is bundled into a structured evidence package tagged by source.

    Names and locations are stripped before any model sees it. We score the work, not the person.

    PII redacted before scoring
  3. Stage 03 · Score

    Score every question against the evidence

    For each question we run an observation check: one yes or no question, the full evidence bundle, and a single instruction, which is to answer only if you can quote the exact text that justifies it.

    The engine returns an answer, a citation, and one sentence of reasoning. With no quote, the only legal answer is insufficient evidence.

    One question · one citation · one answer
  4. Stage 04 · Verify

    Verify every citation, and punish hallucinations

    Models can fabricate plausible quotes, so we do not take their word for it. Every citation goes through a deterministic validator, pure code with no AI in the path, which checks that the quote actually appears in the evidence.

    If it cannot be found, the answer is downgraded to insufficient and a citation violation is logged. Those violations are a leading indicator of prompt drift.

    Deterministic validator · no AI in the path
  5. Stage 05 · Aggregate

    Aggregate honestly, and never fake confidence

    We do not average our way to a clean number when the evidence is not there. Each criterion comes back assessed when there is enough evidence, partially assessed when there is some signal and missing pieces, or not assessed when the simulation never surfaced it.

    We say which, plainly, rather than rounding a gap into a score.

    Not assessed is a valid result
  6. Stage 06 · Reproduce

    Make every score reproducible, forever

    Every evaluation writes a structured audit line: which prompt versions generated the questions and scored them, which model snapshot produced the answer, which seed was used, and how many citation violations occurred.

    Six months on, replay the exact evaluation against the exact evidence and get the same score. That is not a marketing claim, it is how the pipeline is built.

    prompt version · model snapshot · seed · violations
The controls

What we built so the model cannot drift.

Five engineering decisions that turn a clever language model into a defensible evaluator. These are the spine of the pipeline, not toggles.

  • Versioned prompts

    The instructions we give the engine are file versioned and stamped onto every score. Changing a prompt means a new version, and a new test gauntlet.

  • Pinned model snapshot

    We do not ride the latest model. We pin to a specific snapshot, recorded with every score, so a vendor update cannot silently change a score you already shipped.

  • Self consistency sampling

    For high stakes criteria we run each check several times in parallel with independent seeds and take a majority vote. A flake in one lane cannot bias the result, and ties are surfaced rather than averaged away.

  • PII redaction

    Names and locations are stripped from the evidence bundle before it ever reaches the evaluator. The model scores the candidate's work, not who they are.

  • Built for ADET compliance

    Engineered to clear NYC Local Law 144, the Illinois AI Video Interview Act and the Colorado AI Act. Questions referencing protected characteristics are rejected at generation, demographic and stylistic signals are blocked at scoring, and a paired candidate bias audit gates every prompt change.

Common questions

What buyers ask us, with honest answers.

How do I know it is not just a black box?

Every score breaks down into the questions that produced it, the literal quote from the work that justified each answer, and the engine's one sentence of reasoning. If you can read the evaluation page, you can audit the score. We never show a number without the evidence behind it.

What happens when the AI gets it wrong?

Two things. Because every score is cited, you can see why it went wrong: it cited the wrong passage, missed an artifact, or stopped at one quote. And overrides are first class. The reviewer corrects the score with a comment that becomes part of the permanent audit trail, and patterns of overrides are how we improve the questions over time.

How do you prevent demographic bias from creeping in?

Three layers. Names and locations are stripped before the evaluator sees the evidence. We run a paired candidate bias audit, the same evidence with different surface signals, and require the score delta to stay within tolerance before any prompt ships. And the questions are tied to operational competencies rather than communication style or self presentation.

Can I reproduce a score from six months ago?

Yes. Every evaluation persists the prompt version, model snapshot, seed and evidence bundle. The same inputs through the same pinned model produce the same outputs. If a vendor releases a new model, your old scores are unaffected, bound to the snapshot they were produced under.

Is the AI making the decision?

No. The engine produces an evidence backed assessment for each criterion and the human makes every call. The platform never auto rejects, auto advances, or hides anyone from a reviewer. The engine's job is to make your job easier, not to replace it.

What if a simulation did not surface enough evidence?

That criterion is marked not assessed, plainly. We do not average missing evidence into the score or pretend to a confidence we do not have, and you can re run with a different simulation if you need more signal.

Want to see this on a real operator?

We will walk you through a live evaluation, every observation, every citation, every override, for a role you are working on right now.

Book a walkthrough Back to Storm

Or reach us directly: contact@yolexlabs.com