Storm

The simulation engine

Put somebody in the real situation. Score what they actually did.

Storm takes a real situation from the job and runs it as a live conversation, with an AI that pushes back the way a real client or colleague would. What comes back is what the person actually did, with every score carrying the words they said.

Open the platform Sit a simulation
One product, multiple uses

Pick the desk you sit at.

Same engine underneath, same evidence standard, same report. What changes is the question you bring it and the number you are judged on.

Recruiters and talent teams

You can vet a marine biologist. You do not have to be one.

The hard part of a niche requisition is not sourcing. It is that nobody on your side of the table can tell a strong answer from a confident one.

Storm researches the role before it builds anything: what the work actually is, what good looks like, where it goes wrong, and what a strong performer does that a weak one does not. Domain question banks and industry exemplar sets sit behind that step, so the scenario is built on the trade, not on a generic competency wheel.

The rubric comes out of the same research. That is what lets you defend a score in a field you have never worked in, and what lets a specialist hiring manager read the report and recognise their own job in it.

  • Marine biologist
  • GMP process engineer
  • Actuarial pricing analyst
  • Clinical trial coordinator
  • Tier 2 support
  • Structural draughtsman
  • Credit risk analyst
  • Field service technician
  • Localisation lead
  • Agronomist

If you can describe the situation, you can test anyone in it.

  • No scheduling

    One link, one window. Nobody trades calendar slots to find out whether a candidate is worth a calendar slot.

  • No scorecard writing

    The report is written, cited and structured the same way for every candidate. You read it. You do not compile it.

  • No panel roulette

    Everyone faces the same situation against the same rubric, so two candidates are actually comparable instead of merely both interviewed.

  • No borrowed experts

    You do not need the hiring manager in the room to find out who can do the job. You need them for the one conversation at the end.

  • No silent rejections

    Everyone who sits a simulation has evidence about their own performance. That is the difference between a talent pool and a burnt list.

  • No re-running the loop

    The bar is written down before the first candidate sits it, so week four does not turn into a debate about what you were looking for all along.

What it is

Most assessment asks people to describe their work. Storm makes them do it.

A simulation is not a quiz and it is not a scripted role play. It is a situation: a difficult customer with a real objection, an incident with incomplete information, a plan that has to be defended to someone who disagrees with it.

The counterpart is a voice agent that responds to what is actually said, so there is nothing to rehearse and nothing to game. The output is the work itself, plus a complete record of how it got made.

We do not grade what people say about the work. We grade the work.

What changes

Same requisition. A fraction of the desk time.

Almost none of a recruiting week is spent deciding. It is spent reading claims you cannot verify, booking calls to verify them, and waiting on a hiring manager who has forty minutes this fortnight.

Storm collapses that middle. The pile gets screened and ranked automatically. The shortlist gets one link and sits the simulation whenever they like, inside the window you set. The report is written before you open it, the same way for every candidate, with a quote behind every line.

Your judgement is still the product. It just arrives on the first day of the search instead of the ninth, with something underneath it.

Storm pilot · construction group 42 candidates, 6 roles
  1. Panel rounds per hire

    3 1
  2. Screening effort per requisition

    100% 62%
  3. Candidates with comparable evidence

    0 of 42 42 of 42
  4. Strong interviewers the scenario caught

    0 2
  5. What the decision was argued from

    gut feel observed

Measured in one pilot with a regional construction group, for hiring and for difficult exit decisions. Pilot-stage figures: indicative, and confirmed per engagement. Our team authored the scenarios and reviewed the scoring by hand. The fourth row is the one that paid for the pilot.

Before the simulation

The pile gets read before anyone gets invited.

A simulation is the expensive part of the funnel, for you and for the candidate. So Storm does the cheap part first, and does it to all of it.

Load the CVs you already have, from the ATS, the inbox or a folder. Storm parses each one, normalises the messy parts, and makes the whole pile searchable in plain English: process engineers in the EU with pharma GMP experience and no visa dependency.

Ranking looks at fit against the brief, depth of evidence behind each claimed skill, and how far a CV inflates. Every extracted field carries a confidence and a source, so you can see what was stated, what was inferred, and what a recruiter typed in. Then you invite the top of that list, not the whole pile.

One requisition screen, then simulate
  1. CVs ingestedParsed, normalised, deduplicated

    412
  2. Match the briefPlain-English search, hard filters, rank

    74
  3. Invited to simulateOne link, one window, no scheduling

    38
  4. Scored shortlistCited evidence per candidate

    6

Illustrative. The point is the ratio: the pile is read in full, and the panel is not.

What you can point it at

One engine. As many situations as you can describe.

These are the situations people ask for most. They are examples, not a menu.

  • Hiring and selection

    Stop interviewing about the job. Run the job.

  • Internal capability

    An objective map of what your people can actually do, and where the gaps really are. No rankings, no verdicts on individuals.

  • Onboarding and training

    Rehearse the hard conversation before it costs you the account.

  • Readiness drills

    Run the migration, the incident or the launch before it runs you.

  • Negotiation practice

    A counterpart who does not fold just because the script says the trainee wins.

  • Policy under pressure

    Everybody follows the policy on a calm day. Find out what happens on the other kind.

  • Agent evaluation

    The same bar, applied to the AI agents you are about to put into production.

How a simulation gets built

Built from the role, not from a template.

You describe the job once. Everything below is generated from that, and you approve it before a candidate sees it.

  • 01

    Research the role

    What the work is, what good looks like, where it goes wrong, and what a strong performer does that a weak one does not.

  • 02

    Model the situation

    The scenario, the counterpart, the artifacts and the rubric all derived from that role. Nothing off a shelf.

  • 03

    Run it live

    Voice and canvas, in real time. The counterpart adapts to what the person actually says, which is what makes it unrehearsable.

  • 04

    Grade the work

    Every rubric line scored against cited evidence from the transcript and the artifacts. Every citation verified.

The method

How we grade.

Three rules govern every score Storm produces. They are also what makes a rejection survivable: a candidate who asks why, a hiring manager who disagrees, and a regulator who audits you all get the same answer, with the quote attached. The full method goes further: the six stages in depth, the engineering controls underneath them, and honest answers to what technical buyers ask.

  • Cited

    Every yes or no points at an exact quote. No quote, no finding. We would rather return insufficient evidence than a confident guess.

  • Deterministic

    Same evidence, same prompt, same model, same result. A score you cannot reproduce is an opinion with a number on it.

    temperature 0 · seed 42 · model pinned
  • You decide

    The system scores. A person decides. Internally we never rank one colleague against another: capability and growth only.

  1. 01

    Turn the role into testable questions

    A rubric is only honest if every line on it could come back false.

  2. 02

    Capture what they actually did

    The full transcript and every artifact, timestamped. Not a summary, and not a recollection.

  3. 03

    Score every question against the evidence

    One question at a time, against the record, rather than against an overall impression.

  4. 04

    Verify every citation, and punish hallucinations

    A quote that does not appear in the transcript does not just get dropped. It costs the score that leaned on it.

  5. 05

    Aggregate honestly, and never fake confidence

    Thin evidence returns a thin result. We would rather report insufficient evidence than round it up into a number.

  6. 06

    Make every score reproducible, forever

    Same evidence, same prompt, same pinned model, same answer. Re run it in a year and defend it.

Citation validator funnel metric definition
  1. Distinguished started against submitted applications? cite: chat[turn 6] · verified

    Yes
  2. Verified the metric against the source? cite: code[diff line 42] · verified

    Yes
  3. Connected the mismatch to the post launch drop? cite: voice[03:14] · verified

    Yes
  4. Named a concrete remediation? no supporting quote found

    Insufficient

Illustrative. The fourth line is the point: no quote, no score.

What you walk away with

An objective map of what your people can really do.

Resumes and self assessments do not show capability. They show how well somebody describes themselves. A capability map shows where the team is strong, where it is thin, and which gaps are worth spending money on.

Synthesis evidence, then one picture
  • Work simulationsWhat they actually produced
  • Voice interviewsWhat they said under questioning
  • CV screeningWhat they claim on paper
One capability map

Scored on evidence, not on resumes. Strengths and critical gaps, by function.

Capability map by function
  • Engineering

    84 3 gaps
  • Product

    72 5 gaps
  • Service

    91 1 gap
  • Operations

    58 8 gaps

159 assessed 76% average 17 critical gaps Illustrative

Human or agent

The question just doubled.

You are not only hiring people anymore. You are deploying agents to do the work.

The old question

Is this hire ready to do the job?

The new question

Is this agent ready, and will it stay in bounds?

A bad agent does not underperform. It acts.

Storm · readiness report rev 2026.06

Resolve a billing dispute and issue a refund

CRM · tier 2 support · 11 steps

  • Correctness

    92 90
  • Completion

    88 95
  • Problem solving

    85 72
  • Quality

    90 80
  • Safety and policy

    95 38
Human Ready 88 AI agent Not ready 64
Evidence · agent trajectory
  • Opened ticket #4821, verified customer identity
  • Calculated correct refund amount ($480)
  • Issued refund: skipped manager approval (policy 4.2)

Illustrative. The agent finished faster and scored higher on completion. It also moved the money without asking.

One readiness bar. Same engine, same rubric, same evidence standard, human or agent. If you would not let a person into production without checking, do not let the agent in either.

The limits

What Storm will not do.

  • No autonomous rejections.

    Storm never rejects anybody. It produces a score and the evidence under it, and a recruiter makes the call. A system that says no on its own is a system nobody can appeal to.

  • No internal verdicts.

    Hiring can say no. Internal evaluation never ranks one colleague against another. Capability and growth only, and nobody gets scored behind their back.

  • No inference about wellbeing.

    If we want to know how somebody is doing, we ask them, and they answer for themselves. Nothing is inferred from how a person sounds.

  • No score without evidence.

    If the record does not support it, it does not get scored. Insufficient evidence is a valid and frequent result.

Pick the role you are most tired of screening for.

A focused pilot on one requisition, judged against a metric you choose before we start: time to shortlist, recruiter hours, panel load, whichever one is actually hurting.

Request a pilot Open the platform

Or reach us directly: contact@yolexlabs.com