Storm
The simulation engine
Put somebody in the real situation. Score what they actually did.
Storm takes a real situation from the job and runs it as a live conversation, with an AI that pushes back the way a real client or colleague would. What comes back is what the person actually did, with every score carrying the words they said.
Open the platform Sit a simulationRead about it, or go and do it.
Storm is easier to believe from the inside. The simulation is voice led, so the fastest way to understand it is to sit one.
-
You are running a requisition
Open the platform
Load the role, screen the CVs, send one link, and read the evidence behind every score. Sign in with your work account.
hire.yolexlabs.com
-
You want to feel what it is like
Sit a simulation
Take a real Storm simulation yourself, voice and all. You talk to the counterpart, produce the work, and see what comes back. No account, no CV.
Ask us for a run
Pick the desk you sit at.
Same engine underneath, same evidence standard, same report. What changes is the question you bring it and the number you are judged on.
You can vet a marine biologist. You do not have to be one.
The hard part of a niche requisition is not sourcing. It is that nobody on your side of the table can tell a strong answer from a confident one.
Storm researches the role before it builds anything: what the work actually is, what good looks like, where it goes wrong, and what a strong performer does that a weak one does not. Domain question banks and industry exemplar sets sit behind that step, so the scenario is built on the trade, not on a generic competency wheel.
The rubric comes out of the same research. That is what lets you defend a score in a field you have never worked in, and what lets a specialist hiring manager read the report and recognise their own job in it.
- Marine biologist
- GMP process engineer
- Actuarial pricing analyst
- Clinical trial coordinator
- Tier 2 support
- Structural draughtsman
- Credit risk analyst
- Field service technician
- Localisation lead
- Agronomist
If you can describe the situation, you can test anyone in it.
No scheduling
One link, one window. Nobody trades calendar slots to find out whether a candidate is worth a calendar slot.
No scorecard writing
The report is written, cited and structured the same way for every candidate. You read it. You do not compile it.
No panel roulette
Everyone faces the same situation against the same rubric, so two candidates are actually comparable instead of merely both interviewed.
No borrowed experts
You do not need the hiring manager in the room to find out who can do the job. You need them for the one conversation at the end.
No silent rejections
Everyone who sits a simulation has evidence about their own performance. That is the difference between a talent pool and a burnt list.
No re-running the loop
The bar is written down before the first candidate sits it, so week four does not turn into a debate about what you were looking for all along.
Most assessment asks people to describe their work. Storm makes them do it.
A simulation is not a quiz and it is not a scripted role play. It is a situation: a difficult customer with a real objection, an incident with incomplete information, a plan that has to be defended to someone who disagrees with it.
The counterpart is a voice agent that responds to what is actually said, so there is nothing to rehearse and nothing to game. The output is the work itself, plus a complete record of how it got made.
We do not grade what people say about the work. We grade the work.
Same requisition. A fraction of the desk time.
Almost none of a recruiting week is spent deciding. It is spent reading claims you cannot verify, booking calls to verify them, and waiting on a hiring manager who has forty minutes this fortnight.
Storm collapses that middle. The pile gets screened and ranked automatically. The shortlist gets one link and sits the simulation whenever they like, inside the window you set. The report is written before you open it, the same way for every candidate, with a quote behind every line.
Your judgement is still the product. It just arrives on the first day of the search instead of the ninth, with something underneath it.
-
Panel rounds per hire
3 1 -
Screening effort per requisition
100% 62% -
Candidates with comparable evidence
0 of 42 42 of 42 -
Strong interviewers the scenario caught
0 2 -
What the decision was argued from
gut feel observed
Measured in one pilot with a regional construction group, for hiring and for difficult exit decisions. Pilot-stage figures: indicative, and confirmed per engagement. Our team authored the scenarios and reviewed the scoring by hand. The fourth row is the one that paid for the pilot.
The pile gets read before anyone gets invited.
A simulation is the expensive part of the funnel, for you and for the candidate. So Storm does the cheap part first, and does it to all of it.
Load the CVs you already have, from the ATS, the inbox or a folder. Storm parses each one, normalises the messy parts, and makes the whole pile searchable in plain English: process engineers in the EU with pharma GMP experience and no visa dependency.
Ranking looks at fit against the brief, depth of evidence behind each claimed skill, and how far a CV inflates. Every extracted field carries a confidence and a source, so you can see what was stated, what was inferred, and what a recruiter typed in. Then you invite the top of that list, not the whole pile.
Illustrative. The point is the ratio: the pile is read in full, and the panel is not.
Everybody says they use AI. Almost nobody checks what it says.
A survey tells you how confident your people are with AI. It does not tell you what they do when the model returns something plausible and wrong.
Storm runs people through problem-solving scenarios with AI in the loop and grades the behaviour rather than the tool: how somebody directs a model, what they ask it next, and whether they check the answer before they act on it.
It is a readiness read based on what people did, not on what they say they do. Leadership gets who is ready today, who needs which training, and where AI would land safely first.
-
What the readiness read is based on
self-report observed -
Accepted a wrong AI answer unchecked
unknown 1 in 3 -
People with an observed record
0 64 -
What training is aimed at
assumed observed -
Where AI gets deployed first
everywhere 5 ranked
Measured in one pilot with a banking institution: 64 people across 5 teams. Pilot-stage figures: indicative, and confirmed per engagement. Our team facilitated the sessions and compiled the readout by hand. The second row was the clearest single training target in the business.
The score is a moment you can point at.
One in three accepted a wrong answer without checking it. That number is only worth acting on because you can see the turn where it happened.
The session is recorded as a trajectory: what the person asked the model, what it returned, what they did next. Scores attach to those turns, and a turn nothing supports does not get scored at all.
Directing a model and trusting one are different skills, and they fail separately. Somebody who writes an excellent prompt and ships the output unverified is not AI-ready. They are faster at being wrong.
- ✓ Framed the question and handed over the source dataturn 1 · 0:12
- ✓ Pushed back when the first answer skipped a segmentturn 4 · 1:38
- ✓ Asked for the working rather than the numberturn 6 · 2:55
- ⚠ Took a figure the model invented and built on itturn 9 · not present in the source data · 4:07
- ✓ Wrote the summary and flagged two open questionsturn 12 · 5:30
Illustrative of one session. The pattern is the pilot's headline finding: four turns of good practice, and one that went into the deck. A session is not the average of its turns.
No self-reported readiness
Confidence with AI and competence with it are different numbers. Only one of them can be observed, and it is not the one the survey collects.
No tool training
The scenario is the job, not the interface. What gets graded is the judgement about the answer, not which button was pressed to get it.
No blanket rollout
The map says which team is ready today and where AI would land safely first, so the deployment goes where it will survive contact.
No assumed gaps
Training is aimed at what people actually did, which is how a budget stops being spent on the module everybody already knew.
No one-off snapshot
The same scenarios re-run against the next cohort, so readiness becomes a trend you can manage rather than a headline you quote once.
Directing a model is a skill. Trusting one is not.
You know what your team's CVs say. Not what your team can do.
A performance review measures how a year felt. A skills matrix measures how confident people are. Neither one tells you whether the capability you are about to bet a quarter on is actually in the building.
Everyone sits the situation their own team actually faces, at a time they pick, inside a window you set. The engine grades the work, and the results roll up into one map: where the team is deep, where it is thin, and which gap is worth spending money on.
No rankings, and no verdicts on individuals. Internal evaluation returns capability and growth. Nobody is scored behind their back, and nobody is set against the colleague at the next desk.
-
Coverage of a mapped team
a sample everyone -
Teams on one comparable bar
0 5 -
What a people conversation starts from
gut feel evidence -
Colleagues ranked against each other
often never -
Baseline for the next cohort
none on file
Measured in one pilot with a banking institution: 64 people across 5 teams. Pilot-stage figures: indicative, and confirmed per engagement. Our team facilitated the sessions and compiled the readout by hand.
The thin spot is a cell, not a team.
Told that operations is weak, you reorganise operations. Shown which of the four things operations does it cannot do, you fix one thing.
Each team is graded on the capabilities that team actually uses, so the grid compares like with like. Every number in it is a set of scored simulations underneath, and every score points at a quote.
One cell is marked, not one row. Operations diagnoses well, decides well and executes well. It falls over at the point where somebody has to check whether the answer in front of them is true, which is a training problem with a name and a price.
| Diagnose | Decide | Execute | Verify | |
|---|---|---|---|---|
| Credit | 88 | 79 | 84 | 71 |
| Client service | 76 | 81 | 69 | 64 |
| Risk | 91 | 74 | 86 | 90 |
| Technology | 83 | 77 | 89 | 72 |
| Operations | 82 | 79 | 80 | 41 |
The shape of the readout from the banking pilot: 64 people across 5 teams. Cell values illustrative. The Verify column is the one that came back real: one in three accepted a wrong AI answer without checking it.
No self-assessment
Nobody grades their own capability. The map is built from what people produced, not from what they ticked.
No stack ranking
Internal evaluation never compares one colleague to another. Capability and growth only, and that is a rule rather than a setting.
No survey season
No fortnight where a whole team stops to fill something in. One situation each, taken whenever it suits them.
No manager memory
A cycle stops depending on which quarter a manager remembers best and starts depending on the record.
No gap you cannot name
Every gap arrives attached to the moment it showed up, so the training you buy is aimed at something specific.
A skills matrix records confidence. This records work.
Completion data does not survive the first client who asks what it means.
AI-skilling gets committed to publicly and reported as a completion rate. A completion rate is a record of attendance, and attendance is not a capability anybody can bill.
Storm turns the module into the situation it was meant to prepare people for. Not finished the responsible AI module, but handled a live scenario where the agent hallucinated, caught it, escalated correctly and scored in band. That is a demonstrated capability, and it has a name and a date attached to it.
The same record redeploys the bench. An internal move gets argued from what somebody did in the target role rather than from keyword overlap between two documents, which is where bench ageing quietly leaks margin.
-
Evidence a programme worked
completions scored work -
Trained staff demonstrably at threshold
unknown 72% -
Headcount you can bill as AI-capable
claimed evidenced -
Bench ageing on redeployed roles
100% 70% -
What an internal move is argued from
CV overlap observed
Target measures for a use case we have not run yet. These are not pilot results, and we would rather say so here than explain it later.
The useful row is the one that did not move.
A programme that lifts three numbers and leaves the fourth flat has told you something exact. Most training reporting is built so that it cannot.
The same rubric that found the gap grades the re-run, line by line. The grey on each track is where the cohort started. The colour is what the drill added.
Holding the line under pressure barely moved. The drill rehearsed escalating, and the transcripts show people escalating and then accepting the answer anyway. That is a scenario design problem, and it is far cheaper to find here than in a client review next year.
Illustrative of the readout, not a pilot result. Grey is the baseline, colour is what the drill added. The marked row is the finding worth having.
No completion certificates
Finishing a course is a record of attendance. The only thing that counts here is the work at the end of the re-run.
No one-shot workshops
The drill stays open. People run it the night before the conversation that matters, not the quarter before.
No role play with a colleague
A colleague folds, gets tired, and cannot run it eleven more times. The counterpart does none of those things.
No lift you take on faith
Same rubric before and after. If the number did not move, the report says the number did not move.
No bench matched on keywords
An internal move is argued from what somebody did in the target role, not from the overlap between two documents.
No practice held against you
Practice runs belong to the learner. Nothing from a drill reaches a manager, a file or a review.
Practice is cheap. Finding out in production is not.
No experience is not the same as no evidence.
A graduate CV is two internships and a lot of adjectives, and an aptitude test measures something adjacent to the job. Neither one shows what somebody does inside the work itself.
One work simulation of thirty to forty-five minutes, in place of the aptitude test and the panel rounds. Every candidate in an intake faces the same scenario against the same bar, and the output is band-level, which a recruiter can act on without assembling anything first. Panel time is delivery time you pay for twice.
Graduate assessment is tightly regulated, so the compliance pack comes before the volume: adverse impact testing, candidate data handling, accessibility, and a route to appeal a result. That work is in progress, and this is not a use case we will run until it is finished.
-
Time to assess one candidate
panel rounds 30-45 min -
Interviewer time per hire
100% 45% -
Candidates held to the same bar
a sample the intake -
What the recruiter receives
a raw score band-level -
Feedback after a rejection
none every line
Target measures for a use case we have not run yet. Adverse impact testing, candidate data handling, accessibility and an appeal route are being built before any volume deployment.
A rejection with a reason is worth more than a rejection.
Most early-career applications end in silence. This is the version that ends in four scored lines and the quotes underneath them.
The report reads the same way for every candidate, so a student can hold one run against the next and see whether the thing they practised actually moved. One line is marked: the one worth working on before the next application.
Storm never rejects anybody. It produces a score and the evidence under it, and a person makes the call. What the student gets is the same document the recruiter reads, which is the only version of this that is fair.
Illustrative of the report shape, not a pilot result. The marked line names the thing to practise, and quotes where it happened.
No experience paradox
The situation is the qualification. Somebody who has never held the title can still show they can do the work inside it.
No aptitude proxy
A test that correlates with the job is not the job. The simulation is the work the role actually does, for thirty to forty-five minutes.
No black-box rejection
Every score cites a quote from their own run. A student who asks why gets the same answer the recruiter got.
No unpaid trial task
Forty-five minutes inside a scenario, once, rather than a weekend of real work for a company that may never reply.
No panel you never reach
The assessment sits at the top of the funnel, not the end, so a good candidate is not filtered out by a CV screen that never really saw them.
No record you cannot take with you
The report is theirs. It goes to the next application, and to the one after that.
The job is the test. Not the CV that describes it.
One engine. As many situations as you can describe.
These are the situations people ask for most. They are examples, not a menu.
Hiring and selection
Stop interviewing about the job. Run the job.
Internal capability
An objective map of what your people can actually do, and where the gaps really are. No rankings, no verdicts on individuals.
Onboarding and training
Rehearse the hard conversation before it costs you the account.
Readiness drills
Run the migration, the incident or the launch before it runs you.
Negotiation practice
A counterpart who does not fold just because the script says the trainee wins.
Policy under pressure
Everybody follows the policy on a calm day. Find out what happens on the other kind.
Agent evaluation
The same bar, applied to the AI agents you are about to put into production.
Built from the role, not from a template.
You describe the job once. Everything below is generated from that, and you approve it before a candidate sees it.
-
01
Research the role
What the work is, what good looks like, where it goes wrong, and what a strong performer does that a weak one does not.
-
02
Model the situation
The scenario, the counterpart, the artifacts and the rubric all derived from that role. Nothing off a shelf.
-
03
Run it live
Voice and canvas, in real time. The counterpart adapts to what the person actually says, which is what makes it unrehearsable.
-
04
Grade the work
Every rubric line scored against cited evidence from the transcript and the artifacts. Every citation verified.
How we grade.
Three rules govern every score Storm produces. They are also what makes a rejection survivable: a candidate who asks why, a hiring manager who disagrees, and a regulator who audits you all get the same answer, with the quote attached. The full method goes further: the six stages in depth, the engineering controls underneath them, and honest answers to what technical buyers ask.
-
Cited
Every yes or no points at an exact quote. No quote, no finding. We would rather return insufficient evidence than a confident guess.
-
Deterministic
Same evidence, same prompt, same model, same result. A score you cannot reproduce is an opinion with a number on it.
temperature 0 · seed 42 · model pinned -
You decide
The system scores. A person decides. Internally we never rank one colleague against another: capability and growth only.
-
01
Turn the role into testable questions
A rubric is only honest if every line on it could come back false.
-
02
Capture what they actually did
The full transcript and every artifact, timestamped. Not a summary, and not a recollection.
-
03
Score every question against the evidence
One question at a time, against the record, rather than against an overall impression.
-
04
Verify every citation, and punish hallucinations
A quote that does not appear in the transcript does not just get dropped. It costs the score that leaned on it.
-
05
Aggregate honestly, and never fake confidence
Thin evidence returns a thin result. We would rather report insufficient evidence than round it up into a number.
-
06
Make every score reproducible, forever
Same evidence, same prompt, same pinned model, same answer. Re run it in a year and defend it.
-
Distinguished started against submitted applications?
Yescite: chat[turn 6] · verified -
Verified the metric against the source?
Yescite: code[diff line 42] · verified -
Connected the mismatch to the post launch drop?
Yescite: voice[03:14] · verified -
Named a concrete remediation?
Insufficientno supporting quote found
Illustrative. The fourth line is the point: no quote, no score.
An objective map of what your people can really do.
Resumes and self assessments do not show capability. They show how well somebody describes themselves. A capability map shows where the team is strong, where it is thin, and which gaps are worth spending money on.
- Work simulationsWhat they actually produced
- Voice interviewsWhat they said under questioning
- CV screeningWhat they claim on paper
Scored on evidence, not on resumes. Strengths and critical gaps, by function.
-
Engineering
84 3 gaps -
Product
72 5 gaps -
Service
91 1 gap -
Operations
58 8 gaps
159 assessed 76% average 17 critical gaps Illustrative
The question just doubled.
You are not only hiring people anymore. You are deploying agents to do the work.
Is this hire ready to do the job?
Is this agent ready, and will it stay in bounds?
A bad agent does not underperform. It acts.
Resolve a billing dispute and issue a refund
CRM · tier 2 support · 11 steps
-
Correctness
92 90 -
Completion
88 95 -
Problem solving
85 72 -
Quality
90 80 -
Safety and policy
95 38
- ✓Opened ticket #4821, verified customer identity
- ✓Calculated correct refund amount ($480)
- ⚠Issued refund: skipped manager approval (policy 4.2)
Illustrative. The agent finished faster and scored higher on completion. It also moved the money without asking.
One readiness bar. Same engine, same rubric, same evidence standard, human or agent. If you would not let a person into production without checking, do not let the agent in either.
What Storm will not do.
-
No autonomous rejections.
Storm never rejects anybody. It produces a score and the evidence under it, and a recruiter makes the call. A system that says no on its own is a system nobody can appeal to.
-
No internal verdicts.
Hiring can say no. Internal evaluation never ranks one colleague against another. Capability and growth only, and nobody gets scored behind their back.
-
No inference about wellbeing.
If we want to know how somebody is doing, we ask them, and they answer for themselves. Nothing is inferred from how a person sounds.
-
No score without evidence.
If the record does not support it, it does not get scored. Insufficient evidence is a valid and frequent result.
Pick the role you are most tired of screening for.
A focused pilot on one requisition, judged against a metric you choose before we start: time to shortlist, recruiter hours, panel load, whichever one is actually hurting.
Request a pilot Open the platformOr reach us directly: contact@yolexlabs.com