how scoring works
A score with receipts.
Three signals build your number: what you said, the code you wrote, and how you said it. Two are measured, one is judged, and every point traces to a moment in your transcript.
readiness = 0.40·interview + 0.35·coding + 0.25·speaking − staleness
Only surfaces you have actually tried count. The weight redistributes across them, so an untouched tab never pads your number, and going quiet for weeks pulls it back down. A formula, not a mood.
Measured, or judged. Never guessed.
Two of the three signals are computed from what you did. One is an expert judgment against your exact role. Every score names which it is, and where it came from.
Graded against what a strong answer for your exact role must hit, only on what you were asked.
- ownership, specifics, numbers
- dodged questions get named
source role rubric · graded, then validated
Run for real against hidden tests in a sandbox, while you explain your approach out loud.
- correctness is pass or fail
- approach probed mid-solve
source sandbox execution · hidden tests
Pace, fillers, pauses, and talk time, computed straight from your speech timestamps.
- pace · 140 to 160 wpm ideal
- talk time · 75 to 80% you
source speech timestamps · no model
Tell me about a time you led a project.
Larpy: “Lead with what you owned, name the decision, and end on a measurable result.”
The number becomes a tier.
Hard bands, identical for everyone, no curve. Scores within a few points are noise; the tier is the signal. Click the legend and play the whole range.
Larpy: “No notes. You owned every call and put a number on it. They’re drafting the offer.”
delivery 9.2
You can't smooth-talk it.
Our standing test: take a strong answer, strip the substance, keep the exact same confident voice. The score has to drop. If fluent nothing still scores well, our grader failed, not you.
Strong answer
“We cut checkout latency 30%. I isolated the N+1 query, added an index, and measured it against a control.”
88
substance.score = 88
no substance
Substance stripped, same fluency
“We really improved checkout. I worked closely with the team and it went great, honestly a big win for us.”
< 80
must drop ≥ 8, or the test fails
Held to the exam-industry bar.
The hardest test of any grader: does it match a human expert? We validate the same way the SAT and TOEFL are validated, and hold one bar: agree with an expert as often as two experts agree with each other.
grader agreement · 0 is a coin flip, 1 is identical
target qwk ≥ 0.70
study in progressWhat we don't claim.
Tools that claim perfection are the ones you should not trust. The honest edges:
One interview is noisy, for humans too. The readiness number across sessions is the real signal, not any single card.
Your accent is not a metric. Delivery scores pace and structure, never how you sound. A rough mic will not tank you either.
"Ahead of X% for this role" is a calibrated estimate. A model's read of where you sit, meant as a guide, not a live database of every candidate.
Code is execution-verified when the sandbox supports it. When it cannot run, we fall back to expert review, and we say so instead of dressing it up as a real run.
The fastest way to trust it is to feel it.
Run one interview and read the verdict. Every number on your card traces back to something on this page.