Interactive Walkthrough · 2026
Release Gatekeeper
An interactive walkthrough of a real AI release pipeline: five decisions, real evidence, and a verdict scored against what actually shipped.
Visit Release Gatekeeper
Most AI portfolios claim evaluation judgment. This walkthrough lets you test mine, and your own, in about three minutes.
Five decisions, real stakes
You play release gatekeeper on a document-automation feature. At each step you get the evidence that was actually on the table: a demo that beats the offline eval, a production replay that inverts the win, a precision threshold with an automation cost curve, and a refusal budget that someone has to sign. The naive call is always available. So are the consequences.
At the end, your calls are scored against the decisions the shipped system actually made, from the shadow-mode replay gate to the compliance-signed 9 percent refusal budget.
Why it exists
Evaluation judgment is the hardest AI PM skill to demonstrate in an interview and the easiest to claim. Every number anchoring the walkthrough is documented in the Decision Docs series on this site, so the simulation and the paper trail check each other.
Play it at alvn.io/evals, then read the experiment brief that replaced our offline evals.