Why we gate graduation on evals, not intuition
A harness that feels ready and a harness that's measured ready are different claims. Notes from building the first custom benchmarks: what makes a task worth testing, and why a passing score should be as boring as a build passing CI.