Skills
Build your eval golden set from real production failures and calibrate the judge to 85–90% human agreement
Teams shipping reliable agents in 2026 run a measurement loop, not vibes: a golden dataset of 200–500 examples drawn from actual production failures (not synthetic prompts), graders they trust, an LLM judge calibrated to 85–90% agreement with a human-annotated reference set, and a CI gate that blocks regressions before release. The economics justify the judge — human eval runs $5–$50/instance at dozens/day, while a calibrated LLM judge costs fractions of a cent and clears thousands/minute.
↳ Follow the thread