Vibe Coding
Jev's headline eval numbers measure agreement with frontier models, not correctness
The same 2026-09-19 writeup takes apart TypeSafe's launch claims from September 15. The "193x faster and 444x cheaper than frontier LLMs" figures come from evals that score agreement with GPT-6 and Fable 5.1 rather than ground truth, so they are a similarity metric wearing an accuracy label. "Zero hallucination" reduces to schema compliance, and the calibrated-probability claim ships with no published calibration curve or paper. Worth holding onto as a general reading rule for the wave of small typed-decision models: ask what the eval's reference label actually is.
Source
↳ Follow the thread