Five model artifacts in the official Ollama library solve zero of 164 tasks and nothing in the pipeline tests for it
Executing 327 quantized code-capable artifacts (305 from the official Ollama library across 15 model lines at every eligible quantization at or under 8 GB, 22 from top HuggingFace community repos) through a calibrated 15-task smoke suite found five silently defective artifacts: four Qwen2.5-Coder-3B conversions and one phi3.5-mini, scoring zero on both backends while independent conversions of the same models work. That is 1.6% of official artifacts and 2 of 29 conversion groups. Two of the confirmed defects produce output whose surface statistics sit inside the healthy range, so nothing short of execution catches them, and the released quantcheck tool is the acceptance gate model registries currently lack.
↳ Follow the thread