Same model weights swing 20 points across OpenRouter providers, and pinning a fallback chain caused an outage
A practitioner writeup measures DeepSeek V4 Flash 0731 at 90.2% down to 75.3% on GPQA Diamond and 81.3% down to 58.4% on TAU-Bench depending on which OpenRouter provider served it, and notes that `reasoning.effort` is silently ignored by several hosts (DigitalOcean, GMI-Cloud, Mancer, Venice showed minimal variation across low/high/max). Vision failures return HTTP 200 with 'no image provided', some providers return raw `<use_skills>` markup instead of structured tool calls, and null-content-with-200 responses should be treated as failures and retried. Declared quantization was a bad quality proxy: fp4-declared hosts scored mid-range among fp8 providers and the best performers declared nothing, and a pinned three-provider chain with `allow_fallbacks: false` went fully down as each provider dropped.
Source
↳ Follow the thread