Fetching from the wire…
Research2026-07-11 · source-backed
This study investigates dropping speculative decoding's lossless guarantee without any training, quantifying speed-ups against controlled capability drift. Standard spec decoding exactly preserves the sampling distribution. Relaxing it buys latency at a small distributional cost that's sometimes worth it. Paired with "Resample or Reroute?" on budget-aware model routing, the theme is that inference-time knobs, not just model choice, decide your cost/quality point.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
Five tiers collapsed into three AI-native ones (Foundation, Advanced, Prime) effective April 9, with the old Standard/Pro/Pro Plus/Enterprise/Enterprise Plus structure hitting end-of-sale July 1. This is the quiet version of the AI pricing shift. Rather than announce per-outco...
Using op-schema-aware seeded fuzzing against a high-precision fp64 CPU reference on 24 Triton kernels, 15 correct and 9 intentionally buggy, the method caught all 9 buggy variants and passed all 15 controls across five GPU classes (arXiv:2606.20128). Standard kernel benchmarks...
On a verifiable protein-function characterization task routed across tools, model choice swamped federation topology, RL-versus-LLM harness, and prompt expertise: Opus at roughly 92 to 94%, o4-mini at 40 to 50%. Federation across institutional boundaries cost almost nothing (a...
Learned KV eviction has a soft-to-hard mismatch: training uses differentiable gates that attenuate contributions, but inference only saves memory when entries are physically removed (arXiv 2608.23296). A controlled 2x2x2 over attention type, learned gating and positional encod...
A 1.5B distilled model trained with GRPO chooses NoThink, Short, or Long at response start, using a shaped reward that makes each mode pay off at a different length plus hard per-mode token caps. Accuracy held at 0.782 against 0.796 baseline while mean length fell from 4,796 t...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.