Fetching from the wire…
Research2026-08-13 · source-backed
Microsoft Research tested continuity, separation, order, enclosure, and knots in both static analysis and interactive planning. Static performance was consistently better than interactive, both well below human. The failure modes differ in a way that matters: static errors are perception failures, planning errors appear after the scene is understood correctly, as models lose track of relationships across multiple actions. Image and video generation tools helped little and often altered topology outright.
Each link below shares sources, entities, or timing with this story.
Flint sits between terse-but-bland chart specs and hand-tuned bespoke visuals, designed so agents can author short specs that still produce expressive, non-generic charts. If you're building agents that generate reports or dashboards, this is a middle path worth a look. Generi...
Microsoft Research dropped a paper that should change how every builder thinks about their agent configuration files. SkillOpt (arXiv 2605.23904) treats a Markdown document as an external parameter of a frozen LLM and applies learning rate, batch, and momentum concepts in text...
SigLIP2-so400M vision encoder plus Phi-4-mini-instruct via a lightweight adapter, three-stage training, DAPO-based RL refinement, co-trained classification and grounding heads (Microsoft Research). First on the ReXVQA leaderboard at 94% as of August 2026. The number worth star...
Orchard Env is a Kubernetes environment service supplying reusable isolated components for data collection, RL rollouts, and evaluation without per-domain modification. The differentiating claim is harness-native training: a lightweight proxy records a real harness's own model...
Uber's COO Andrew Macdonald told Business Insider what a lot of engineering leaders are thinking but won't say publicly: "Getting harder to justify money spent on tokenmaxxing." The backstory: Uber's CTO revealed the company burned through its entire 2026 Claude Code budget by...
Microsoft Research showed Qwen3-4B with a "Skeptical-Agent" outperforms 32B models and approaches 235B single-attempt performance. A 50x+ model size compression through inference-time self-refinement. Practical evidence that you can trade model size for inference-time compute...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.