Fetching from the wire…
Research2026-08-08 · source-backed
TutorMoments replays real decision points from 462 de-identified one-on-one math tutoring transcripts with 1,500+ teacher-marked moments and thousands of annotations from 27 educators. Across seven models, default assistant behavior was to over-scaffold, making problems easier instead of letting students reason. Performance improved substantially once the prompt explicitly named the scaffolding-versus-rigor trade-off, though variance stayed wide. Transferable well beyond education: if your task requires the model to hold back, you have to say so explicitly.
Each link below shares sources, entities, or timing with this story.
The infrastructure writeup details the platform behind AI2's Earth-observation foundation models, pretrained on roughly 10 TB of multimodal satellite data. The wildfire-risk mapping job compressed an estimated 4,737 hours of serial compute into roughly 30.5 hours at fractions...
Engineering post-mortem on Hugging Face walking through what broke building a production agent rather than a demo. First-party agent-engineering retrospectives from a research lab are rare, since most agent content is vendor marketing wearing a lab coat. Read it next to The Pr...
Skill-α (arXiv 2608.01678) reframes skill generation as RL over sequential edits, decomposing skill construction into individually evaluable changes. The novel signal is a rollback reward that scores each modification by comparing downstream task execution using the original s...
A macOS-only native app built on Rust and GPUI, the GPU-accelerated framework behind Zed, that normalizes sessions, transcripts, tool activity, and checkpoints from existing CLIs into one provider-neutral model (waku.sh). Checkpoints live as hidden git refs so you rewind the w...
A paper from Stanovsky's group (arXiv:2606.16576) tests whether LLM agents can uncover a hidden deterministic finite automaton through membership and equivalence queries. Performance drops sharply as the automaton grows, and trajectory analysis exposes recurring failures in qu...
The paper (38 upvotes, July 17) identifies world-action models whose internal rollout of what should happen is accurate while the emitted action diverges from that rollout. Framed for robotics, but the diagnostic generalizes: any time an agent produces a correct plan and then...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.