Fetching from the wire…
Skills2026-08-03 · source-backed
Add one line to your agent instructions: before claiming a fix, demonstrate the validation command fails on the original buggy state. The BSG-VA paper measured 46% of passing checks as carrying zero bug-discriminating information, and roughly a third of the 7.8-point improvement from full bug-contrast feedback came from the reminder alone. Cheapest reliability win on this list.
Each link below shares sources, entities, or timing with this story.
An r/LocalLLaMA post (215 upvotes) ran identical inputs through Qwen 35B-A3B and Gemma 26B-A4B and found the tokenizers diverge almost entirely on code: 2.6x apart on HTML/JS, but 1,025 vs 1,039 tokens on a 55-line instruction document. Near-identical on prose. That's a near-3...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
CAFE (arXiv 2608.24794) makes corrective feedback an in-trajectory intervention the agent chooses to request, using one shared-parameter model alternating between search-agent and critic roles. Online RL shapes request returns from a prompt-level call-versus-skip success gap;...
Thinkingbox is an MCP-compatible sandbox with isolated sessions, full execution traces, and outcome evaluation against terminal backend state, carrying 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank internal IT and consulting support (arXi...
First large empirical study of static prompt-configuration files, across 11,427 repos, plus qualitative coding of 65 sampled files into a 65-code codebook (arXiv 2608.10622). Adoption emerged fast from mid-2024 but clusters in small, low-activity, single-maintainer repos. Cont...
The standard criterion for LLM-generated bug reproduction tests, fails on buggy code and passes on the golden fix, turns out to be insufficient. Many F→P tests are "lax": they reproduce the symptom while still admitting plausible-but-wrong patches. Worse, co-generating the tes...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.