Stack layer / Contrast
Harness-Layer Auto-Research Cut Agent Token Traffic 44.7-49.0% at Equal Task Performance
arXiv 2609.20519
Stack layer
When2Think replaces uniform length penalties with instance-level difficulty control, ending the efficiency tax
arXiv / HuggingFace Daily Papers
Contrast
Showing the teacher the answer adds less to self-distillation than the distillation itself does
arXiv / HuggingFace Daily Papers
Policy dependency / Stack layer
Agents hallucinate tools that do not exist, a 675B model does it as often as a 7B one, and merging MCP servers adds new failure surfaces
arXiv 2609.19425
Policy dependency / Stack layer
A Cheap Read-Only Verifier Captures Nearly All the False-Pass Benefit of a Full Planning Stack
arXiv 2609.20474
Stack layer
On-Demand Attention: a lightweight recall head lets a model decide when long-context reads are worth paying for
arXiv
Stack layer
Cairn makes agent memory collective: query the community's reputation for a tool before calling it, submit evidence after
arXiv
Stack layer / Threat pattern
James Mickens argues chain-of-thought monitoring can never be a sound security control
arXiv