Fetching from the wire…
Agents2026-08-08 · source-backed
arXiv 2608.05906 keeps a dual-polarity memory of verified corrections and observed dead ends for Text-to-SQL repair: 66.34% to 69.79% on Spider, 47.35% to 48.44% on BIRD. Then the authors say the quiet part: paired analysis supports the Spider gain but is weak on BIRD, MERIT isn't reliably separated from untyped dynamic retrieval on either benchmark, and Reflexion-style memory reaches 51.24% on BIRD at higher cost. Negative-result honesty in an agent-memory paper is rare enough to be worth reading for the methodology alone.
Each link below shares sources, entities, or timing with this story.
A new system gives simple queries a fast single-pass while triggering exploration and verification for complex joins. Achieves new SOTA on Spider and BIRD text-to-SQL benchmarks. Addresses a real production problem: uniform compute wastes money on easy queries and underperform...
Danish Foundation Models trained it from scratch on 161 datasets. Across 20 benchmarks spanning English, math and code, and Danish, it beats the original HRM-Text 1B, sets a new Danish state of the art, and competes with Qwen 3.5 4B and Gemma 4 E2B. Weights are on Hugging Face...
A July 2 evaluation tested semantic chunking against simple approaches on long structured academic theses using RAGAs, and the sophisticated method didn't win. Performance varied more with document formatting, preprocessing, and query type than with chunking strategy. The auth...
This is the most directly usable paper of the day and it does something rare: it bolts onto an existing agent without touching it. Ledger is a deterministic runtime wrapper that distills an agent's completed interactions into explicit state. What has been observed, what has be...
July 17, Product Hunt's #1 product was Unabyss for Claude: shared memory across all apps and LLMs, 598 votes. July 18, #1 was ZooData: "the data layer for AI agents," 606 votes. Neither is an application. Both are substrate. (Product Hunt) One launch is noise. Two consecutive...
Four new frameworks in one day — each with 6K-12K stars — signal that agent memory has graduated from "feature" to "product category": - memU (12.3K stars) — Memory for always-on proactive agents with filesystem-like hierarchy and proactive intent prediction. PostgreSQL + pgve...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.