Fetching from the wire…
Research2026-08-08 · source-backed
DataSpace benchmarks data agents on 410 cross-language tasks over 7,439 artifacts totaling 15.01GB across CSV, JSON, SQLite, Markdown, PDF and video, validated by 11 domain experts. Six frontier multimodal models across five frameworks: best accuracy only 66.34%, and harness choice alone produced a 15.36-point spread independent of model. The consistent failure across all six backbones was multimodal evidence integration and joins, which is precisely what every "chat with your data" product claims to do.
Each link below shares sources, entities, or timing with this story.
Everyone is building summarize-and-evict context management. Compaction, rolling summaries, hierarchical memory, vector-store recall. The entire agent-memory category assumes the answer is to throw away history intelligently. PRO-LONG (arXiv 2607.20064) keeps the complete stru...
claude-mem hit 80,189 stars at v12.6.4, with 1,840 commits and 109 contributors. It hooks five agent lifecycle events to capture observations, compresses them through Claude's agent SDK into SQLite, and reinjects relevant context on new sessions. No manual tagging. One npx com...
Created August 3, ~3,209 stars/day, the highest velocity of anything created in the last two weeks. Rust core with Node.js, Python and WebAssembly bindings, converting Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV and PDF, claiming median sub-5ms conversion and an 81/1...
MemPalace (~53.6K stars) claims best-benchmarked, against mem0 (~57.8K) and claude-mem (~80.8K) (GitHub). Agent memory went from experimental nicety to a competitive subcategory with published benchmarks. Pair this with the local-first angle: Mnemo offers a Rust + SQLite + pet...
Release notes almost never say "we removed a feature because the evals said it doesn't work." Deep Agents v0.7, landed July 29, says it twice. LangChain removed the base system prompt entirely. Trimmed tool descriptions by 43%. Net effect: base input tokens dropped roughly 65%...
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.