Fetching from the wire…
Research2026-06-19 · source-backed
Agent-Orchestrated Adaptive RAG (arXiv:2606.05658) finds agentic enhancements are not universally beneficial. Dynamic query decomposition gained +0.17 MRR on a structured DevOps benchmark but degraded ranking precision on a multi-hop benchmark, and the self-reflective loop only improved citation accuracy at a real latency cost. The takeaway is cost-aware orchestration: gate decomposition and reflection per query type and domain, and measure the latency tax before you turn them on everywhere. "Add more reasoning" is not a strategy, it's a tax you have to justify per query.
Each link below shares sources, entities, or timing with this story.
PCAS: Policy Compiler for Secure Agentic Systems — The first paper to provide measured enforcement results for agent policy compliance (48% to 93%). Uses dependency graphs and Datalog-derived policy language with a reference monitor intercepting all actions. Three case studies...
arXiv 2608.00765 compresses retrieved docs into query-conditioned visual representations, sidestepping the trade-off where hard compression is query-aware but weak and soft compression is strong but needs costly offline encoding. Beats both baselines across varying retrieval d...
The study extracted 130 clean atomic state transitions from 707 real issues in SWE-bench Lite and Verified. Plain RAG scored 0.57-0.59 answer accuracy; an LLM reranker didn't help and added latency, about 18 seconds against 2.1. A (subject, relation, object) supersession memor...
A July 2 evaluation tested semantic chunking against simple approaches on long structured academic theses using RAGAs, and the sophisticated method didn't win. Performance varied more with document formatting, preprocessing, and query type than with chunking strategy. The auth...
Two numbers from this paper should change what you do with your .claude/skills directory this week. First: 65.7% of the benefit from agent skills comes from procedural anchoring. Explicit knowledge injection accounts for 4.5%. Second: expand the skill pool from 5 items to 100,...
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.