Fetching from the wire…
Research2026-08-14 · source-backed
arXiv 2608.13404 analyzed 5,968 IaC-Eval scenario timelines across 15 configurations, tracking 30 CIS Benchmark check IDs for cases where a passing check fails after a repair iteration. Under strict detection, 3.3% of scenarios regress, resource restructuring is the root cause 79.0% of the time, and regressing transitions show 2.6x more code churn (Cohen's d=0.90). 36.6% of standard-mode regressions self-correct within an average of 1.2 iterations. Iteration 3 is the identified optimal stopping point. That's the rarest kind of paper output: an actual number to put in a config.
Each link below shares sources, entities, or timing with this story.
Mohamed Jouini evaluates seven agentic strategies on IaC-Eval v2, 186 AWS/Terraform tasks with Rego v1 intent policies (arXiv 2607.20478). ReAct with MCP or ChromaDB-backed RAG lifts Qwen2.5-Coder 7B from 14.0% to 45.7%; iterative refinement on verifier feedback reaches 62.9%...
Earlier methodologies misclassified standard-library modules as hallucinations (arXiv 2608.22652). Testing seven inference-time defenses across eight models in five families and four languages, RAG reduced the hallucination rate in 18 of 32 model-language configurations, while...
ORCA-bench pairs a live OpenTelemetry-instrumented microservice system (six days of metrics, logs and traces via Prometheus, Jaeger and OpenSearch, plus full source access) with 1,079 RCA tasks varying report specificity and co-occurring faults. Best result across five frontie...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
Most benchmarks test single-episode solving, and memory benchmarks test fact retention. Neither checks procedural reuse, whether an agent can convert a solved session into a reusable search/debug/verify routine (arXiv). Under a Train/Extract/Test protocol with held-out tasks,...
Learned KV eviction has a soft-to-hard mismatch: training uses differentiable gates that attenuate contributions, but inference only saves memory when entries are physically removed (arXiv 2608.23296). A controlled 2x2x2 over attention type, learned gating and positional encod...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.