SourcesDREAM Deep Research Evaluation with Agentic MetricsarXiv·medium signalXBlueskyLinkedInCopy linkAWS evaluation framework for deep research agents measuring multi-turn reasoning, tool use patterns, and research quality.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerPattern: the agent config layer is being treated as an unmanaged dependency graph, and three independent sources said so this weekarXivStack layer / ContrastA knowledge graph for what-to-do: procedure triplets that self-evolve by contrasting failed trajectories with successful onesarXiv 2609.09153Stack layer / Update threadGander Splits a Full-Duplex Omni Agent Into a Realtime Cerebellum and a Reasoning BrainarXiv 2609.08977Stack layer / ContrastA spec-first agent framework taxonomy: persuasion, front-loaded structure, or controls the agent cannot editarXiv 2609.09671Stack layer / ContrastA capability-scoped harness cut prompt-injection execution from 33-47/75 runs to 3/75 without asking the model to spot malicious textarXiv 2609.08371Stack layer / Update threadFive model artifacts in the official Ollama library solve zero of 164 tasks and nothing in the pipeline tests for itarXiv 2609.05881Stack layer / Threat pattern16% of 3,171 public agent-harness setups carry a confirmed security defect, and 3.8% ship a skill that pre-approves your shellarXivStack layer / Follow-up threadSnowflake's HybridDeepResearch shows frontier models hit only ~50-54% Pass@8 when an answer needs both SQL and web searcharXiv