Stack layer / Follow-up thread
Benchmark Radar Ships a Daily-Updated Catalog of 1,283 AI Benchmarks With 12,916 Numeric Score Observations
arXiv 2609.11115
Policy dependency / Follow-up thread
MOSAIC Picks a GraphRAG Traversal Policy per Query and Beats the Best Fixed Policy by 9.96 Points
arXiv 2609.11065
Policy dependency / Stack layer
RIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persisting
arXiv 2609.12127
Stack layer
Agent Benchmarks Carry a Double Measurement Confound: Moving Execution Decisions Off the Scaffold Turns a Flat Leaderboard Into a Spectrum
arXiv 2609.09218
Stack layer
Every LLM Judge Grades Stronger Models More Leniently, and a Label-Free Ensemble Tracks an Oracle Within 0.5 Points
arXiv 2609.12002
Stack layer
RoofLang Lets an Optimizer Agent Search Inference Architectures Instead of Profiling One Stack, Finding 6.2-50.1% Gains on B300
arXiv 2609.12551
Stack layer
GraphAHA merges equivalent programs into one node so test-time search statistics are shared, beating the strongest baseline in 18 of 20 cases
arXiv 2609.12757
Stack layer / Threat pattern
Unlearning Methods That Pass TOFU and MUSE Still Leak the Secret on 22-86% of Queries Once the Model Is an Agent
arXiv 2609.12808