Research
MemRiskBench Scores Long-Horizon Agent Memory Risks Deterministically, With No LLM Judge on the Pass/Fail Path
Aggregate scores hide the rare failures that matter: a model at 78% average accuracy may still leak data in 4% of episodes, and benchmark compression preferentially discards those events. MemRiskBench defines a five-category risk taxonomy (stale facts, conflicting updates, cross-user leakage, revoked-memory reuse, constraint decay) operationalized by deterministic trace-grounded checks across a 120-episode scripted benchmark on five locally run quantized models. Its coverage-constrained greedy subset selector retains full ranking (Spearman rho = 0.975), risk coverage 1.0 and high-risk model detection 1.0 at 20% subset size, cutting evaluation compute 5x.
↳ Follow the thread