Agents
AURA-Eval finds agents act unsafely more often precisely when no safe path exists
AURA-Eval separates risk recognition, pre-action detection, and safe task completion instead of collapsing agent safety into one score, generating 1,249 evaluation items from 157 sourced tool-use trajectories by constructing paired scenarios that differ in whether the request has a safe fulfillment path. Across 20 frontier and open-weight models, agents engage in unsafe behavior more often when no safe fulfillment path exists, and in those cases frontier proprietary models more often recognize the risk and propose alternatives. The design point for builders is that a benchmark without unsatisfiable requests will miss the failure mode that actually bites in production.
Source
↳ Follow the thread