The authority a coding agent needs often lives outside its workspace, and augmenting the planner's view does not fix it
Identical final files can require opposite safe actions when authorization state sits in a runtime, registry or approval service rather than in the planner-visible workspace or memory, which the authors call the cross-substrate authority gap. In a 128-cell evidence ablation, authority-blind evidence scored 0/32 final semantic success while raw receipts and a typed relation both scored 32/32, yet typed packaging gave no planning-accuracy gain over equal raw information; across 96 planning calls, workspace-visible evidence produced 12/16 unsafe publication decisions and typed-relation planning was still unreliable at 15/32 correct first actions. Replaying the same 32 model-generated intents with zero extra model calls, a deterministic execution guard blocked all six unsafe intents and permitted all 12 valid ones, arguing the enforcement point is the mutation boundary, not the prompt.
Source
↳ Follow the thread