An Audit-Grade Agent Run Capsule Replays Every Model Response but Only Completes 2 of 10 Tool-Using Workloads
NovaFabric records an agent run without modifying agent logic into a portable Run Capsule with a fifteen-entity schema, sealed with a DSSE signature, RFC 3161 timestamp, Merkle log, and redaction attestation, built on OpenTelemetry, DSSE/in-toto, and W3C PROV rather than new cryptography. Mocked replay served every model response from the capsule in 10 of 10 cases, but only 2 of 10 tool-using workloads completed because tool-response substitution is missing, and declared-stream completeness measured 0.652. REST ingest was lossless across a 314-machine ten-region run but capped at 61.6 req/s with p99 of 26.8 s due to per-worker serialization, and the team reports six defects found in their own system and evaluation corpus.
↳ Follow the thread