Policy dependency / Stack layer
AWS benchmarks show cheapest-per-token is the wrong metric: a pricier model cost 8x less per correct answer
AWS Machine Learning Blog
Policy dependency / Stack layer
AgentCore Evaluations ships 16 evaluators for the failure mode infrastructure monitoring cannot see
AWS Machine Learning Blog
Policy dependency / Stack layer
AWS shows MCP Apps running on Bedrock AgentCore, putting interactive HTML widgets inside ChatGPT and Claude from one server
AWS Machine Learning Blog
Policy dependency / Stack layer
Pattern: Anthropic's own bar is that Claude-written production code gets reviewed harder than human-written code
Simon Willison
Policy dependency / Stack layer
A replay of 68,266 real Claude Code requests says plain LRU beats the clever KV-cache policies
GitHub
Policy dependency / Stack layer
Willison and Alex Garcia ship Datasette security releases from their first frontier-model code audit, using three models
Datasette Blog
Policy dependency / Stack layer
Salesforce frames its agent stack as an 'Enterprise AI Harness' with a separate AI Control Plane
Salesforce
Policy dependency / Stack layer
Boris Cherny: production code written by Claude should clear a higher bar than human code, and he lists the guardrails Anthropic runs
Simon Willison's Weblog