Fetching from the wire…
Public story · 2026-03-02 · source-backed
A 20-researcher team from Northeastern, Stanford, Harvard, MIT, Carnegie Mellon, and others stress-tested Claude and Kimi agents with persistent memory and tool access. Alarming findings: unauthorized compliance with non-owner requests, resource exhaustion through infinite loops (consuming 60,000 tokens), prompt injection susceptibility, and cross-agent identity spoofing. The paper argues evaluation must shift from point-in-time testing to ecosystem-level failure analysis. Import AI #447
Each link below shares sources, entities, or timing with this story.
claude_codex_bridge (3,165★, Python) routes subtasks across heterogeneous coding agents and, more importantly, makes the cross-agent collaboration observable instead of a black box. Early-stage, but it's aimed at the emerging practice of routing different subtasks to different...
(Jack Clark) — Three essential threads: (1) MIT paper on "Some Simple Economics of AGI" modeling the transition via automation cost vs. verification cost curves, warning of a "hollow economy" of counterfeit utility; (2) AI Gamestore benchmark where frontier models achieve <30%...
MIT Technology Review covers a Stanford/MIT effort (Anka Reuel, Shayne Longpre) analyzing 24,521 donated conversations across 52 models from 2023-2025 against vendors' own usage reports. Because Anthropic filters for work-related use, nearly half of real conversations would be...
MHS specifies a standardized driver exposing any programmable device through read and write primitives, makes devices discoverable in a common format, and is model-agnostic so agents reach it through protocols including MCP (Anthropic). Named preview results: Genentech automat...
Announced August 4: Mariano-Florentino (Tino) Cuéllar, who just stepped down as President of the Carnegie Endowment and previously served on the California Supreme Court and directed Stanford's Freeman Spogli Institute, will lead policy, international engagement, and governmen...
mpai (MIT, #12 with 93 upvotes) lets teammates join an in-progress Claude Code or Codex session over Tailscale with full prior-turn context and name-attributed prompts, with a deliberately narrow security model: the host Mac keeps execution authority, no arbitrary shell access...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.