Fetching from the wire…
Public story · 2026-08-24 · high
The test ran 1,802 Oracle scripts through Kiro, Gemini, and Copilot, then swapped specs between agents to see what broke.
Why now: The cross-agent numbers come from a paper posted to arXiv in August 2026, before most teams have audited spec handoffs between agent vendors.
Gemini's SQL syntax validity collapsed to 2.33% when it worked from a migration spec written by a different AI coding agent. Token F1 came out to 0.035, AST similarity to 0.015, in a test that swapped specs between agents to see what survived the handoff.
The numbers come from a paper posted to arXiv on Oracle-to-PostgreSQL migrations. Researchers ran 1,802 Oracle PL/SQL scripts, pulled from a corpus of 1,006 files, through three tools: Amazon Kiro, Gemini, and Copilot. Then they cross-tested, feeding a spec one agent wrote into a different agent's execution step. The worst pairing was Gemini running on a Kiro-authored spec.
Rewriting the foreign spec into Gemini's own format before execution closed most of the gap.
The gap matters for anyone chaining different vendors' coding agents on the same migration job, since this test covered only a few of many possible agent pairings. It's a narrow slice of the broader spec-interchange problem, but the collapse was steep enough to notice.
Each link below shares sources, entities, or timing with this story.
A GitHub Issue. No code, no credentials, no access. Just a paragraph of English that tells an AI agent to copy your private repo into a public comment. That's GitLost, and it works whether the agent runs on Copilot, Claude, Gemini, or Codex. (Noma Security) Noma Security discl...
19,659 stars since February, ~119/day, 1,410 forks. Different angle from the token-compression proxies: rather than compressing what goes to the model, it sandboxes tool output completely while persisting session memory and enforcing routing via MCP plus hooks across Claude Co...
Agentic Storefronts let merchants sell directly within ChatGPT, Copilot, and Gemini. Buyers never leave the AI interface to complete transactions. Commerce is becoming a distribution layer inside AI, not a destination. PYMNTS
The first systematic study of deceptive UI impact on LLM web agents, accepted at IEEE S&P 2026, tested against real e-commerce, streaming, and news dark patterns. Gemini 2.5 Pro: 65.78% susceptibility. Claude 3.7 Sonnet: 53.79%. GPT-4o: 51.26%. Guardrail models and prompt post...
v0.10.0 (~84.8k stars, Apache-2.0) ships no agent of its own and drives whichever CLI you already have, Claude Code, Codex, Cursor, Copilot, OpenClaw, Gemini, Kimi, Qwen, Cline, plus BYOK OpenAI-compatible endpoints, via od mcp install <agent>. It produces single-page HTML pro...
Announced July 30, Gemini 3.1 Flash-Lite and 3.5 Flash join Cohere and Meta options, with Oracle explicitly framing model selection as per-scenario price-performance. The incumbent ERP vendor is conceding the model layer entirely and defending the data and workflow layer. That...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.