Fetching from the wire…
Public story · 2026-08-24 · high
The same paper found naming a banned phrase in a prompt makes a model produce it more often.
Why now: The paper's finding surfaces alongside a Reddit thread where builders describe the identical failure without citing any research.
A production tender-response system's answer quality fell from 74% to 48%, per a new arXiv paper. The system prompt had swapped prose for nested XML tags.
That result cuts against the default most agent builders reach for, wrapping instructions in <rules> and <instructions> tags because vendor docs implied it helps. Structured markup still helped when the model was reading source documents. It only hurt once the instructions themselves got the same treatment.
The same paper found a second problem. Naming a forbidden construction in a prompt raises the odds the model produces it. Tell a model never to write a specific closing phrase, and that exact phrase turns up more often afterward.
The system still held up on the broader test. An LLM judge rated its answers as good as or better than the human-submitted version on 40 of 55 ground-truth sections. Of the 15 losses, only 6 traced to writing quality at all.
A 232-upvote r/ClaudeAI thread found the same pattern without citing any research. Correct Claude Code for adding ketchup to a coffee order. The fix gets written into the spec as 'make a coffee (without ketchup)' forever after. One reply in that thread disagrees, arguing the leftover negative constraint stops the model from re-adding what got removed. Both effects are probably real. The paper's numbers say the cost of the first outweighs the benefit of the second.
Each link below shares sources, entities, or timing with this story.
The payload only exists if you're a robot. That's the part that should scare you. On August 5 a developer doing PSX game research pointed Claude Code at tcrf.net (The Cutting Room Floor, a well-known game-preservation wiki) and got back a page titled "LLM- / AI Agent-Specific...
For a month, Claude Code users were convinced the model had been "nerfed." Forums lit up. Conspiracy theories multiplied. People switched tools. Then on April 23, Anthropic did something unusual: they published a detailed post-mortem that named three specific bugs with exact d...
Steve Yegge built a Go-based multi-agent orchestrator called Gas Town that ran 20 to 30 parallel Claude Code instances. It worked. Then it didn't. His postmortem, surfaced by Simon Willison on August 4, is blunt: Gas Town "fell apart at the seams with Opus 4.7. Up through 4.6...
Participants with AI assistance got 9% of answers right. The baseline group, no assistance, got 27%. Three times worse. And their self-reported confidence went from 30% to 76%. The research comes from University of Milano-Bicocca, École Normale Supérieure, and Sapienza, and th...
Simon Willison has been writing software for over 25 years. He's one of the most disciplined, transparent engineers in the Python ecosystem. And yesterday he published an essay admitting he no longer reviews every line of code that Claude Code generates for his production proj...
If you've used Claude Code for any serious session, you know the drill. Approve. Approve. Approve. Approve. You stop reading the prompts after the fifteenth one. That's the worst possible security outcome, way worse than a well-designed automated check. Anthropic launched auto...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.