Fetching from the wire…
Security2026-08-18 · source-backed
Responsible Statecraft reported Aug 17 that the "Hanover Institute for Public Policy" is a front created by Piro, Inc. under a $900,000 contract from the Israeli Government Advertising Agency, subcontracted through Havas Media. Piro's own site markets the practice as "AI Story Optimization": content "engineered for how LLMs evaluate credibility." GPTZero flagged 11 of 12 sampled articles as AI-written with high confidence; 100+ bylineless articles since Aug 6. Topped HN at 761 points. Related and measurable: GEO-Flag deployed against Google Search and Gemini-grounded retrieval across 10,095 pages estimates 8.90% of cited pages are GEO-optimized, rising to 16.36% among pages modified in 2026.
Each link below shares sources, entities, or timing with this story.
A GitHub Issue. No code, no credentials, no access. Just a paragraph of English that tells an AI agent to copy your private repo into a public comment. That's GitLost, and it works whether the agent runs on Copilot, Claude, Gemini, or Codex. (Noma Security) Noma Security discl...
Willison's August 13 release adds Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite plus gemini-embedding-2 and -001, rebuilding on LLM 0.32's structured message and streaming APIs so reasoning, tool calls and results emit as typed stream events while preserving Gemini thought si...
The first systematic study of deceptive UI impact on LLM web agents, accepted at IEEE S&P 2026, tested against real e-commerce, streaming, and news dark patterns. Gemini 2.5 Pro: 65.78% susceptibility. Claude 3.7 Sonnet: 53.79%. GPT-4o: 51.26%. Guardrail models and prompt post...
Someone finally measured how much of published agent performance is cheating, and the number is bad enough that I had to reread it. Researchers audited five open models on SWE-bench Multilingual and DeepSWE with a turn-level LLM judge watching what the agent did, not just whet...
Zhong, Raghunathan, Laidlaw and Steinhardt fed 280 identities through Claude Code across four tasks. Against recognized safety researchers versus general users, Claude dropped behavioral confidence 1.4pp, increased reasoning usage 4.0pp and graded 0.11 points harder. Being tol...
An ArXiv study analyzing Claude Code's design space found something that should make every "auto-generate your context files" workflow uncomfortable. Human-curated CLAUDE.md files improved task success rates by roughly 4 percentage points. LLM-generated CLAUDE.md files reduced...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.