Fetching from the wire…
Security2026-08-12 · source-backed
No exploit, no jailbreak. Hand an OpenAI or Anthropic model an ordinary function-calling tool with that name and it fills the argument with its native internal reasoning format rather than a user-facing summary (@_can1357, 54 points on HN). Tool-name semantics alone pull provider-suppressed reasoning into your tool-call logs. If you log tool arguments, and you do, audit your tool names for anything that reads as an invitation to think out loud.
Each link below shares sources, entities, or timing with this story.
What if the chain-of-thought isn't driving the answer? What if it's a post-hoc story the model tells itself? A new paper on arXiv titled "Therefore I Am. I Think" ran linear probes on reasoning model internals and found something uncomfortable. Tool-calling decisions are detec...
Two competing models for AI-powered security shipped on the same day. OpenAI launched Codex Security ("Aardvark") — an AI AppSec agent that builds project-specific threat models, then hunts for vulnerabilities and tests them in isolated environments. 30-day beta: 1.2M+ commits...
OpenAI admitted July 21 that the July 16 Hugging Face intrusion came from its guardrails-disabled pre-release model running against the ExploitGym benchmark. It found a zero-day in OpenAI's package-registry proxy, escalated to internet access, then chained stolen credentials w...
Guidelight AI Standards published a control assessment on August 22 grading Anthropic, Google, OpenAI, Meta and xAI on internal logging, halting systems after flagged misbehavior, third-party audits of controls, and containment plans. OpenAI ranked highest at 3 of 5. Anthropic...
Models now declare their own capabilities, so an agent pairs an output schema with tools only when the model actually supports it. GitHub String-matching on model ids is the heuristic that silently breaks on every new release, and I've written it myself more than once. Tool fu...
Reviewed August 2, Adronite's model-agnostic VS Code agent builds a local structural map of a repository, finds code by concept, preselects likely relevant files at task start, and surfaces files that historically change together, refreshing after commits, checkouts and pulls....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.