Fetching from the wire…
Public story · 2026-08-17 · high
Models cited gender or ethnicity in their own stated reasoning in at most 0.03% of cases, so self-report cannot catch the bias.
Why now: The audit's design was prespecified before a single response was scored, which is what makes the reversal hard to write off as prompt luck.
A randomized audit ran seven LLMs through 40,068 doctor-recommendation choices and found race and gender swayed picks by up to $14 a visit, per arXiv.
That money moved for reasons the models almost never admit to. Researchers found gender or ethnicity showed up in the models' own stated reasoning in at most 0.03% of responses.
Reputation still drove the biggest swings. A doctor rated 3.9 to 4.7 stars gained 31.4 points of choice probability over lower-rated peers. Being listed first alone was worth about $11 in perceived value.
The demographic effects ran opposite the direction found in prior human-audit studies of real patients. Female-signaled names gained 2.5 points of choice probability. Hispanic-, South Asian-, and Black-signaled names gained 1.3 to 2.9 points over White-signaled names.
The study built in three personas and nine paraphrases across 3,024 choice sets.
A transparency rule built on asking a model to explain itself catches none of this, because the bias lives in outcomes, not in stated reasoning. Auditing healthcare AI for fairness means testing what a model actually recommends across demographic swaps, not what it says about why.
The audit's design was prespecified before a single response was scored, which is what makes the reversal hard to write off as prompt luck.
Each link below shares sources, entities, or timing with this story.
LangChoiceBench covers 28 projects across seven software areas where Python is a poor default, run against 25 LLMs. Python stays heavily over-selected, recommendation-implementation consistency is low, and smaller open-weight models show stronger bias. Analysis of 9,826 reason...
arXiv 2607.24174 (July 27) generated adversarial log entries from real attack traces and got multiple state-of-the-art LLMs to classify traces containing clear indicators of compromise as benign. The defensive gift: the natural-language explanations emitted alongside the class...
The framing is Gricean: an uncertain cooperative speaker retreats up the specificity hierarchy, trading informativeness for truthfulness. On a T-REx-based benchmark varying entity familiarity and referent specificity, model activations do encode whether a referent falls inside...
The July 30 changelog closed the hosted model-catalog and playground service that let developers prototype against multiple LLMs from GitHub directly. If you prototyped against Models endpoints, this is a migration event, not a skim. The surrounding changelog items (Copilot up...
Test-time training for long-context LLMs is highly sensitive to which spans you train on. Random spans degrade accuracy because most are irrelevant (arXiv:2607.09415). S-TTT has the model first identify relevant evidence passages, then run adaptation only on those, for up to 1...
Models navigate to the correct file for 92%+ of required deletions but cut the exact target line only 52% of the time, and 29% of passing patches wrap dead code in a conditional instead of removing it. Grep the diff for newly added if guards around code the task said to delete...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.