Rewriting a Resume Without Changing Its Evidence Flips Up to 41% of LLM Screening Decisions
Competence-Preserving Resume Perturbations (arXiv 2609.16517, submitted 15 Sep 2026) builds occupation-grounded candidate profiles at controlled competence levels, renders each into multiple presentations varying in wording, structure, polish, and extraction quality, and uses a deterministic validation gate to exclude any variant that alters the underlying qualification evidence. Across six open instruction-tuned LLM conditions, Llama-3.1-8B with its native chat template had the strongest validity at 0.781 yet reversed 29.6% of matched pairwise decisions; Mistral-7B-v0.3 reached 0.644 validity with a 41.4% flip rate. Native chat formatting improved validity for several chat-tuned models but did not remove the instability, so screening validity and presentation stability are separate properties.
↳ Follow the thread