DeskDeepSeek V4 Benchmark LeaksHHumai Blog·high signalXBlueskyLinkedInCopy link90% HumanEval, 80%+ SWE-bench. Engram memory. Targets consumer hardware. Open-weight. Expected Feb 17.↳ Follow the threadStack layer / ContrastReflexion-Style Verbal Memory Sometimes Lowers Success Versus Plain Retry, and Replay Experiments Show WhyarXiv 2609.12404Stack layer / Threat pattern787,562 Function Pairs Show AI Code Is Half the Size of Human Code With Different Defect Classes, Not FewerarXiv 2609.12708Stack layer / Update threadsmolbenchmark ranks sub-8GB models by tokens per joule and thermals on hardware you already ownyuvrajsingh-mist.github.io (via r/LocalLLaMA)Stack layer / ContrastSplitting Decode by Attention Type Instead of by Operator Buys 31-56% More Tokens Per JoulearXiv 2609.13134Stack layer / Update threadSplitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops FailarXiv 2609.12839Policy dependency / Stack layerA replay of 68,266 real Claude Code requests says plain LRU beats the clever KV-cache policiesGitHubPolicy dependency / Stack layerSGLang Hit With Unauthenticated Pickle RCE via /update_weights_from_tensor, the Fourth Critical Inference-Stack CVE in Four WeeksCERT Coordination CenterPolicy dependency / Stack layerRIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persistingarXiv 2609.12127