ResearchCalibrated Speculative Decoding: Frequency-Guided Selection Achieves 2.33x Peak ThroughputarXiv·medium signalXBlueskyLinkedInCopy linkConventional speculative decoding suffers from frequent false rejections when draft models produce semantically correct but lexically divergent tokens. CSD adds Online Correction Memory (aggregates historical rejection patterns as rescue candidates) and Semantic Consistency Gating (probability-ratio verification instead of exact token matching). Achieves 2.33x peak throughput speedup over baselines.SourceSource pagearXiv↳ Follow the threadStack layer / ContrastReflexion-Style Verbal Memory Sometimes Lowers Success Versus Plain Retry, and Replay Experiments Show WhyarXiv 2609.12404Stack layer / ContrastPre-registered ablation shows removing an LLM verifier stage from an offensive-security agent shifts median reported findings from 0 to 2 per runarXivStack layer / Threat patternMemRiskBench Scores Long-Horizon Agent Memory Risks Deterministically, With No LLM Judge on the Pass/Fail PatharXiv 2609.14976Policy dependency / Stack layerSeqMoE Reaches 80.22% of Full-Load MoE Performance With Only 45% of Experts Resident in Device MemoryarXiv 2609.12978Stack layer / Threat pattern787,562 Function Pairs Show AI Code Is Half the Size of Human Code With Different Defect Classes, Not FewerarXiv 2609.12708Stack layer / Threat patternAgent Frameworks Detect Dangerous Plan Steps and Then Execute Them Anyway; Fewer Than 20 Lines Closes the GaparXiv 2609.15293Stack layer / ContrastSplitting Decode by Attention Type Instead of by Operator Buys 31-56% More Tokens Per JoulearXiv 2609.13134Stack layer / ContrastDistilled Byte Models Match a Token Model's Accuracy on One-Sixth the Data and Cut Logit Storage to a FiftharXiv 2609.12303