Policy dependency / Stack layer
SeqMoE Reaches 80.22% of Full-Load MoE Performance With Only 45% of Experts Resident in Device Memory
arXiv 2609.12978
Stack layer / Contrast
Splitting Decode by Attention Type Instead of by Operator Buys 31-56% More Tokens Per Joule
arXiv 2609.13134
Policy dependency / Contrast
An audit of a production agent pipeline found a reported p99 latency of 2,147,483,647 ms, the signed 32-bit maximum, from a lifecycle clamp
arXiv 2609.12017
Stack layer / Contrast
Reconstructing Facts From Feed-Forward Residual Vectors Answers Single-Fact Questions in Two-Million-Token Contexts
arXiv 2609.12686
Stack layer / Contrast
Odin Runs All 32 Llama-3-8B Transformer Layers Under FHE in 366 Seconds on One H100, 4.51x Faster Than THOR
arXiv 2609.12378
Stack layer / Update thread
Splitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops Fail
arXiv 2609.12839
Policy dependency / Stack layer
RIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persisting
arXiv 2609.12127
Stack layer
RoofLang Lets an Optimizer Agent Search Inference Architectures Instead of Profiling One Stack, Finding 6.2-50.1% Gains on B300
arXiv 2609.12551