DispatchUlysses Sequence Parallelism: 12x Longer Sequences on 4x H100Hugging Face Blog·high signalXBlueskyLinkedInCopy linkSnowflake Arctic protocol. SP=4 reduces memory 3.3x, 3.7x throughput at 64K tokens. Integrated into HF Accelerate, Transformers, TRL.SourceSource pageHugging Face Blog↳ Follow the threadShared entity / ContrastSaaStr: 82 of the 100 Fastest-Growing AI-Native Startups Have Technical CEOs, Against 49% in the 2013 Unicorn ClubSaaStrStack layer / Contrastbartowski measured which tensors actually break under quantization, and Q3_K_M got 15% smaller for itHugging Face / bartowski (via r/LocalLLaMA)Stack layer / ContrastOdin Runs All 32 Llama-3-8B Transformer Layers Under FHE in 366 Seconds on One H100, 4.51x Faster Than THORarXiv 2609.12378Stack layer / ContrastSplitting Decode by Attention Type Instead of by Operator Buys 31-56% More Tokens Per JoulearXiv 2609.13134Stack layer / Update threadoh-my-pi v18.1.20 makes every interactive session self-host for remote attach, and documents why encrypted reasoning breaks on proxied gatewaysGitHubPolicy dependency / Stack layerA replay of 68,266 real Claude Code requests says plain LRU beats the clever KV-cache policiesGitHubPolicy dependency / Stack layerRIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persistingarXiv 2609.12127Policy dependency / Stack layerSGLang Hit With Unauthenticated Pickle RCE via /update_weights_from_tensor, the Fourth Critical Inference-Stack CVE in Four WeeksCERT Coordination Center