Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
Compiled 2026-07-21 · source-backed
vLLM v0.29.0 makes FlashInfer all-reduce default for TP CUDA groups
Source findingvLLM v0.29.0 adds CUDA graph memory profiling for KV cache auto-sizing
Source findingvLLM released agentic-api as stateful Rust server layer for agent applications.
Source findingTRL v1.11.0 replaced custom vLLM server with thin layer over vLLM serve
Source findingNetraRuntime kernel optimizations achieved 78,498 output tokens per second, 2.16x vLLM throughput.
Source findingmarin integrates vLLM in its training infrastructure.
Source findingvToken integrates with vLLM with token-table indirection.
Source findingvLLM shipped day-0 support for Qwen3.8-2.4T-A95B on NVIDIA hardware.
Source findingvllm.cpp demonstrated token-identical output with vLLM at equal or faster speeds
Source findingvLLM shipped Decode Context Parallelism (DCP) for long-context KV cache optimization.
Source findingvLLM achieved 25K tok/s per GB200 GPU using FlashInfer backend.
Source findingvLLM uses FlashInfer GDN backend for prefill kernel optimization.
Source finding