OSSxaskasdf/ntransformer — Llama 70B on Single RTX 3090 via NVMe-to-GPU BypassGitHub·high signalXBlueskyLinkedInCopy linkC++ CUDA inference engine streams model layers through GPU via PCIe. Optional NVMe-to-GPU bypass. Show HN 380 pts.SourceSource pageGitHub↳ Follow the threadPolicy dependency / Stack layerA Show HN MCP Proxy Puts Policy Where the Model Cannot Reach ItGitHub (Show HN)Policy dependency / Stack layerUnsloth 0.1.805-beta doubles Qwen3.8-Flash-Next and GLM-5.3-Flash inference with MTP on by defaultGitHubStack layer / Contrastllama.cpp traces Qwen3-tts NaN output to an F16 FFN weight overflowing on a 1.5e5 activation peakGitHubStack layer / Update threadpydantic-ai 2.37.0 adds glm-5.3-flash and maps Z.AI's non-standard finish_reason valuesGitHubStack layer / Update threadClaude Code adds a Containment Escape rule to auto mode and blocks plugin symlink path escapesGitHub (anthropics/claude-code)Stack layer / Update threadPattern: an Anthropic default-model flip now propagates through the whole tool chain inside 24 hoursGitHubStack layer / Update threadMCP Rust SDK 3.2.0 coordinates OAuth refreshes through credential stores so concurrent clients stop racing each other's tokensGitHubStack layer / Update threadLangSmith SDK 0.12.0 makes trace sampling deterministic and fails closed when an anonymizer throwsGitHub