Sources
Splash serves coding agents on one Mac at 2x the next-fastest engine by specializing kernels per model
incoai/splash launched 2026-09-18 (183 stars in a day, Apache 2.0), a local inference engine for Apple silicon that claims 2x the decode speed of the next-fastest engine on Qwen3.8-27B on a 48GB M5 Pro, and 282ms to first token with a 32K context cached. Its argument is anti-generality: kernels, draft model and memory plan are specialized per served model, which is why there is nothing to configure. It speaks OpenAI Chat Completions, OpenAI Responses and Anthropic Messages with streaming, tool calls, JSON Schema output, images and inline PDFs, and ships shortcuts to launch opencode, claude, codex or hermes against it. Requires M3 or newer, macOS 26.4, and 36GB unified memory.
↳ Follow the thread