SourcesTaalas HC1 Custom Silicon 17000 Tokens Per SecondTaalas·high signalXBlueskyLinkedInCopy linkCanadian startup serving Llama 3.1 8B at 17,000 tok/s on model-specific ASIC. 73x H200 speed, 10x less power. TSMC N6, M raised from Fidelity.SourceSource pageTaalas↳ Follow the threadShared entity / Stack layerPoisoning RAG context drops accuracy from 77.9% to 43.5%, and number swaps only bite once they hold the majorityarXiv 2609.09243Stack layer / ContrastApple's A20 Pro widens the iPhone memory bus to 96-bit, putting on-device inference near 115 GB/sr/LocalLLaMA (corroborated by Notebookcheck and 9to5Mac)Stack layer / Threat patternVibe Coding Cut Task Time 27% and Raised Security Vulnerabilities in the Same TrialarXiv 2609.09560Policy dependency / Stack layerCROSS-CATEGORY: Three Independent Agent-Action Gates Shipped in 48 Hours, All Judging the Command Against Stated IntentProduct Hunt, github.com/AGGIB/Stroq and rewarelabs.com (three independent sources; the 72% figure is Reware's own)Stack layer / ContrastEdge0 runs a 35B MoE on Apple Silicon in 2.9 GB of active memory by streaming experts off SSDGitHubStack layer / Update thread100 LLM agents running a simulated town economy for 26 weeks froze the money supply: 0.3% of prices ever changed, and memory deletion made no differencearXivStack layerA solo WebGPU engine doubled 1-bit 27B decode speed in two days to 30 tok/s on a 6 GB laptop GPUr/LocalLLaMA (huggingface.co/mentriaai/Bonsai-27B-mentria)Stack layerKernel fusion took GLM 5.3 Flash from 29 to 40 tok/s on an M3 Ultra by removing idle GPU gaps, not by better matmulsr/LocalLLaMA (github.com/IngeniousIdiocy/ds4)