Fetching from the wire…
Skills2026-06-28 · source-backed
Set use_dora=True in PEFT's LoRAConfig with the 2026 starting recipe (r=16, target_modules='all-linear'). DoRA decomposes weights into magnitude and direction and applies LoRA only to direction, yielding +3.7% on LLaMA-7B and +1 to 4.4% on larger models with zero added inference cost (Spheron). Free accuracy.
Each link below shares sources, entities, or timing with this story.
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
LoRA/QLoRA adapts behavior but leaves a same-size, same-cost model. If you want cheaper inference, distillation gives you a permanently smaller one. A distilled Llama 3.1 8B trained on a 70B teacher captures 90 to 95% of quality at ~10% of inference cost. The hybrid move: dist...
PortLLM claimed training-free, data-free transfer of LoRA patches onto updated base models, but only over short horizons and without theoretical grounding. This study runs 10 continual-pretraining steps on Mistral, Gemma, and Qwen and finds portability persists long-run, meani...
The NYT reported on July 17 that the June 2026 proposal is structured as monthly installments with an early-exit clause for either side, and would sit alongside Anthropic's existing $45B three-year SpaceX GPU deal from May. Meta fell about 6% intraday before closing down 2%. T...
arXiv 2609.03900 compares a periodic hierarchy against cumulative replay over a 24-month Wikidata stream, varying evaluation month, replay rank and query formulation. On Qwen2.5-1.5B the hierarchy's 5.0-point advantage over rank-8 replay becomes an 11.6-point deficit against r...
A report on Zuckerberg's internal AI all-hands, including a meeting reportedly interrupted by an employee, surfaced confusion in Meta's direction (Wired). It adds to a run of stories questioning whether the Llama/superintelligence reorg has a coherent plan. For builders depend...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.