Fetching from the wire…
Models2026-08-18 · source-backed
OpenMOSS (Xipeng Qiu's group, 32 authors) released MOSS-VL on Aug 15, built on gated cross-attention so it can ingest incoming video frames during generation, with visual tokens kept outside the decoded sequence. 66.0 on OmniMMI Proactive Alerting against a 37.5 baseline, time-to-first-token 5.1x faster than Qwen3-VL-8B. All five checkpoints, the staged training curriculum, and inference code released. At 430 upvotes it's by a wide margin the most-upvoted paper on HuggingFace Daily Papers today.
Each link below shares sources, entities, or timing with this story.
The authors define goal-directed execution as four repeated behaviors: selecting goals, constructing task-relevant state, maintaining fidelity to higher-level objectives, verifying completion against the environment. Post-training Qwen3.5-122B-A10B on 363 long-horizon multi-to...
SecOPD fine-tunes a defense using token-level feedback during on-policy distillation rather than the sequence-level signal prior work used. Against PISmith adaptive injections on Qwen3.6-27B it reports 9.0% attack success where Meta-SecAlign, the previous state of the art, sit...
Three rounds of LoRA self-training on Qwen3-8B against a frozen control turned up seven systematic measurement failures, including a ledger showing capability changes on a model that was never trained, largely an artifact of inference batching. arXiv After a per-problem exact...
Deng et al. built 120 real-case-grounded tasks across 20 business scenes in six financial domains, running four self-evolving scaffolds on a shared Qwen3.7-Max backbone against paired non-evolving controls. Letta posted the highest evolved score (91.65) and fewest compliance i...
Mobius-v0 restructures the transformer into one globally shared Memory (FFN) holding knowledge vectors plus multiple Reasoners (self-attention) that repeatedly query it, using hidden states as cache and carrier (arXiv 2608.14290). Trained from scratch, a 7B Mobius matches a 7B...
Coloring positive words green consistently pushes VLM sentiment predictions positive, to the point that models fail to weigh negative words in the same text (arXiv 2608.14286). The authors trace it to color-induced changes in the vision encoder's latent representations, and sh...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.