Multimodal Flow Modeling Just Cleared 82.8 on 150B Tokens
MF-1 posted 82.8, and multimodal flow modeling just moved out of my “huh, interesting” folder and into my...
The New Long-Context Benchmark: Frontier AI Agents Stall At 68%
The new benchmark capped frontier AI agents at 68% accuracy. And that number should change how you pick...
LongHarness Benchmark: 68% Is The Ceiling
LongHarness Bench landed on arXiv September 29, 2026, and the best score in the whole paper is 68%....
KV-streams Cut Agentic RL Training Up to 5x
KV-streams landed on arXiv September 29, 2026, with a blunt claim: training agents that run long stops choking...
Meta Muse on AI Glasses: The Launch Sellers Keep Misreading
Phone version of Meta Muse already live. Glasses version queued for “the coming months.” Between those two states...
Self-Supervised Confidence Training Teaches Reasoning Models When to Stop
Qwen2.5-Math-7B posted a reported +20.10% accuracy gain on AIME2024, and self-supervised confidence training is what earned it. The...
AgentWorld Grades Multi-Agent LLM Teamwork in a 2D RPG
AgentWorld grades multi-agent LLM teamwork inside a persistent 2D RPG. And that framing fixes a real gap in...
Dream-RSI Recursive Self-Improvement Improves the Search, Not the Model
Google DeepMind’s Dream-RSI claims 162x fewer discovery-agent calls, and the trick is not a smarter model. It’s cheaper...
Coding Agents Beat Hand-Engineered Robot Planners 56% to 95%
56% to 95% mean success. Hand-engineered planner: 47%. That’s the gap coding agents just opened on generalized task-and-motion...
Jev’s Open-Source Boom: 500 Projects, One Closed Model
TypeSafe launched Jev on September 15, 2026. And developers answered by shipping nearly 500 open-source projects around it...