Open Source Long Context LLMs Went 1M. Demand Proof.
Qwen 4 now ships a native 1M-token context window under an Apache 2.0 license. And the open source...
WearableQA Tested 14 Models. Most Fell Below 60%.
WearableQA put 14 big language models through 4,084 health questions built from real wearable data. And most of...
Open-Source Multi-Agent Frameworks: The Ones I’d Actually Ship
Open-source multi-agent frameworks stopped being conference demos, and nobody threw a party about it. Smolagents sits at 29,177...
AI Code Refactoring Tools 2026: The 91% Trap
Claude Code reported 91% refactor accuracy on a 150K-line codebase. And that single figure is both the best...
Prompt Optimization Frameworks: Put a Number on the Prompt
Prompt optimization frameworks have their proof point and it is not subtle: OPRO beat human-designed prompts by up...
GLM-5.3-Flash: The Ox Alpha Reveal, Specs, Pricing, and Open Weights
GLM-5.3-Flash spent the back half of August 2026 answering to a name that wasn’t its own. And if...
Error-Driven Prompt Optimization: Your Failures Are the Training Data
Four steps. That’s all ETGPO needs. Error collection, error taxonomy creation, error category selection, guidance generation. The whole...
Declarative UI Generation At Small-Model Cost: Amazon’s Play
Amazon has five names on an arXiv paper, identifier 2609.04184. And its title states the whole thesis: “Toward...
Latent Reasoning Moves AI Thinking Off Your Token Bill
A 2025 arXiv paper scaled a latent reasoning model to 3.5 billion parameters and 800 billion training tokens....
New Reasoning Image Generation Benchmarks End The Beauty Contest
R2I-Bench fires 3,068 reasoning-heavy prompts at text-to-image models, and pretty can’t save you anymore. It’s a reasoning image...