LLM Citations Are Guesses. Source Tracking Makes Them Receipts
C²-Cite raised citation quality by an average of 5.8% and response correctness by 17.4% on the ALCE benchmark....
Procedural Graph Architectures Are Rewiring Self-Evolving Agents
Graph-of-Agents landed at ICML 2025 with a blunt message for anyone shipping LLM agent pipelines: procedural graph architectures...
Open-Source LLM Toolchains Now Run Entirely Local
Somewhere past 130,000 GitHub stars, Ollama stopped being a curiosity and became the front door everyone walks through....
Open-Source Multimodal LLMs Now Take Audio and Video Natively
Open-source multimodal LLMs stopped pretending video is just a pile of images with a transcript stapled on. Qwen3-Omni...
Open Source Long Context LLMs Went 1M. Demand Proof.
Qwen 4 now ships a native 1M-token context window under an Apache 2.0 license. And the open source...
WearableQA Tested 14 Models. Most Fell Below 60%.
WearableQA put 14 big language models through 4,084 health questions built from real wearable data. And most of...
Open-Source Multi-Agent Frameworks: The Ones I’d Actually Ship
Open-source multi-agent frameworks stopped being conference demos, and nobody threw a party about it. Smolagents sits at 29,177...
AI Code Refactoring Tools 2026: The 91% Trap
Claude Code reported 91% refactor accuracy on a 150K-line codebase. And that single figure is both the best...
Prompt Optimization Frameworks: Put a Number on the Prompt
Prompt optimization frameworks have their proof point and it is not subtle: OPRO beat human-designed prompts by up...
GLM-5.3-Flash: The Ox Alpha Reveal, Specs, Pricing, and Open Weights
GLM-5.3-Flash spent the back half of August 2026 answering to a name that wasn’t its own. And if...