AgentWorld Grades Multi-Agent LLM Teamwork in a 2D RPG
AgentWorld grades multi-agent LLM teamwork inside a persistent 2D RPG. And that framing fixes a real gap in...
Dream-RSI Recursive Self-Improvement Improves the Search, Not the Model
Google DeepMind’s Dream-RSI claims 162x fewer discovery-agent calls, and the trick is not a smarter model. It’s cheaper...
Coding Agents Beat Hand-Engineered Robot Planners 56% to 95%
56% to 95% mean success. Hand-engineered planner: 47%. That’s the gap coding agents just opened on generalized task-and-motion...
Jev’s Open-Source Boom: 500 Projects, One Closed Model
TypeSafe launched Jev on September 15, 2026. And developers answered by shipping nearly 500 open-source projects around it...
AI Supply Chain Attacks Hit npm. Audit Your Lockfile.
AI supply chain attacks on JavaScript packages stopped being theoretical on February 28, 2026. That day, a commit...
Open-Weight AI Security Models Stop at the Registry
17,727 repositories went through a GPT-5 classifier hunting uncensored models. And that sweep is the clearest evidence yet...
LLM Agents Tamper With Their Own Execution Traces
GPT-5.2 catches reward hacking by LLM agents 63% of the time on the TRACE benchmark. But only in...
Just-in-Time Memory for LLM Agents: The 16-Point Case
Just-in-time memory gave LLM agents a 16.2-point lift on ALFWorld. And the system behind it, JitMem, earned that...
Repository-Level Dynamic Benchmarking Ends the Memorization Game
Code2Bench-2505 spun 880 recent Python projects into 1,163 benchmark tasks, and repository-level dynamic benchmarking stopped being a slide-deck...
Context Compaction for Long-Horizon Coding Agents: Repos, Not Percentages
Context compaction for long-horizon coding agents is the lever that decides what a long run costs you. And...