AI Coding Agents Stopped Autocompleting. Devin Started It.

    AI coding agents did not slowly stop being autocomplete. They stopped on March 12, 2024, the day Cognition posted Devin to X.

    Quick definition if you are arriving cold: an AI coding agent takes a whole task. A bug report, a feature request. And works it end to end, running commands and editing files instead of finishing your sentences for you. Autocomplete completes a line; an agent is supposed to complete the ticket.

    Cognition introduced Devin as “the first AI software engineer” and reported it resolved 13.86 percent of SWE-bench issues without assistance (Cognition’s announcement).

    The model holding the record before it resolved 1.96 percent unassisted.

    And 4.80 percent with a human lending a hand.

    That gap is the whole shift in one breath. Writing code got cheap. Checking it did not.

    How Devin changed AI coding agents

    Autocomplete guesses the next few tokens and gets out of the way.

    Nice work if you can get it.

    Devin’s pitch was a different category of thing.

    Cognition called it “an autonomous agent that solves engineering tasks through the use of its own shell, code editor.

    And web browser.” A shell means the agent runs the commands instead of asking you to. A browser means it goes and reads the docs instead of waiting for you to paste them into a prompt.

    The announcement squeezed the contrast into a single line: “Instead of just autocompleting tasks, it can write entire programs on its own.”

    The test behind that line, SWE-bench, pulls issues from real open-source repositories on GitHub and asks the system to close them.

    Find the bug. Edit the right files. Pass the project’s own tests. That is a other activity from predicting code.

    Cognition also said Devin had passed practical engineering interviews at leading AI companies and completed real jobs on Upwork.

    Those claims come from the vendor’s launch post. And they should be read with that in mind — I have nothing beyond the post to check them against. The benchmark number is the one I hold onto, because at least something measured it.

    What 13.86 percent admits

    One number matters to me in every agent launch, and it is never the number in the headline.

    Devin resolving 13.86 percent of SWE-bench issues unassisted, against 1.96 percent before, is a real leap.

    It is too a confession: on a benchmark made of real GitHub issues, most problems went unsolved, with nobody in the loop to catch the miss.

    Vendors quote the first half of that sentence.

    You live with the second half daily.

    When an agent writes code you did not write, the time you used to spend writing goes to reviewing instead. And reviewing unfamiliar code is slower than writing familiar code ever was. Reject a plausible-looking wrong fix in thirty seconds and you are fine. Ship it, and you pay later.

    For a small shop, this reframes the purchase entirely.

    You are not buying an engineer. You are buying throughput that arrives with a review bill attached. And the review bill never makes the pricing page.

    Do you know how long it takes you to audit a merge request you did not write? Most people do not, and that is exactly the number that decides whether any of this pays off.

    AI coding agents moved into Slack

    By March 2026, the interface had changed again.

    The @Agenhq profile on X described a full loop triggered straight from chat: message it in your Slack channel and it plans the task, writes the code, fixes CI.

    And opens the merge request — “End-to-end. Autonomously. For anyone on your team” (@Agenhq).

    Ticket queue in, pull request you did not write out.

    OpenAI’s Codex leaned the same direction, on volume. It was “built to allow you to start many sessions at once.

    So you can have multiple agents working in parallel,” and “designed to be used by senior engineers” for “adding features or fixing bugs autonomously” (Dan Shipper’s thread).

    Notice who it is for.

    Senior engineers. The exact people equipped to audit what comes back.

    Parallel agents break arithmetic that single-agent demos hide. Four agents at once means four merge requests, which means your review queue is now the constraint on the entire system. Every vendor shipped a code-writing agent.

    I have not seen the review agent that catches up.

    The best open-source AI coding agents in 2026

    In August 2026, an X post rounded up ten coding agents and said, flat out, “All 10 are open source” (the list).

    The names show how wide the surface has gotten:

    – OpenHands. Autonomous software development
    – SWE-agent. Built to solve real GitHub issues automatically
    – Cline. Works in the IDE and the command line, with Model Context Protocol support
    – Goose. Coding, automation, and tool execution
    – Crush. A terminal agent with multi-model and MCP support
    – Aider. An AI pair programmer living in your terminal and Git workflow
    – OpenCode — model-agnostic
    – Qwen Code. A CLI agent built around Qwen models
    – GPT Engineer. Building and iterating on software

    That is nine. I cannot give you the tenth, as the post did not name it either.

    And I am not going to invent an entry to make the count tidy.

    The spread is the point.

    When the field is this crowded and this open, no vendor’s brand is the decision anymore. Your workflow is. If you live in the terminal, start with Aider or OpenCode; if you live in the IDE, start with Cline; if your team lives in Slack, test the Agen pattern. The tools cost nothing. So the real price is your judgment about which tasks deserve autonomy and which do not.

    Common questions about AI coding agents

    What is an AI coding agent?

    Software that takes a full engineering task.

    Not a half-typed line.

    And works it to a finished change, using its own shell, editor.

    And browser rather than waiting on your keystrokes.

    Are AI coding agents free?

    The open-source ones above carry no license fee; you pay for model usage and, less obviously, for review time. That second line item is the one that surprises people.

    What is SWE-bench?

    A benchmark assembled from real issues in open-source GitHub projects, where a system must find the bug, patch the right files.

    And pass the project’s own tests. Devin’s 13.86 percent unassisted resolution came from it.

    The shift already happened; the numbers above are the receipts. What has not happened, for most operators, is building the oversight muscle to match. One concrete move this week: pick a single agent off that list, aim it at a real issue in a repository you know cold. And review the diff as if a stranger opened it. Time the review honestly. That number, not the token bill, is your true per-task cost.

    I track this shift weekly from the builder’s side, failures included.

    Subscribe if you want the next installment, since the agents are not waiting for your review process to catch up.

    Sources

    – Cognition’s Devin announcement
    – @Agenhq on X
    – Dan Shipper’s thread on OpenAI Codex
    – The ten-agent list

    Leave a Reply

    Your email address will not be published. Required fields are marked *