Coding Agents Grew Up
One model reportedly coded on its own for sixteen days against real software projects, and a major platform shipped its first full coding agent. The line from autocomplete to unattended multi-day work has been crossed, and the human's job changed with it.
One model reportedly coded alone for sixteen days on real projects. We have crossed from autocomplete to coworker, and the human's job changed with it.
Two stories this month, same direction. One model reportedly worked on real software projects for sixteen days straight, largely on its own. And a major platform shipped its first full coding agent: not an autocomplete, not a chatbot with a terminal, but an agent meant to plan changes, write them, and validate them.
Sixteen days. Let that sit. We went from "it finishes your line" to "it finishes your sprint" in about three years. The line between autocomplete and coworker has been crossed, and it was crossed quietly, in production, on real code.
what actually changed
Not the code. The job around the code.
When the model wrote lines, the human wrote programs. Now the agent writes programs, and the human does something else: specifies what "done" means, reviews what got built, and owns the result. The work moved up a level. Specifying precisely, reviewing carefully, taking responsibility. Those were always the hard parts of software. They are now the whole job.
This is the part the discourse keeps missing. The fear is "the agent replaces the programmer." The reality so far is "the programmer becomes the person who is accountable for what the agent built." Accountability did not get automated. It got concentrated. Fewer hands on keyboards, same weight on shoulders. More weight, if anything, because now you are signing off on work you did not watch being written.
what breaks
Sixteen days of unattended work sounds like magic until you ask what sixteen days of unattended mistakes looks like.
Context rot. Over a long run, the agent's understanding of the goal drifts. Day one it is building what you asked for. Day twelve it is building what it thinks you asked for, three reinterpretations deep. Long-horizon work needs checkpoints throughout, because drift compounds. Waiting until the end to check is how a small wrong turn becomes a wrong destination.
Confident wrongness at scale. An agent does not get tired, which means it does not slow down when it should. A human programmer gets a bad feeling at 11pm and stops. The agent keeps going at 3am, generating clean, well-formatted, completely wrong code, and every hour of wrong work is an hour of cleanup later.
The review burden. Someone has to read sixteen days of output. Reviewing code you did not write is already the hardest part of senior engineering work. Reviewing code no human wrote, at machine volume, is a new discipline, and most teams do not have it yet. The bottleneck moved from writing to reading. Reading is slower.
the honest read
This is real progress, and it changes the economics of building software. A small team with good agents can now attempt things that used to need a big team. That is genuinely democratizing, and I do not say that lightly.
But the teams winning with coding agents are not the ones who fired their reviewers. They are the ones who got serious about specification and review, the two disciplines the industry has been underinvesting in for decades. The agent is only as good as the "done" you define and the eyes you put on the output.
Autocomplete made typing faster. Coworkers change the org chart. We are in the second one now. Act like it.