update
Apr 17, 2026
By Teun
AI token budgets are inflating code output, not productivity
New data from developer analytics firms suggests AI coding tools are increasing code volume, but not necessarily useful output. Companies such as Waydev, GitClear, Faros AI, and Jellyfish say engineers are approving more AI-generated code while also spending more time revising it later.
AI coding tools are generating more code, but that is not the same as generating more useful code. New data from developer analytics firms such as Waydev, GitClear, Faros AI and Jellyfish suggests that some teams are approving a higher volume of AI-written changes while later spending more time cleaning them up.
The pattern has picked up enough attention that some people in software teams have started using the term “tokenmaxxing” for the practice of pushing AI assistants to produce as much output as possible. The label is informal, but the underlying behavior is easy to understand: if a tool can write code quickly, it can also fill repositories with more lines, more files and more changes to review.
That can make productivity look better on the surface. A developer who accepts a large number of AI suggestions may appear to be moving faster, especially if the only metric being watched is output volume. But the firms tracking engineering activity say the more important question is what happens after the merge, when other engineers have to read, test and repair the code.
Waydev, which tracks developer activity, and GitClear, which analyzes code changes, have both reported signs that AI assistance is changing the shape of commits and reviews. Instead of a clean productivity boost, the data points to a rise in churn, the software-industry term for code that is quickly changed again after it lands.
Faros AI and Jellyfish, which also focus on engineering analytics, have pointed to a similar gap between quantity and quality. Teams may ship more accepted AI-generated code, but the follow-on work can increase as engineers revise generated snippets, untangle edge cases or replace code that does not fit the rest of the system.
This matters because software development is not judged by how much code is written. It is judged by whether the code works, stays maintainable and does not create new problems for security, testing or operations. A feature that arrives faster but requires repeated cleanup can add hidden cost instead of removing it.
The issue is especially visible with AI coding assistants and agents, which are designed to draft functions, suggest fixes and, in some cases, make multi-step changes across a codebase. These tools can save time on boilerplate and routine tasks, but they also produce output that needs human review. That review burden can grow when developers accept large amounts of generated code without fully checking how it fits into the existing system.
Some engineering leaders have been pushing for metrics that go beyond raw output, including the amount of rework, the number of follow-up commits and the time spent in review. That is because AI can inflate productivity dashboards without improving the underlying work.
The concern is not that AI coding tools are useless. It is that the easiest thing to measure, how much code gets produced or accepted, is often the least meaningful thing to measure.
As more companies roll out AI helpers across development teams, the central question is shifting from how much code the tools can write to how much extra work that code creates after it is merged into the main branch.