Harvard Study Finds AI Coding Tools Create Review Bottlenecks Without Increasing Software Output
An analysis of 700 software organizations reveals autonomous coding agents drive up code volume while lengthening review times by nearly 50 percent.

While artificial intelligence tools have dramatically increased the volume of code generated across engineering teams, they have failed to accelerate the delivery of finished software due to human review bottlenecks, according to a study first reported by Ars Technica (https://arstechnica.com/ai/2026/10/ai-coding-agents-generate-more-code-but-not-more-software/).
Conducted by Harvard University researchers Fiona Chen and James Stratton, the study evaluated aggregated operational data from engineering analytics platform Jellyfish. The dataset encompassed 300 million individual work events—such as commits and pull requests—and issue-tracking records across more than 700,000 employees at over 700 software firms from 2021 through March 2026.
To evaluate how AI affected productivity, the researchers tracked the adoption timelines of both autocompletion assistants and autonomous coding agents, which independently write and submit code based on prompts. Using difference-in-differences regression analysis, they compared organizational productivity metrics before and after the introduction of these tools.
The findings show that adopting autonomous AI agents caused raw code output to surge. Across the sampled firms, agent deployment resulted in an average 30 percent increase in total lines of code written, a 20 percent increase in total commits, and a 23 percent increase in pull requests. However, this output failed to yield higher software throughput: resolution rates for Jira-tracked Issues and Epics showed no statistically significant change, with no compositional shift in task size or complexity.
Instead, the added code created significant friction in the verification pipeline. The average duration between pull request submission and merging increased by 49 percent after teams introduced AI agents. Furthermore, the share of pull requests requiring revisions nearly doubled, and the average number of reviewer comments per pull request climbed by 35 percent.
To cope with the heavier review load, companies shifted engineering labor, leading to a 14 percent increase in the share of employees conducting code reviews. When cross-referencing Jellyfish activity with LinkedIn data, the authors found no statistically significant employment reductions attributable to AI adoption.
Automated code review tools have provided minimal relief so far. Although 80 percent of examined firms used some automated review tooling by March 2026, AI agents accounted for only 23.3 percent of review comments and 10.8 percent of pull requests, leaving human developers responsible for the vast majority of verification work.
By March 2026, 95 percent of the firms in the study had introduced AI coding agents. The researchers noted that engineering organizations are still adjusting workflows, suggesting that the balance between faster initial drafting and extended code review times may evolve as teams refine how and when to deploy autonomous agents.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



