Do AI coding agents actually make senior engineers faster?

Microsoft measured 24% more merged pull requests, but longer merge times and doubled reviewer workload complicate the productivity claim.

Do AI coding agents actually make senior engineers faster?

Microsoft measured 24% more merged pull requests from regular agent users. Merge times grew, and reviewer workload roughly doubled. The terminal looked fast. The engineering system wheezed. Claude wrote the feature, tests and pull request before I finished an espresso. Bellissimo. I spent 40 minutes reconstructing undocumented assumptions, checking whether the tests proved anything and finding a redundant abstraction. The code looked finished. My afternoon disagreed. During Microsoft’s four-month early-2026 rollout of Claude Code and GitHub Copilot CLI, regular users merged 24% more pull requests per engineer per day. Yet reviewer workload roughly doubled, human-review coverage fell from 89% to 68%, and AI-authored PRs took 22% longer overall to merge, according to Microsoft research summarized by TechRepublic. That matters more than stopwatch demos. Senior engineering starts after the magic typing animation.

My take: AI rapidly manufactures senior-engineering artifacts, then dumps context and accountability on overloaded humans.

PR volume is a gorgeous bad dashboard

Microsoft measured a 24.0% rise in merged PRs per engineer per day, likely between 14.5% and 33.7%. A placebo test moving the rollout earlier found no similar jump.

After years of productivity claims based on eight developers doing timed JavaScript, four months of enterprise telemetry feels luxurious.

Seniority mattered less than expected: individual contributors and principal engineers gained similarly. Agents helped experienced developers submit more code.

I had confidently expected the opposite.

But Microsoft measured merged PRs, not customer impact, security or long-term maintainability. More reviewable containers can beautify a dashboard while delivery sounds like my old Fiat climbing a hill outside Ivrea.

A separate enterprise study tracked 802 developers and 196,212 pull requests at a mid-sized company from January 2024 to April 2026. In mid-2025, its CTO announced a “2x mandate”, measured by merged PRs per engineer per month.

The dashboard delivered: throughput rose from 21.2 to 44.3 merged PRs per active developer, or 2.09 times the pre-mandate baseline.

Then came the invoice. AI-authored PRs took roughly 20% longer after the first human review and 22% longer overall to merge. Human-review coverage fell as automated AI review expanded.

My publishing infrastructure once reported soaring referrer traffic. Google Search Console showed roughly 99% of that “audience” was bots.

Lovely chart. No readers.

PR volume creates the same trap: reward the pipeline’s cheapest stage and everyone produces units for someone else to inspect. It’s judging a restaurant by plates leaving the kitchen without checking whether they reach tables.

Very efficient. Nobody ate.

Polished PRs can counterfeit experience

Agent-generated pull requests look eerily competent: clean descriptions, plausible names, passing tests and confident repository explanations.

I appreciate the polish. For five minutes, it can impersonate architectural understanding.

On July 13, 2026, Kartik Ghanshyambhai Pansuriya, Ehsan Ghorbani, Deepak Singh and Eman Abdullah AlOmar published research predicting acceptance and review effort for human and agent pull requests. Using submission-time information, their tree-based models predicted acceptance with F1 scores above 0.95.

Inputs included textual clarity, metadata, repository context, timing signals and lightweight diff statistics. The researchers wrote that “acceptance prediction is feasible from early signals”; text quality and metadata were among the strongest predictors.

Machines generate those signals for pennies. A tidy description once implied organized thinking. Now it may mean someone typed /pr and refilled a water bottle.

Review effort was harder to predict. Comment counts and time-to-merge depended heavily on reviewer availability, local workflow and team habits. The model recognized mergeable-looking artifacts but struggled to estimate their human cost.

The surface is legible. The organizational cost appears when a senior engineer opens the diff.

Another study examined 567 Claude Code PRs across 157 open-source projects. Although 83.8% eventually merged, only 54.9% needed no changes. Humans revised the other 45.1%, especially for bug fixes, documentation and project-specific standards.

Neither study proves juniors gain more speed than seniors.

The reviewer must still decide whether those choices survive contact with the codebase.

I fall for the presentation too: neat fixtures, comprehensive-looking cases, green checkmarks. My pulse drops for three seconds. Then I remember the model wrote both implementation and exam.

Plausibility is excellent theater.

One principal engineer, infinite queue

Maliha Noushin Raida and Daqing Hou studied 25,264 agentic PRs across 2,361 popular GitHub repositories in their July 2026 paper, Early Adoption of Agentic Coding Tools by GitHub Projects. Most contributions skipped elaborate human-agent collaboration.

One person handled them.

A single developer reviewed or modified the agent’s work in 78.9% of cases. Including PRs accepted unchanged by one reviewer, one-person oversight covered nearly nine in ten.

Raida told Help Net Security that even busy projects followed this pattern: “Even among the most active small repositories (i.e., small teams with more than 30 agentic PRs), the majority of agentic PRs continued to follow a single-reviewer workflow.”

She added: “This suggests that increased agentic activity did not necessarily lead to more distributed review practices.”

Tape that above every adoption dashboard. Generation scales horizontally. Trust keeps arriving at one human desk.

Adoption remains early. The median repository in Raida and Hou’s dataset produced only one or two agentic PRs over three months; just 25 projects matched a working developer’s pace under the reference benchmark.

Companies celebrate extra motorway traffic. The tollbooth still has one employee.

Single-reviewer PRs merged at 81.2%, versus 80.3% under multi-reviewer or committer workflows. Raida said: “The merge rate … was very similar between the two collaboration patterns: 81.2% for single-reviewer pull requests and 80.3% for multi-reviewer/committer pull requests.”

That says little about later reverts, follow-up fixes or durability; the study measured acceptance within its window. More reviewers also add coordination overhead. Seven people on Zoom do not guarantee wisdom.

The problem remains: contributions cost almost nothing to create, while review stays concentrated.

Senior engineering includes invisible scar tissue: spotting a doomed abstraction, remembering a customer promise in a three-year-old Slack thread, or knowing why an ugly workaround survived after the clean solution failed in production.

GitHub’s contribution graph has no square for scar tissue.

A clean local implementation can still cause product disaster. The danger lies between components, beyond management’s latest line-count metric.

A senior engineer working on a laptop, surrounded by code snippets and AI tools, illustrating productivity in tech.

Legacy code eats the demo for lunch

Enterprise data showed the strongest AI-related output growth in newer repositories. Legacy codebases gained little regardless of seniority; principal engineers and individual contributors followed similar patterns.

The repository mattered more than the engineer’s title.

New codebases have cleaner conventions and fewer hidden dependencies. Mature systems contain unfinished migrations, undocumented customer exceptions and compromises retained because every “obvious” replacement started another fire.

I grew up in Ivrea, Olivetti’s town, and studied computer engineering at Politecnico di Torino. Engineers taught me an Italian lesson: if an old machine has a strange metal bracket, assume somebody painfully learned to keep it.

Legacy software is mostly strange metal brackets.

According to Codacy’s analysis, LinearB’s 2026 Software Engineering Benchmarks Report found agentic AI PR review pickup times 5.3 times longer than for unassisted PRs. AI-assisted PRs waited 2.47 times longer.

These are LinearB figures cited by Codacy, not Codacy’s data. Still, they match Microsoft’s review pressure: finished-looking diffs arrive faster than humans can gain enough context to judge them.

Trust remains low. Stack Overflow’s 2025 developer survey put trust in AI accuracy at 29%. Plausible code may take longer to inspect than broken code: syntax errors shout; a billing-logic mismatch waits quietly until Friday evening.

My nonna would disown this fish analogy, but a chef spots spoiled fish immediately. One suspicious smell in a beautiful fillet takes longer because dinner depends on judgment.

Generated code poses the same verification problem: I must assess intended behavior, architectural fit and forgotten edge cases.

Addy Osmani and Jason Gorman call this comprehension debt: the widening gap between code a team owns and code its humans understand. When requirements change or production breaks, everyone reconstructs reasoning nobody performed.

In mature products, my advantage is not typing speed. It is knowing which innocent change wakes billing, breaks an enterprise integration or destroys Sunday lunch with an incident.

Where agents genuinely earn their keep

Microsoft found large, persistent gains, especially among frequent users.

Developers using agents at least five days per week gained above 50%; three-day users gained roughly 15%. The overall effect persisted throughout the four-month study.

That is serious acceleration. Denying it would look ridiculous.

Tool choice also mattered at Microsoft: in comparable weeks, Copilot CLI users saw roughly 2.2 times the PR lift of Claude Code users. Microsoft’s researchers warned that its internal environment prevents universal product rankings, so no Champions League table from one company’s telemetry.

Mature organizations may control the risk. Gearset CEO Kevin Boyle wrote in TechRadar Pro that “Seventy-six percent of enterprise teams are reviewing AI-generated work at least as rigorously as human-written work.”

That includes 33% applying stricter checks. Although 82% of surveyed teams used AI during building, only 58% used it during release, where production exposure rises.

Good. Teams use agents where errors are easier to catch, then tighten human control near production.

Acceleration is credible when tasks are bounded, repositories clean, tests meaningful and review capacity planned. Then I can delegate boilerplate, migrations or contained features and spend the savings on contextual decisions.

I use agents this way constantly.

A food processor makes chefs faster at chopping onions. I still want the chef choosing the fish and running Saturday service.

Piano, amico.

Buy verification before another pile of licenses

Review capacity is infrastructure. Before buying more agent licenses, I want reviewer pickup time, post-review merge latency, PR size and production incidents on the dashboard.

I also want rework rates. License use measures enthusiasm for generating code, not whether software safely reaches customers.

CircleCI’s 2026 data, cited by Codacy, covers more than 28 million CI workflow runs across over 22,000 organizations. Feature-branch throughput rose, but median main-branch throughput fell nearly 7%, with main-branch success dropping to 70.8%.

Activity grew. Safe delivery did not keep pace.

A minority escaped: main-branch throughput increased 26% while feature-branch activity rose 85%. Codacy attributes this to stronger automated checks, cleaner review signals and clearer merge policies.

Move senior judgment upstream. Before generating a giant diff, write the execution plan, constraints and acceptance criteria, then have another human inspect risky assumptions. Reviewing ten lines is cheaper than reverse-engineering intent from 1,400 generated lines.

Keep PRs small enough for a tired human. “The agent produced all of this together” explains the mess; it does not earn one merge button.

Deterministic tools should handle formatting, linting, type checks, secret detection, dependency scanning, SAST, coverage requirements and complexity thresholds before senior review.

Humans can inspect architecture, business behavior, reversibility and cross-team impact. I will not spend senior attention hunting missing semicolons in 2026. We have machines for that, grazie.

Pansuriya and his co-authors found acceptance highly predictable but review effort hard to forecast because availability and team workflow mattered. Even strong models inhabit messy companies full of calendars, incidents and lunch.

Founders should replace the “2x developer” dashboard with one harsher metric: production outcomes per hour of senior review.

If output doubles while senior review hours triple, the impressive demo was financed by the attention of the company’s most expensive people.

Accountability will have a human price tag

By July 2028, nearly every serious engineering team will have broadly comparable coding agents. Models will improve, but availability will offer little competitive advantage. Most companies will generate more code than they can safely verify.

The scarce engineer will understand the system well enough to approve changes and accept production responsibility.

Companies measuring PR volume will imagine an army of synthetic developers. Many will have built ten assembly lines with one quality inspector.

Whenever an AI rollout claims a 30% productivity gain, show reviewer hours, merge latency, rework and production outcomes beside it. If the gain disappears there, senior engineers funded the launch with noise-canceling headphones and quietly lost the will to live.

By 2028, code generation will be a commodity budget line. Engineers able and willing to sign “ship it” will cost more every year.

Frequently asked questions

Do AI coding agents actually make senior engineers faster?

AI coding agents can increase senior engineers’ pull-request output, but they do not automatically improve end-to-end delivery speed. Microsoft measured 24% more merged pull requests among regular users while AI-authored pull requests took 22% longer overall to merge and reviewer workload roughly doubled.

Why do AI-generated pull requests take longer to review?

AI-generated pull requests can look polished while leaving assumptions, architectural fit, business behavior and edge cases for humans to verify. Review effort also depends on reviewer availability and team workflow. Faster code generation therefore creates more reviewable artifacts without expanding the scarce human capacity needed to approve them.

Where are AI coding agents most useful?

AI coding agents provide the clearest acceleration on bounded tasks in clean repositories with meaningful tests and planned review capacity. They are useful for boilerplate, migrations, repetitive plumbing and contained feature work, while humans retain responsibility for architecture, business behavior, reversibility and cross-team impact.

Sources

Related reading

Luca

Luca

Luca by the way is the personal blog of Los Angeles based entrepreneur Luca Capula. A true Italian who lives between Torino and LA.

More posts →