# Counting the Cost of AI Code Review: Bugs Found, False Alarms and People's Time

> AI code review's main cost is human triage time, not tokens. Whether it pays depends on the precision and severity of findings, as curl and an ICSE 2025 study show.

Source: https://sheng.page/en/posts/ai-code-review-cost-benefit/ · Published: 2026-10-03 · Author: Sheng

Daniel Stenberg, who maintains curl, had been complaining about AI-generated vulnerability reports for a long time. In May 2025 he said publicly that the project had yet to see a single valid security report written with AI help. The reports looked convincing, properly formatted and confident in tone, and the maintainers had to read each one, try to reproduce it and reply, only to find it was not real.

In September the same year the other side of the story showed up. The security researcher Joshua Rogers ran several AI code-scanning tools against curl, sorted through the output and sent in a large batch of reports, and by early October around 50 fixes based on them had been merged. Stenberg called them "Actually truly awesome findings". He added that most were small mistakes and nits of the kind a static analyser finds, still worth fixing, and that several were quite impressive. His conclusion: "Powerful tools in the hand of a clever human is certainly a good combination."

A few months later, at the end of January 2026, curl ended its bug bounty anyway. Stenberg said the quality of submissions had collapsed, with plenty of obvious slop and even the reports that didn't look AI-written getting worse, and that the project needed to stop itself drowning.

Both are AI finding bugs. One produced dozens of real fixes, the other produced noise the maintainers couldn't keep up with. The models weren't that different. What differed was who did the work of telling real from fake. Rogers had filtered his findings before sending them, while the flood of anonymous reports pushed all of that filtering onto the maintainers. Bringing AI code review into a team comes down to the same sums.

## 1. Tokens are not the main cost

When a team adopts AI review, the cost everyone sees is the API bill. The one that gets missed is the time people spend deciding what to do with each comment.

An industry study presented at ICSE 2025, "Automated Code Review In Practice", looked at a company that introduced automated review based on Qodo PR Agent. It covered ten projects and 238 practitioners, with a closer analysis of three projects, 4,335 pull requests in total, 1,568 of which went through automated review. 73.8% of the automated comments were resolved, and most practitioners reported a minor improvement in code quality. The price was that average pull request closure time went from five hours 52 minutes to eight hours 20 minutes, and people reported faulty reviews, unnecessary corrections and irrelevant comments.

The numbers cut both ways. With more than seven in ten comments resolved, the tool was mostly saying something reasonable. With closure time up by two and a half hours, every comment needed someone to read it, think about it and respond, whether or not it turned out to matter. Resolved isn't the same as a bug found either, and the abstract doesn't break down how many of those comments pointed at problems that would actually have caused trouble.

## 2. A rough formula

Broken down, one round of AI review looks roughly like this:

```text
benefit = real problems found × cost of each one reaching production
cost    = model spend + total comments × human triage time per comment × hourly cost
        + changes made because of false alarms (things changed that shouldn't have been)
```

Total comments times triage time is usually much bigger than the model spend. Here is a worked example with made-up numbers, for illustration only. One review produces 40 comments, 8 of them real problems, a precision of 20%. Each comment takes six minutes on average to triage, four hours in total. If one of those eight is a bug that would have corrupted data, the four hours were well spent. If all eight are about naming and formatting, the four hours would probably have been better spent writing code.

So there is no general answer to whether AI review is worth it. It depends on two ratios: precision decides how much noise people have to read, and the severity of the real problems decides whether reading it was worth it.

## 3. Adding a verification layer

Since human triage is the expensive part, the obvious move is to let the models filter first. One set of agents looks for problems, another set takes each finding and tries to prove it wrong, and only the findings that survive go to a person.

Extending the made-up numbers: each of the 40 comments gets three verifying agents, 120 extra model runs. Twelve comments survive, seven of them real. Human triage drops from 40 comments to 12, and from four hours to about one hour 12 minutes, while precision rises from 20% to nearly 60%. There are two costs. The model spend goes up several times over, and one real problem was thrown out during verification.

The second cost is the easier one to miss. Verifying agents make mistakes too. They may not see the full call chain, or they may be talked round by a guard clause that looks plausible, and mark a real problem as invalid. With a verification layer the list people see gets much shorter, and whatever was dropped from it really does go unread. A safer approach is to sample the rejected findings regularly, so you at least know roughly how often this layer throws out real problems.

Whether it pays off can be read straight off the formula. The human time saved has to exceed the extra model spend plus the expected cost of the real problems it drops. It pays best when people are expensive, comments are many and precision starts out low. When there are few comments to begin with, or the code is something people will read line by line anyway, an extra layer is just extra spend.

## 4. Which code is worth reviewing this way

Working back from the cost of a problem reaching production, the places worth the investment are fairly specific:

- Code that touches money, permissions or data consistency, such as debits, reconciliation and authorisation checks, where one bug costs far more than a few more model runs.
- Concurrency, retries and state transitions, which human reviewers tend to miss. Models are more patient than people with questions like "what if these two requests arrive at the same time".
- Old code nobody has looked at closely, and large amounts of AI-generated code that people have only skimmed.

Formatting, naming and wording in docs are cheaper to hand to a linter and a formatter, and need no human judgement.

## 5. How to tell whether it works

Without records, all you have is a feeling. At a minimum, record for every AI comment whether it was valid in the end, how severe it was, how long it took to decide and, if valid, who fixed it. After a few dozen pull requests you can work out your own team's precision and the average cost of each real problem, instead of quoting a vendor's figures.

It's also worth comparing the same pull requests against human review: what AI found that people missed, and what people found that AI missed. Rogers noted that none of the tools he tried caught a known infinite-loop bug in a different project. Tools have different blind spots from each other, and from people.

## Notes

- In both cases at curl it was AI finding bugs. What separated a flood of noise from around 50 merged fixes was who paid for telling real from fake.
- The main cost is the time people spend on each comment. In the industry study, pull request closure time went up by two and a half hours.
- The benefit depends on two ratios: precision decides how much noise gets read, severity decides whether reading it was worth it.
- An adversarial verification layer raises precision and saves human time, at the cost of several times the model spend and some real problems thrown out. Sample the rejected findings regularly.
- Money, permissions, concurrency and state transitions are where it pays most. Leave formatting and naming to linters.
- Record validity, severity and triage time for every comment, and work out your own team's numbers.

## Further reading

- The Register, [Curl project, swamped with AI slop, finds not all AI is bad](https://www.theregister.com/2025/10/02/curl_project_swamped_with_ai/): what Rogers found with AI tools, how many fixes were merged, and Stenberg's comments.
- The New Stack, [Drowning in AI slop, cURL ends bug bounties](https://thenewstack.io/drowning-in-ai-slop-reports-curl-ends-bug-bounties/): when and why curl ended its bug bounty. The article also notes that no figures on the share of slop were published.
- Umut Cihan et al., [Automated Code Review In Practice](https://arxiv.org/abs/2412.18531): the ICSE 2025 SEIP industry study, the source of the 73.8% resolution rate and the longer closure times.
- Wikipedia, [Code review](https://en.wikipedia.org/wiki/Code_review): defect detection rates for human review and formal inspection, useful as a baseline.

> Ideas and technical judgement by Sheng; drafted with Claude · cost figures in the worked example are illustrative.