The Road to Next — your interactive course for Next.js with React

Toward a Self-Improving Agentic Code Review Loop

Robin Wieruch •

Earlier this year I wrote about how we let AI agents review code against our documented patterns, because the review bottleneck had moved from writing code to reading it. That worked, and then it created the next problem: once every developer has a review agent, who reviews the reviewers?

Read More
Agentic Code Review: Pattern Matching for AI

As a freelance AI engineer, I introduced a review process to a product team over the last three weeks. Every developer owns a personal AI review skill. They fire it at each other’s pull requests and walk away. The author’s agent answers. And a scorecard, which I ran for the first time after those three weeks and plan to run weekly, tells us whose review skill is actually any good. This post walks through the loop, the numbers from the first three weeks, and where it still has weak spots.

To give you an idea of the scale first, this is what the review skills of five developers reported in those three weeks, and what became of it:

Three weeks of AI review findings
Every inline finding the review skills posted on the team's pull requests, by what happened to it.
789
findings reported by AI review skills
259
confirmed fixed (33%)
245
more with the code changed (31%)
31
refuted as wrong (4%)
259Fixed, the author confirmed the fix in a reply
245Code changed, no reply, probably fixed, not confirmed
30Valid, deferred, accepted for a follow-up
77Declined, not worth the change
31Refuted, the finding was wrong
147No verdict, unanswered, still open or unclear

For roughly two out of three findings the code changed afterwards, and one in twenty-five was plain wrong. No reviewer read these findings before they went out. Agents wrote them, and in many cases an agent on the author’s side fixed or answered them.

Every Developer Brings Their Own Reviewer

The obvious way to do AI review is one shared reviewer for the whole repository. One prompt, one configuration, maintained by whoever cares most. We deliberately did not do that.

Instead, every developer keeps a review skill on their own machine: a set of instructions their coding agent follows when asked to review a pull request. One developer’s skill hunts for duplicated code, another one reads the API contract first, a third one runs an adversarial pass. Nobody is asked to align. The reasoning is the same as for human review: five people who think alike find the same five things. A single shared reviewer also overfits. It gets tuned against the findings the team already knows to look for, and it stops surprising anyone.

One shared reviewer
It looks where it was tuned to look. Run it five times and it finds the same things five times; the rest of the change stays unseen.
Five reviewers, each their own
seen by one reviewerseen by two
Every skill looks somewhere else. They overlap a little, and together they see most of the change.
A conceptual picture, not a measurement.

The second rule is fire and forget. When a pull request opens, every teammate is invited to point their review skill at it. The agent posts its findings as inline comments, and the developer who started it does not look at them again. That sounds careless, and for a human review it would be. But it is the only way the volume works. A reviewer who has to read their own agent’s review first is back to being the bottleneck.

Each comment says which model wrote it, so nobody mistakes it for a colleague’s opinion:

text
[Model Name]:
> [SUGGESTION] `user-list.tsx:56` The empty check is repeated
> in every list component. The base list already knows whether
> it is loading, so it can make this decision itself.

Fire and forget still leaves one chore: someone has to fire. Right now every developer keeps track of the open pull requests themselves, which ones they already reviewed and which ones are ready for another round. The first developers on the team are automating that with one top-level agent running in a loop. Every 15 minutes or so it checks for pull requests it has not reviewed yet, and for each new one it launches a subagent that runs the whole review.

The alternative would be to host everyone’s review skill in the cloud, so that reviews arrive on their own the moment a pull request opens. I do not think that works for us. A review skill is only half of a reviewer. The other half is the harness around it, the developer’s own agent setup, and that does not move to the cloud with a prompt file. It would also put every skill in one shared place, which works against the diversity we wanted, and it would move the token cost away from the developer who chose to run the review.

The Author’s Agent Answers

Fire and forget moves the work to the author. A pull request can collect twenty, thirty, sometimes fifty findings from different agents, and no human reads fifty findings carefully.

So the author answers with an agent as well. It takes each finding, checks it against the actual code, and decides: fix it, decline it because it is not worth the change, or refute it because the finding is wrong. Then it leaves a reply in the thread, again marked as written by a model. A fix names the commit. A refutation brings the evidence.

text
> Non-issue: the field is required by the schema, so an empty
> value is rejected before this branch can run.

Written by Model Name

The answering agent has the harder job of the two, and it cannot be a shallow one. What I have seen is that a single agent walking through fifty findings in one session fills up its context window long before the end. From then on it does not judge findings anymore. It waves them through or waves them off. What works better is handing each finding, or a small batch of related ones, to its own subagent that starts fresh, reads the code, and comes back with a verdict. I can recommend that to the team, but I cannot prescribe it. How a developer’s agent answers is as much their own setup as how it reviews.

One agent answers everything
context windowfull
judged against the codewaved through or waved off
Every finding it reads stays in the session. Partway through, the context is full, and the later findings get a verdict without a real look at the code.
A subagent per batch of findings
judged against the codecontext used per subagent
Each subagent starts fresh, reads the code for its few findings, and hands back verdicts. No session ever gets near its limit, so the last finding gets the same attention as the first.
A conceptual picture, not a measurement.

If the first round produced many findings, a second round is fair game after the fixes landed. One rule turned out to matter a lot here: a later round has to read the earlier threads first. Without that, one agent recommends moving to approach A, the author moves, and the next agent recommends moving back to B. Two agents without shared context will happily argue a pull request in circles.

The review loop for one pull request
Agents do the volume on both sides. A human still gives the final review.
  1. A pull request opens
    The author asks the team for reviews, as always.
  2. Every teammate fires their own review skill
    Fire and forget: each agent posts its findings as inline comments, and the developer who started it moves on.
    skill Askill Bskill C
  3. The author's agent answers every finding
    It checks each finding against the code, fixes what holds up, and leaves a reply in the thread either way.
    fixeddeclinedrefuted
  4. Another round, if the first one found a lot
    A later round reads the earlier threads first, so it does not argue the code back to where it started.
  5. A human reviews last
    Is this the right change, was it asked for, does the design fit? The pattern matching is already done.
The feedback loop on top
Once a week, a scorecard reads the replies from step 3 and shows how each review skill did. Every developer tunes their own skill from it.

The Human Review Is Still There

None of this replaces the human review. It changes what the human review is for.

After the agents are done, a developer still reviews the pull request and puts their name on it. They are free to use their agent to explore the change, ask questions, check a hunch. But it is not fire and forget anymore. The inline comments are theirs, and they stand behind them.

What they no longer have to do is the pattern matching. Naming conventions, a missing test, a duplicated helper, an unhandled error path: the agents found those. The human question is a different one. Is this the right change? Was this feature asked for? Does the design fit where the product is going? Those are questions an agent answers confidently and badly.

Read More
Your AI Output Is Someone Else's Input

Scoring the Reviewers

Here is the part that made the whole setup interesting to me. Every finding has an author (a review skill) and, ideally, a verdict (the reply). That is a dataset.

After three weeks I sat down with an agent and scraped it: every inline review comment on every pull request in that period, grouped by the developer whose skill posted it. A second pass read the replies under each finding and classified the outcome as valid, declined, or refuted. The result was 789 findings from five review skills. Four replies were too unclear to call. Of the rest that got an answer, 73% were accepted and 8% were refuted.

The averages are not the interesting part. The spread is. The totals above are real. For the breakdown below I made up the names and adjusted the numbers per skill, so that no teammate is recognizable and every outcome a scorecard like this can produce shows up once:

Five review skills, three weeks
Up is more precise, right is deeper. Select a review skill to see how its findings ended. Made-up names, adjusted numbers.
↑ valid share of answered findings
100%50%0%
0valid findings per shared pull request →5
Jonas's review skill: balanced97 findings, 56 answered
48 valid6 declined2 refuted
86% of its answered findings were accepted. A typical review posts 4 findings, and on pull requests shared with other skills it lands 2.9 valid findings each.

Each of the five skills tells a different story:

  • Mara’s skill is never wrong, and it almost never says anything. A typical review posts one or two findings, and on pull requests that other agents also reviewed, it misses most of what the others catch.
  • Jonas’s skill sits where everyone wants to be: reasonably deep and right most of the time.
  • Lena’s skill finds far more than everyone else, but a quarter of its answered findings get declined as not worth the change. It buries the author in nits.
  • Tariq’s skill is deep and often wrong: a quarter of its answered findings were refuted.
  • Felix’s skill is both shallow and imprecise.

Precision alone would have put the wrong skill on top. The skill with a perfect score is the one I would improve first, because a reviewer who only speaks when certain is leaving most of the review undone. This is why the scorecard needs both axes. Precision is how often a skill’s findings are accepted, so a declined nit costs as much as a wrong finding. Depth is how many valid findings it lands when several skills look at the same change.

It also gives every developer something concrete to do. The shallow skill needs to look for more. The one that is often wrong needs a verification step before it posts. The deepest one needs to stop posting every nit. Nobody has to agree on a shared prompt for that. Everyone tunes their own.

This is the kind of verdict the scorecard hands to each developer:

This week's verdict per review skill
One reading of the numbers and one thing to change before next week's scorecard. Made-up names, adjusted numbers.
Mara
Precise but shallow
Every answered finding was accepted, but a typical review posts one or two.
Next week: Look for more. The skill can afford to be wrong sometimes.
Jonas
Balanced
Right most of the time, with a handful of findings in a typical review.
Next week: Keep the precision and push the depth a little further.
Lena
Deep but nitpicky
Finds the most by far and is rarely wrong, but a quarter of its answered findings get declined.
Next week: Stop posting every nit. Keep what is worth a change.
Tariq
Deep but often wrong
Thorough, but a quarter of its answered findings were refuted.
Next week: Verify each finding against the code before posting it.
Felix
Shallow and imprecise
One finding per review, and close to half were declined or refuted.
Next week: Add a verification step first, then widen the search.

Watching a Skill Move

A single scorecard is a snapshot. The second one is where it gets useful, because next week’s scorecard shows whether the change worked. That is the whole idea: a developer does not have to guess anymore whether a tweak to their review skill made it better.

Here are the same five skills again, with an earlier scorecard added as hollow dots. Three weeks of data do not give us a real earlier scorecard for everyone yet, so those earlier positions are illustrative:

How the five review skills moved
The hollow dot is an earlier scorecard, the filled dot is the one above. Up is more precise, right is deeper. Select a review skill to see how it moved. Made-up names, illustrative positions.
↑ valid share of answered findings
100%50%0%
0valid findings per shared pull request →5
Jonas's review skill: balanced97 findings, 56 answered
48 valid6 declined2 refuted
86% of its answered findings were accepted. A typical review posts 4 findings, and on pull requests shared with other skills it lands 2.9 valid findings each.
Against the earlier scorecard: It moved the right way on both axes: deeper and more precise.

Every direction a skill can move in shows up once. Jonas’s skill moved the way everyone hopes for, deeper and more precise at the same time. Felix’s got more precise but stayed shallow. Tariq’s lost precision without finding more. Mara’s barely moved. And Lena’s traded a little precision for a lot more depth.

I have one real move to show, my own, and these numbers are measured and not on the chart above. About ten days before the scorecard ran, I changed my review skill to look harder. Comparing the latest week with the weeks before it, my typical review went from 4 findings to 12.5, while the valid share of answered findings slipped from 81% to 76%. So the skill got a lot more exhaustive and a little less precise.

That is a trade I would take again. I would rather lean toward depth until precision reaches a floor, somewhere around 70%, and only then pull back. The catch is that more findings are only worth it while they are worth acting on, so the next thing to trim is the nits, not the depth. Without the comparison I would not have known that there was a trade at all, or how much each side of it moved.

A Baseline to Measure Against

Five personal review skills are not the only reviewers on our pull requests. GitHub Copilot reviews every one of them as well, and nobody on the team tunes it. That makes it a useful baseline: a reviewer that is the same for everyone, answered by the same agents, scored by the same scorecard.

We loop it, too. The author’s agent works through its comments, pushes the fixes, asks for another review, and repeats that until a round comes back with nothing meaningful. The same two rules apply as for our own skills: a later round has to know what the earlier ones settled, and the loop needs a cap, because a reviewer that is asked often enough will always find something.

In the same three weeks, Copilot posted 385 inline comments, and 255 of them got a reply. Of those, 71% were accepted, close to the 73% of our own skills. But 16% were refuted, twice our rate, and it is shallow: two out of three of its reviews left no inline comment at all, and the ones that did typically left two. What it has over every personal skill is coverage. It reviewed 271 pull requests, while our own skills reached 94, because nobody has to remember to fire it.

So I would keep a baseline reviewer in the loop, whichever product it is, for two reasons. It catches the pull requests nobody got around to. And it gives the scorecard a fixed point: a personal review skill that cannot beat the reviewer everybody gets for free needs work.

The Silent Fix Problem

The scorecard has a weak spot, and it showed up immediately.

What happened to 789 findings
Only a reply tells you whether a finding was right. Everything else is a guess.
401Answered with a reply, the only ones with a verdict
245Code changed, no reply, probably fixed
91Resolved, no reply, code untouched
52Still open, not handled yet

Only about half of the findings were ever answered. Many of the rest were clearly handled: the commented lines changed afterwards and the thread was resolved. But a changed line is not a verdict. Maybe the finding was right. Maybe that code was rewritten for a different reason. To know, an agent would have to read the follow-up commits for every unanswered finding, and that burns far more tokens than reading a one-line reply.

The reply rate differed wildly per author, from almost every finding answered to almost none. That skews the score in ways that are hard to predict. A review skill whose findings mostly land on a careful responder gets every weak finding called out, while one whose findings land on someone who fixes quietly gets neither the credit nor the blame.

So the one rule this process really needs is small: every AI finding gets a reply. It does not matter who writes it, the developer or their agent, and it does not matter how it is worded. We decided against a fixed vocabulary, because aligning every developer’s replying agent is exactly the kind of alignment the setup tries to avoid. A model can read any reply and tell whether it says fixed, declined, or wrong. It cannot read a reply that is not there.

Where the Score Can Lie

I would not trust this scorecard blindly. Two things can bend it, and one thing it cannot see at all.

Findings per review
Each dot is one review of one pull request. The dark line marks the skill's typical review. A review that found nothing posts nothing, so it does not show up. This is the raw count of findings, valid or not. Made-up names, adjusted numbers.
Lena
Tariq
Jonas
Mara
Felix
1251020
shallowfindings in one reviewexhaustive →
Two skills mostly post one or two findings. The goal is to move every row to the right without letting precision drop.

The first one is Goodhart’s law. The moment a review skill is rated by its acceptance rate, the cheapest way to improve the rating is to post fewer, safer findings. The chart above shows what that would look like: two of the five skills mostly post one or two findings per review. Nobody has tuned for the score yet, the scorecard is too new for that, but this is the corner a precision-only score rewards. Where we want to end up is the opposite, exhaustive reviews instead of shallow ones. The scorecard’s second axis is the counterweight, and it only works on pull requests that several skills reviewed.

The second one is the judge. Right now the verdict comes from the author’s agent, and the author is the party that saves work when a finding gets declined. The score does not measure whether a finding was correct. It measures whether two agents agreed. What is missing is a judge outside the loop: something that runs once when a pull request merges, belongs to neither the reviewer nor the author, and gets spot-checked by a human every now and then. We do not have that yet.

And then there is the number I would like to have and cannot get from the threads at all: what did the human review, or production, catch that no agent flagged? That is the only real measure of whether the loop misses things.

Where This Could Go

If the measuring holds up, the direction is a self-improving loop. Review skills get scored every week, developers tune them, precision and depth go up together, and the human review gets shorter because there is less left to find.

The self-improving review loop
Every pass around the loop should leave the review skills a little better than the last one.
Review
Every developer's skill reviews the team's pull requests.
Reply
The author's agent answers each finding: fixed, declined, or refuted.
Score
A weekly scorecard rates each skill for precision and depth.
Tune
Each developer adjusts their own skill where the score is weak.
and back to the next review
Precision goes up, fewer refuted findings
Depth goes up, fewer shallow reviews
Human review shrinks, less left to find

One step further sits a classifier at the end of a pull request that says how confident it is that nothing is left. A small change with few findings and a high score could merge on its own. A large one could earn its way there through review rounds. Architecture and security changes would stay with a human no matter what the score says.

I want to be careful here, because that is an outline and not a plan. A model’s confidence is not calibrated just because it is a number. Before any score merges anything, I would record it silently for a while and compare it with what the human review still found. And there is a quieter risk on the way: when five agents have already approved a pull request, the human review turns into a rubber stamp long before anyone decides that it should.


What convinced me of this setup is that it became measurable. Three weeks ago, “my review agent is pretty good” was a feeling every developer on the team had about their own setup. Now it is a number with known weaknesses, next to four other numbers, and the first thing to fix in the whole loop turned out to be a missing one-line reply. That is a much better problem to have than the one we started with.

Never Miss an Article

Join 50,000+ developers getting weekly insights on full-stack engineering and AI.

AI Agentic UI Architecture React Next.js TypeScript Node.js Full-Stack Monorepos Product Engineering
Subscribe on Substack

High signal, low noise. Unsubscribe at any time.