Code Review Is More Than Bug Detection—and AI Can't Replace the Human Judgment It Requires

There is more to code review than (automatable) detection

A new paper argues coding agents can replace human code review because they match its stated functions: defect detection, style enforcement, knowledge transfer, and awareness. This response pushes back, calling that a 'substitution myth.' It points to what can't be decomposed into functions: a reviewer's genuine confusion as a signal of poor design, skepticism about whether a change is even necessary, noticing what's missing, calibrated attention based on who wrote the code, bidirectional learning, operational context that lives outside repos, and personal accountability. Code review is also coordination, sensemaking, and governance.

The most fundamental issue I have with the article is that it assumes code review is a first and foremost a detection process: you find defects, style violations, security issues, etc., and the assumption is that detecting these faster and cheaper is universally better. But code review is also a coordination process, a sensemaking process, and a governance process.
  1. dimbletimbers

    A defense of human code review I wish I saw more often, especially in light of the concerns people have about cognitive/comprehension debt: comprehension redundancy. At the end, if taken seriously, at least two people understand how the feature works (even if that number is, on average, trending closer to between one and zero). Ideally at least one of the two also comes away with a better understanding of the wider system and how the feature fits into or stands out from that landscape.

  2. n4r9

    There's been a lot of talk about the purpose of code review recently. It makes sense in the face of AI. Heres a link that was submitted a little while ago: https://mathstodon.xyz/@mjd/115096720350507897

    And in response I wrote a non-exhaustive checklist of things that a code review can look for:

    - Does it functionally achieve what it sets out to (as per tacker issue or PR description)?

    - Does it have extraneous code? Leftover debug prints, private API keys etc...

    - Does it have any obvious defects? Memory leaks, un-handled edge cases, security flaws, obsolete API calls, etc...

    - Could it be more understandable? Add/remove abstractions, better variable/method names, more/less functional etc...

    - Is the style consistent with the codebase and/or style guidelines?

    - Are there obvious performance improvements? Hashset instead of list, lazy evaluations, etc...

    - Is it sufficiently well tested?

    I think LLMs are okay at most of these, and worst at the first.

  3. clintonb

    I am contemplating code review within my own organization, and the question I return to is:

    > Does this organization prioritize human learning?

    That has been my primary motivator for code reviews. I want to teach and learn from others, especially given the decreasing levels of collaboration due to increased AI usage.

    The sad truth is that all of my feedback just goes straight to agents. Maybe 10% is reacted to by a human, so I’m left wondering if there’s any value to a real review aside from poorly training robots to do my job, and further atrophying the abilities of my team members.

  4. metalspot

    The true reason why code review is universal is that it provides a liability shield for negligence. Negligence is interesting. It has nothing to do with whether or not you ship something broken. As long as you follow a process that attempts to not ship something broken, then you are not negligent.

    Engineers played along with this farce because code review served valuable team collaboration, coordination and management functions, about which the author of the article is correct.

    Understanding a system by reading code is harder than understanding a system by writing code.

    If AI can generate code at 100X, 1000X, or 10000X human capacity (no ceiling here), and you are gated on code review as your mechanism for system understanding, then a team's productive output will barely increase.

    If companies want to compete in the world of AI generated code, human code review has to go. The only question is, what replaces it?

    Continuing to apply human code review to AI generated code is negligent, if you are shipping at AI generation speed, with that as your only gate, and no other systems and processes to validate correctness and limit risk.

    On the engineering side we can adapt easily.

    Code review was never about finding bugs. When we do code review the first thing we check is: "do the tests pass?" Then we look at the change and the test coverage added for it and ask: "does the test coverage adequately demonstrate the functionality of the code?" The we ask: "What is the scope and potential […]

  5. ChicagoDave

    I was just at the Explore DDD conference in Denver and a portion of Friday was sitting at the cafe tables informally discussing the impact of GenAI on software engineering with notable people.

    Most of these people were deeply concerned that if we lean into using GenAI for “everything” that our collective knowledge will dissipate.

    I was the vocal contrarian. There are many historical examples of humans obfuscating knowledge to simplify progress.

    Does anyone solder their own microchips at scale anymore? No. We have highly sophisticated robots and machinery to do that work with extraordinary outcomes.

    In software engineering, if you remove “coding” as a discipline you’re left with all the other aspects of designing software which I contend can be retargeted in college CS curriculum.

    The leap isn’t about code reviews. It’s about design reviews and that’s where better outcomes are served regardless of whether GenAI is involved or not.

    I have a roughly year old codebase at https://github.com/ChicagoDave/sharpee/ that is designed by me, but generated by Claude Code with my own skills and agents as guardrails. I’m fairly certain the code I extract from Claude doesn’t require human review, but the design of the system and its changes are continually reviewed by me.

    My contention is that we “collectively” are still trying to discern where the AI/human line is and most are still “holding” that line to human interactions.

    Let it go. Define what part you do need human decisions on and foc […]

  6. bob1029

    One of my clients has an automatic "best practices" AI robot that runs each time you create a PR. It is pure downside. Even the developer responsible for creating it admits as much.

    However, for some weird reason it's still in place. This is the part that actually concerns me. Ignoring the bullshit comment is trivial. The quiet and relentless accumulation of entropy is happening everywhere. This is why GitHub crashes at noon every business day.

  7. grumpy_coder

    Code review was pretty rare until Github launched in 2008 and made it very easy to add to your process.

    I honestly didn't see as much improvement in code quality as people assume. Friends approve their friends PRs. Most coding cultures really push back on refusing a PR.

    Automated testing in CI was the much bigger improvement, at least the code compiles now.

  8. fwlr

    If you will tell me precisely what it is that my machine cannot do, then I can always prompt my machine to do just that.

    - John von Altman

    Articulating “what humans can do, that AI cannot” is a mug’s game. If you specify it well enough, they just paste your text into their /goal prompt box and ralph loop their agent swarm until it produces something too exhausting to distinguish from doing the thing. If you don’t specify it well enough, then you’re just doing human-centric magical thinking to move the goalposts etc etc.

More from this day

2026-09-27