I'm tech lead and I basically don't review PRs, and I tell people this, with a caveat - if you can tell me what you specifically want me to review, for what specific purpose, I'm happy to!
So "can you review this bit for race conditions" is great, love it. This forces people to actually think about what in their code they should be suspicious of, if anything.
"Can you review this" [link to PR] is getting a rubber stamp because humans have never been good enough at "just spotting bugs" to make this worth it and now any LLM is better than a human.
And for the purpose of understanding - PR time is too late. I have not reviewed a PR ("for real") in a long time and yet I could tell you how every system my people have built works down to a very fine level of detail. And it's because _we talk to each other!_ We don't just chill in the same slack channel and code independently, we all value each others brains and want each others inputs because we know it will improve our product and we value what perspectives others will bring.
Trying to learn via PR is a sad substitute for real collaboration and teamwork.
As for finding bugs, what happens if you miss them? Do they go out to production and potentially lose user data? Finding bugs in PR is a big problem. For a start it shows your automated tests aren't good enough, and secondly it shows the devs aren't checking their code works well enough.
If you do them at all, PRs should be a gate for checking whether the code meets the team's quality bar, not if it even works. The team should be able to deliver working code without them.
If PRs are a notable part of my architecture defense, I'm going to work on investing in the team instead of reviewing PRs.
Nice way to avoid the responsibility for any team fuckup: it's not me, I only teach them, they decide themselves.
Welcome to the tech industry. Its all about shifting responsibility in case something goes wrong. Thats the only reason companies use third party software in the first place, to have a scapegoat...
Which, incidentally, is a really enjoyable way to work.
Sneding a PR should not be for understanding or just for rubber stamping. It’s about getting someone to look at your approach and helping you find flaws or proposing ideas that could make it better.
When I review PR, the primary question is: For the stated problem, is the diff a good solution? Sometimes I don’t know enough about the problem, so I just try to see if the code has glaring mistakes (mispellings, styles,…) but those are just comments, not suggestions.
And in response I wrote a non-exhaustive checklist of things that a code review can look for:
- Does it functionally achieve what it sets out to (as per tacker issue or PR description)?
- Does it have extraneous code? Leftover debug prints, private API keys etc...
- Does it have any obvious defects? Memory leaks, un-handled edge cases, security flaws, obsolete API calls, etc...
- Could it be more understandable? Add/remove abstractions, better variable/method names, more/less functional etc...
- Is the style consistent with the codebase and/or style guidelines?
- Are there obvious performance improvements? Hashset instead of list, lazy evaluations, etc...
- Is it sufficiently well tested?
I think LLMs are okay at most of these, and worst at the first.
Is there already a pattern or code on in in the existing codebase that handles this functionality,
Do we really need net new code to achieve this functionality?
Can existing code be extended or abstracted to more cleanly implement this feature or functionality.
I don’t think I have ever even once seen an LLM solve a problem related to overengineering by simply removing the overengineering. They always choose to add more epicycles and further compound the complexity.
- Is the change architecturally right?
Particularly the latter LLMs seem still pretty useless at.
Although I do think that LLMs have made it much easier to justify writing low-value code which can make this more common now.
But AI has engendered a collapse in developers’ ability to actually do that. Those of us who are stuck on the vibecoding bandwagon have lost the comprehensive understanding of the systems under our care that we need to understand and explain the quality and maintenance implications of a change.
Worse, if you happen to lose your mind and suggest the initial development cost is anything more than ~zero, your friendly neighborhood Claude keener will publicly shame you for not having sufficient faith in the Glorious Agentic Future. Product leadership will then have no choice but to side with them, not necessarily because they agree, but because they, too, are aware that we’re still in the phase of the hype cycle where openly questioning said hype is a career-limiting move.
> Does this organization prioritize human learning?
That has been my primary motivator for code reviews. I want to teach and learn from others, especially given the decreasing levels of collaboration due to increased AI usage.
The sad truth is that all of my feedback just goes straight to agents. Maybe 10% is reacted to by a human, so I’m left wondering if there’s any value to a real review aside from poorly training robots to do my job, and further atrophying the abilities of my team members.
Previously the knowledge needed to discern that sort of thing would be disseminated through both design and code review sessions. But plan mode and AI code review largely put an end to that.
So we put our heads together and came up with some new policies about project management and how we use AI, and things have steadily getting better since then.
(Though, in fairness, the one guy who seems to actually enjoy getting paged after hours seems to be having less fun.)
If you will tell me precisely what it is that my machine cannot do, then I can always prompt my machine to do just that.
- John von Altman
Articulating “what humans can do, that AI cannot” is a mug’s game. If you specify it well enough, they just paste your text into their /goal prompt box and ralph loop their agent swarm until it produces something too exhausting to distinguish from doing the thing. If you don’t specify it well enough, then you’re just doing human-centric magical thinking to move the goalposts etc etc.Engineers played along with this farce because code review served valuable team collaboration, coordination and management functions, about which the author of the article is correct.
Understanding a system by reading code is harder than understanding a system by writing code.
If AI can generate code at 100X, 1000X, or 10000X human capacity (no ceiling here), and you are gated on code review as your mechanism for system understanding, then a team's productive output will barely increase.
If companies want to compete in the world of AI generated code, human code review has to go. The only question is, what replaces it?
Continuing to apply human code review to AI generated code is negligent, if you are shipping at AI generation speed, with that as your only gate, and no other systems and processes to validate correctness and limit risk.
On the engineering side we can adapt easily.
Code review was never about finding bugs. When we do code review the first thing we check is: "do the tests pass?" Then we look at the change and the test coverage added for it and ask: "does the test coverage adequately demonstrate the functionality of the code?" The we ask: "What is the scope and potential impact of this change?" "What is the deployment and rollback plan and how will we monitor and detect defects after deployment?"
Code review was never about the code. It made the lawyers happy and provided a vehicle for doing the things that actually make systems work.
You’re going a bit hand wavy for an answer by redefining the term into something that fits what you’re promoting.
and is this a solved problem? If not, then the bottleneck is right here, if it is solved, then yeah we shouldn't need anymore software engineers other than the elites
We have all the linters, tests, and AI writing code for us. I don’t need the left hand to tell the right hand it did a good job. I’m very certain my code runs when I push the PR.
What I need now is architectural, long-horizon and business perspective.
Lately I've seen some pretty glaringly obvious issues caught during the preliminary automated code review, and the issues seem to be coming from individuals who don't actually understand what the code is doing.
For the rest of us who are actually using the whole stack effectively, the automated code review is essentially a CI gate to protect the repo from the devs who don't know what they're doing.
But how does automated AI code review help, here? Doesn’t it just reinforce that they don’t need to look at it (or change their habits), because the AI review will catch the issues?
> What I need now is architectural, long-horizon and business perspective.
That's exactly what these tools are now good at. They have a huge gap when fixing these issues properly but they can spot these issues no problem
They also have shockingly weak ability to identify business acceptance criteria that are completely missing in the implementation or test coverage.
„The indent is wrong here“
„Comments should end with a period“
Because this kind of feedback is and was always easy.
Still not solved? Guess it was really about the commas and not the value delivered anyway, so do whatever you feel like.
What you actually need to do is have a manager lead assert that it’s happening in a top down way and just run the formatter with the defaults. Let people argue case-by-case on what to change after that.
Gpt-zero scores "human", and I've always found it to be a better judge
Something is missing in the new ai bot review paradigm we’ve all sleepwalked into.
I’ve been building Archme.io for this reason. PR reviews for the age of AI