Verification Is the New Grunt Work
I keep watching developers review AI-generated code the same way they reviewed their own code five years ago. Line by line. Diff by diff. Checking indentation, naming, whether the variable matches the one three files up. That review method was built for a world where a human typed every character and needed to check their own work. It was never built for the volume we're producing now, and it's cracking under the weight.
The grunt work moved#
Developers are now publicly saying they don't write code by hand anymore. Not 'I use autocomplete' — full stop, they don't type it. That statement would have sounded like malpractice three years ago. Now it's just a description of where the job went. The worker AI writes the PR. The developer's job is no longer producing the diff.
This isn't a productivity story about typing faster. It's a bottleneck story. The hard part of building software used to be writing correct code. That part is disappearing for a huge share of the work — CRUD endpoints, UI wiring, migrations, the stuff that made up most of a sprint. What's left, and what's actually hard now, is deciding what the system should be: architecture, organization, where the seams go, how the parts stay cheap to change. That's taste. It doesn't get automated away because it was never about typing in the first place.
How review used to work#
At byldr, review used to mean fine-grained history tracking — who touched this line, when, why, what it looked like three commits ago. That level of granularity made sense when code moved slowly and a single engineer's judgment covered the whole change. You could hold the entire diff in your head. You could ask 'why this line' and get a real answer from git blame and a Slack thread.
That model assumes a rate of change the AI-agent era doesn't have. When a worker agent turns out PRs touching a dozen files in the time it takes to read the first one, line-level archaeology stops being a review method and starts being a way to fall permanently behind. You can't blame-trace your way through a codebase that's regenerating faster than you can read it.
The verification loop#
What we've moved to instead doesn't try to prove the code correct by inspection. It tries to prove the application correct by use. A verification agent — Grok or something equivalent — walks every path a user could actually take through the app: signup, the weird edge case in checkout, the settings page nobody clicks, the error states. It's not reading the diff. It's using the product the way a very thorough, very patient user would, and logging what breaks.
- 01Worker AI ships a PR
A feature, a fix, a refactor — same as before, except no human typed it.
- 02Verification agent walks every path
Every route, every user flow, every state transition it can reach. Not a static scan — actual empirical traversal of the app as built.
- 03Issues get logged to a task tracker
Any tracker works. The only requirement is the ability to add a task, prioritize it, and read it back.
- 04Developer fixes
Judgment goes here, not in the diff review. Is this a real bug, a design decision, or noise?
- 05Re-walk
Repeat until the agent can't find anything new. Behavior is guaranteed empirically, not proven functionally.
The task tracker is the part people overthink. It doesn't need to be Linear or Jira with custom workflows. It needs three operations: add, prioritize, read back. That's the whole interface between the verification agent and the human closing the loop. Anything more elaborate is ceremony that slows down a process whose entire value is speed.
Why this isn't just faster QA#
This is a different epistemics, not a faster version of the old one. Line-by-line review and unit tests try to prove code correct by reasoning about it — you read the logic, you construct the test cases you think matter, you convince yourself it's right. Empirical path-walking doesn't try to convince anyone of anything. It runs the actual paths and reports the actual failures. It finds the bugs your unit tests didn't think to check for, because it isn't limited to the scenarios a human anticipated when writing the test.
The tradeoff is real: you lose the deep understanding that comes from reading every line. You gain coverage a human reviewer was never going to achieve at this volume anyway. I'd rather have a system that's been walked through every path and had the breakages fixed than a system where a tired engineer skimmed a 40-file PR at 6pm and approved it because the tests were green.
- Architecture — where the seams go, what's a service boundary and what isn't
- Organization — how the codebase stays legible as it grows past what one person can read
- Efficiency — which parts of the system are worth being clever about, and which aren't
- Judgment on what the verification loop surfaces — bug, tradeoff, or expected behavior
That's the short list of what's left for a human. Notice none of it is typing, and none of it is reading every line someone else's agent wrote. It's all taste — decisions about shape and tradeoff that an agent can't make for you because they require wanting something specific for the system, not just producing correct output.
Stop reviewing AI code like it's handwritten code.
Direct an agent to walk every user path through the app end to end. Log every issue into a tracker that can add, prioritize, and read back tasks. Fix. Re-walk. Repeat until the agent stops finding anything new. That loop replaces line-by-line inspection because the old method assumed a volume of code that no longer exists — there's too much of it now, and it was never the part that needed your judgment most. Spend that judgment on architecture, organization, and efficiency. That's the job now. Everything else is the new grunt work, and it belongs to the agent that's willing to click through every path in your app a thousand times without getting bored.
If your team is still drowning in AI-generated PRs under an old review model, this is the fix — and it's smaller than it sounds.
Book a call →