PracticeWorking with AI: Delegation & Oversight
Two reviewers, three times the PRs
The artifact
From: Nadia Okonjo, VP Engineering To: you (CEO) Subject: Review capacity — need a decision this week
I need a call from you on this because all three options cost something and two of them cost money.
Where we are. We turned on agent-assisted development across both squads in March. Merged PRs went from ~40/week to ~130/week. Roughly 70% of those are now agent-authored — I can tell because our agent tooling stamps
Co-authored-by:on the commit, so the attribution is reliable.Everything on the delivery dashboard looks great. Cycle time is down 44%. Merge rate is up. We shipped the billing migration a month early.
The problem. Priya and Tomas are the only two people who review anything touching payments, auth, or the data pipeline. That was 12 PRs a week each in February. It's now 31 and 29. Priya told me last week she has "stopped pretending" to read the large ones end to end, and Tomas has started approving anything where our AI reviewer left no comments — which is about 60% of what reaches him.
I want to be clear that neither of them is doing anything wrong. They are doing the only thing available to them.
What I can't tell you. Whether quality has actually dropped. We had 3 customer-visible defects in Q1 and 4 in Q2, which is noise at this volume. Our incident count is flat. I genuinely do not know whether we are fine or whether we are accumulating something that surfaces in six months.
Option 1 — Automate the review. Buy Sentinel (the code review product Raj found). $4,100/mo at our seat count. Their pitch is that a human rubber-stamping machine-validated code is "a passenger pretending to drive," and their published benchmark shows their agent catching 62% of high-severity bugs when the reviewing model differs from the authoring model — vs 54% when it's the same model. Their recommendation is auto-approve on everything except a tier they call "material," with humans on that tier only.
Priya's reaction: "62% means two in five serious bugs walk. Their own number says they can't do this."
Raj's reaction: "Two in five caught is better than what we have now, which is Tomas approving because there's no comment."
Option 2 — Hold the line. Cap merges at what two people can actually review. Realistically that's ~60/week against 130 now. We'd be turning off capacity we already paid for and Marcus will ask why roadmap dates moved. I'd also expect attrition risk from the squads, who like shipping.
Option 3 — Build verification instead. Fund two engineers for a quarter on integration and property tests, staging soak, and better runtime alerting, so that "is it correct" is answered by machinery instead of by reading. Costs a quarter of two engineers and it pays back nothing this quarter. Tomas is the one who wants this and he's also the person I'd have to pull to do it.
What I need from you: pick one, or tell me the thing I'm not seeing. I have a board update on the 14th and "we're looking at it" is not going to survive contact with Marcus.
— N.
Your task
Decide. Make the call, then tell Nadia why. Say what you're accepting as the cost. If you think the choice as framed is wrong, say what you'd do instead — but still commit to a course of action, not to a further investigation.
Step 1 — make the call
What's your call for Nadia?
How confident are you?
The call isn't scored — there's no right option. What gets scored is what you write next.