Product Judgment When the Draft Is Free
Chapter 7 · 16 min read
On this page
Start here
A Slack thread, Tuesday afternoon.
Eight months ago the team rejected a feature — bulk supplier onboarding via CSV import. The rejection memo is still linked in the thread. It rejected the feature on demand: three customers had asked, all three had fewer than fifty suppliers, and the manual flow took them about twenty minutes a year.
Today an engineer has built a working prototype in two days with a coding agent, and posted a demo video. It works. It is genuinely good.
The thread is enthusiastic. Someone writes: this used to be a quarter of work and now it's a week — surely we just ship it?
Everyone can see the prototype. Nobody has reopened the memo.
And here is what makes this hard rather than obvious: refusing free value looks like stubbornness. The prototype exists. The person who built it is pleased and did good work. Saying no costs you something socially, and the argument for yes — it's basically free now — feels like new information.
It is new information. It is just information about a different question than the one the memo answered.
So: when the cost of producing something collapses, what actually changed about whether to do it?
Core concepts
Which risk actually moved
Chapter one gave you four risks: value, usability, feasibility, business viability. This chapter is largely about what happens to that set when one of them gets dramatically cheaper to resolve.
The CSV import was rejected on value — three customers, twenty minutes a year. The prototype is evidence about feasibility. Feasibility moved. Value did not move at all. Nothing about the customer situation changed between the memo and Tuesday.
How to use it. When a previously-rejected idea returns with a new argument, ask one question before anything else: which of the four risks does the new information bear on, and which risk did we reject it on? If those are different risks, the new information does not reach the original decision. It may still be worth revisiting for other reasons, but it has not answered the objection.
And then price what is still not free. The build got cheap. These did not: surface area in the product, support load, documentation, the next migration that has to carry this feature along, and the attention of whoever maintains it. A feature that costs a week to build and then exists for six years is mostly not a build cost.
How to spot it. The argument's form. "We can build this in a week now" is a statement about cost. If it is being used to answer an objection about demand, the sentence has quietly changed the subject and everyone in the thread is agreeing with it.
Fluency stopped being evidence
For about forty years, a well-structured strategy document was weak evidence that thinking had occurred. Not proof — but producing a polished twelve-page analysis with a clean risks table used to require enough work that the polish correlated with the substance.
That correlation broke. The reading habit did not.
So the practical problem: you now receive documents that look exactly like the good version. Appropriate headings. A risks table. Phased plan. Confident, well-calibrated-sounding prose. And in which no claim is sourced and no number is traceable.
How to use it. Read for traceability rather than structure. Concretely:
- Identify the two or three claims the conclusion actually rests on. Most documents have surprisingly few — the rest is context and elaboration.
- Ask where each one came from, and follow it.
- Mark everything else explicitly as unverified rather than pretending you checked it.
That third step is the one people skip, and skipping it is what turns a fast read into a false sense of having reviewed something.
Why you cannot just check everything. Generating a plausible claim is now nearly free; checking one costs roughly what it always did. Where those two costs diverge that sharply, the volume of unverified plausible material grows without bound and your attention becomes the binding constraint. Attempting uniform scrutiny means either burning a day or doing it superficially and calling the document checked — and the second is worse than not checking, because it produces confidence.
How to spot the trap. Fabrications are shaped — by the generation process, not by anyone's intent — to pass exactly the cheap checks a busy reader runs. A plausible source. A real-sounding figure. A link that resolves. A link resolving is not verification; the link resolving to a document that makes the claim is.
How to spot it in your own reading. Notice when you have formed a favourable impression of a document and cannot name a specific thing in it you checked.
The verification cost decides what to delegate
Whether to hand a task to a model looks like a question about capability. It is mostly a question about arithmetic, with three terms:
- how long the task takes a human,
- how often the model gets it right,
- how long it takes to check the output.
The third term is the one people omit, and it is usually the one that decides.
How to use it. Estimate all three. Delegation wins when the total expected cost — model attempt, plus verification, plus rework when it is wrong — beats doing it directly. Two consequences that are not obvious:
- A task that is fast to do and slow to verify should stay human, even at high model success rates. Much of discovery synthesis is this.
- A task that is slow to do and fast to verify should be delegated even at mediocre success rates.
And the leverage point: verification speed is a variable you can change by design. Asking for output in a checkable form — with sources attached, with the underlying quotes, structured so the load-bearing claims are separable — moves a task from the first category to the second. That is a decision you make when you delegate, not a property of the task.
How to spot the gap. A workflow proposal with time-saved estimates and no verification cost anywhere. The saving is computed against doing the work; the real comparison is against doing the work and checking it.
More output at constant outcome
Chapter one's distinction returns, because it is where the whole thing lands.
Teams adopting AI tooling frequently ship substantially more while their outcomes stay flat. This is not mysterious once you have the vocabulary: the constraint was never how fast the team could produce.
How to use it. When output rises and outcomes do not, resist the two easy readings — that the tools are not working, or that the outcomes will follow with time. The likelier reading is that the team is solving the wrong problem faster, and that production speed was never the bottleneck.
How to spot it. Velocity metrics up, release counts up, behavioural metrics flat or unmeasured. It looks like success in every artifact the team produces about itself.
Back to the Slack thread
Which risk moved: feasibility. Which risk it was rejected on: value. The prototype does not reach the original decision.
What is still not free: the import path becomes a permanent surface, someone supports it, and every future change to supplier data has to carry it.
The honest reply in the thread is not "no." It is: the memo rejected this on demand, and demand hasn't changed. What would change my mind is evidence that supplier counts have grown — has anyone looked? That is answerable in an afternoon, and it treats the prototype as what it is: good work that answered a different question.
Worked examples
The competitive analysis nobody had read
A head of product receives a fourteen-page competitive analysis before a strategy offsite. It is organised, well written, and has a summary table of nine competitors with positioning, pricing, and recent product moves.
She has forty minutes, not a day.
She does not read it front to back. She identifies what the recommendation rests on, which turns out to be two claims: that the two largest competitors have both shipped a specific capability in the last six months, and that a named analyst firm projects the segment growing at 28% annually.
Claim one: she checks both changelogs. One is right. The other shipped something adjacent and narrower, and the difference matters because the recommendation depends on the market having converged.
Claim two: the analyst report exists and the firm is real. The 28% figure is not in it. The report gives a different figure, for a differently-defined segment.
Everything else in the document is probably fine. She marks it unverified and says so at the offsite, in one sentence: two of the load-bearing claims are wrong or unsourced; the rest I haven't checked and neither has anyone else.
The point is not that she caught two errors. It is the shape of the twenty minutes: she did not audit fourteen pages, she found the two loops the argument hangs from and pulled on them. That is a repeatable procedure, and it is the only one that fits inside a real calendar.
The workflow that saved time and cost more
An engineering manager proposes using a coding agent for a category of routine migration work. The case: these tasks average four hours each, the agent completes them in about twenty minutes, they do roughly fifteen a month. Projected saving: about 45 hours a month.
The number is honest and the agent genuinely does the work.
The reviewer asks the missing question: how long does it take to verify one?
Nobody has measured it. When they do, it is about ninety minutes — these migrations touch data in ways that are not obviously correct from reading the diff, and the failure mode is silent. Success rate is around 80%, and the 20% need most of the original four hours to untangle because the agent's version has to be understood before it can be replaced.
Recompute: 20 minutes plus 90 minutes verification, plus 20% of four hours rework, comes to roughly 2.7 hours against four. Still a saving — about a third, not 92% — and the saving is real.
But the shape of the work changed. The engineers now spend most of that time verifying rather than doing, and verification of someone else's work is both more tiring and worse at building the understanding that makes the next migration faster. That cost does not appear in any month's numbers and shows up two quarters later as a team that is slower at the thing it stopped practising.
The decision they make is the interesting one: adopt it, and change the output format so verification gets cheaper. The agent is required to produce a before/after data comparison alongside the diff, which drops verification to about 25 minutes. That single change is worth more than the original delegation, and it is available only because someone asked the third question.
The team that shipped twice as much
A 40-person product org adopts AI tooling broadly. Two quarters later: pull requests per engineer roughly doubled, cycle time down 40%, release frequency up. Everyone can feel it.
Outcomes are flat. Activation, retention, expansion revenue — all within noise of where they were.
The two comfortable explanations arrive immediately. Adoption takes time, the outcomes will follow. And we're measuring the wrong things.
The third explanation is the uncomfortable one and it is usually right: the constraint was never production. The team was already shipping more than it could evaluate, and the bottleneck was — and still is — knowing which things are worth building. Doubling throughput against that bottleneck produces exactly this: more output, same outcome, and a larger product to maintain.
What makes it hard to see from the inside is that every artifact the team produces about itself shows improvement. Velocity is up in the dashboard, the retro is positive, individual engineers report feeling more effective, and they are not wrong. Nothing in the team's own reporting surfaces the problem, because the problem is the absence of a measurement nobody was making before either.
The diagnostic is one question asked of the last quarter's work: for how many of these shipped items can we state the behaviour it was meant to change and what actually happened? If the answer is "a few," the throughput was never the constraint.
Case studies
A team shipping substantially more at flat outcomes
An invented composite, built from the pattern rather than any real company.
The situation. A B2B product company, 38 in product and engineering, £11m annual revenue. Twelve months ago it adopted AI coding assistance broadly, plus model-assisted discovery synthesis and a generated-first approach to PRDs.
Twelve months on:
- Pull requests per engineer per week: 4.1 → 8.3
- Median cycle time: 6.2 days → 3.7
- Features shipped per quarter: 14 → 26
- Discovery interviews synthesised per quarter: 20 → 65
- Activation rate: 31% → 32%
- Net revenue retention: 104% → 103%
- Quarterly expansion revenue: flat
Engineering satisfaction is up. The board deck leads with velocity. The CEO has asked why revenue has not followed, and the honest answer is that nobody knows.
Diagnosis one: which constraint was relieved. Production doubled. Before the adoption, was production the binding constraint? The team's own record says no — the previous year's retro flagged that they shipped fourteen features and could articulate a target behaviour for four of them. The constraint was deciding what to build, and that is exactly what did not get faster.
Diagnosis two: what happened to discovery. Interviews synthesised went from 20 to 65 per quarter, which sounds like the decision-quality constraint being relieved too. It is not. Interviews conducted went from 20 to 24. The rest of the increase is re-synthesis of existing material, plus sales-call recordings processed in bulk.
Meanwhile: how many people on the product team spoke directly to a named customer last month? Two. Twelve months ago it was nine.
The synthesis volume tripled while the contact that generated the signal fell by three-quarters. That is chapter two's premature-automation pattern operating at org scale, and it is invisible in every metric the team tracks — the discovery number went up.
Diagnosis three: what happened to the PRDs. Generated-first PRDs are longer, better structured, and arrive faster. Sampling ten from the last quarter: nine have a risks table, and in eight of them every row is a delivery risk. Six state a success metric; two state a baseline. None names a disconfirming observation.
The previous year's handwritten PRDs were worse documents and better decisions — shorter, patchier, and more likely to contain a sentence like "we're not sure anyone wants this, here's the smallest way to find out." That sentence does not survive a generation step, because it is not what a confident PRD sounds like.
What is actually happening. Three things, compounding:
- Throughput doubled against a bottleneck that was never throughput.
- The discovery signal that would have told them what to build got thinner while its reported volume tripled.
- The documents that used to carry uncertainty now arrive fluent, so the uncertainty is no longer visible at the point of decision.
Each is defensible on its own. Together they describe an organisation that got much faster at converting weakly-grounded decisions into shipped product.
What a good response looks like. Not "stop using the tools" — the cycle-time improvement is real and worth having.
Restore the contact. A floor on direct customer conversations per product person per month, tracked as its own number, unmergeable with synthesis volume.
Make uncertainty survive the draft. Require in every PRD, in the author's own words: which of the four risks this addresses, which it assumes, and what would change our mind. Generated content is fine everywhere else. These three are not delegated, and the reason is not process hygiene — it is that these are the judgment half, and the whole failure is that the judgment half stopped being written down.
Measure the thing that did not move. For each shipped feature, the target behaviour and the observed change. Expect this to be uncomfortable for two quarters, because the honest answer for most of the 26 will be "we didn't measure."
Consider shipping less. The most counterintuitive move and possibly the highest value: cap features per quarter below capacity, and spend the difference on evaluating what shipped. This will be unpopular, and the argument against it — that capacity is being wasted — is the productivity paradox restated as a plan.
Where this could still go wrong. The contact floor becomes a compliance metric and people book calls to hit the number. The "what would change my mind" line gets generated too, which is a real risk and only partly mitigated by requiring it in the author's own words. And the reduced-shipping experiment gets cancelled after one quarter, because the cost is immediate and visible while the benefit is slow and shared.
What the case teaches. The productivity paradox is not a claim that AI tooling does not work. It is a claim about which constraint it relieves. If a team's binding constraint is knowing what to build, doubling how fast it builds converts the same judgment into more product — which is worse than the previous state, not better, because the surface area is permanent and the judgment did not improve.
What great operators do
Demand the raw material behind any synthesis. Summarisation removes exactly the specificity that made evidence evidence, and it now happens automatically. Tell: the write-up links to or quotes actual conversations and behaviours. A findings document with no traceable underlying observation is an opinion citing itself.
Keep human contact with users unautomated the longest. Direct contact is the source of every non-obvious insight in discovery, and it is the easiest thing to automate away — which is why teams do it first and then wonder why they stopped learning. Tell: someone on the team spoke to a named customer this week and can recount the conversation. Summaries of summaries do not count.
Ask "what would change my mind?" and answer it in writing. The cheapest defence against reading every result as support, and the line least likely to survive a generated draft. Tell: the document names the observation that would reverse the decision. Its absence beside a confident recommendation is the signature of a conclusion that preceded its analysis.
Common failure patterns
Fluency mistaken for rigour
What it looks like. A well-structured, confidently-worded document — appropriate headings, a risks table, a phased plan — in which no claim is sourced and no number is traceable. It reads like the good version of this document.
Why smart people do it. Format was a real signal for decades: producing a polished strategy doc used to require the thinking that justified it. That correlation broke, but the reading habit didn't, and reading habits are not consciously held beliefs you can simply update.
The correction. Read for traceability rather than structure. Pick the three load-bearing claims and ask where each came from. Treat "it reads well" as carrying no information about whether the analysis happened.
Cheap-to-build therefore worth building
What it looks like. A previously-rejected idea returning with "we can build this in a week now" as the new argument. The value case is unchanged; only the cost estimate moved.
Why smart people do it. The cost drop is real and dramatic, and it feels like new information. It is — about feasibility, which was rarely the binding risk. It also arrives attached to a working prototype and an enthusiastic colleague, which makes disagreeing socially expensive.
The correction. Ask which of the four risks the new capability actually changed. If the idea failed on value, it still fails on value, and it now also carries maintenance, support, and surface-area cost.
Output mistaken for progress
What it looks like. A retro or update listing shipped work as accomplishment. Velocity up, releases up, the team visibly productive. No behavioural metric appears, or one appears without a baseline.
Why smart people do it. Output is measurable, attributable, and immediate; outcomes are noisy, lagging, and shared. Under pressure the measurable thing wins. With AI tooling the output number rises sharply, which makes the pattern both more tempting and more expensive.
The correction. Pair every shipped item with the behaviour it targeted and the observed change. Where none was measured, say so — an unmeasured launch is an unevaluated bet, not a success.
Make the call
The launch memo that reads perfectly is the checkpoint, and it is deliberately the same scenario chapter one used. That is the course closing a loop rather than an oversight: the memo you critiqued with a four-risk checklist is now something you can read for traceability, and the two readings find different things.
Specifically: the memo notes it was drafted with an AI assistant and then edited. Chapter one's reading finds a risks table containing no product risks. This chapter's reading asks which of its claims are load-bearing and where each came from — the win-rate assertion, the market-context paragraph, the beta enthusiasm quotes — and finds that its polish carries no information about whether any of the analysis happened.
If you did it in chapter one, do it again and compare. Noticing what you missed the first time is more useful than a fresh scenario.
Also planned: the Slack thread from the start of this chapter — the rejected feature returning because a prototype took two days — as a scenario in its own right. It is in production and does not exist yet.
Where to go next. This chapter's concepts are the entry point to Working with AI, which takes the delegation question, the verification asymmetry, and the craft of judging machine-produced work as its whole subject rather than as one chapter. Its course is not written yet; its scenarios are live now — start with Automate it or keep it human, which is the delegation equation as a decision you have to make and defend.