Reading the Numbers You're Given
Chapter 3 · 13 min read
On this page
Start here
The slide says: Activation up 30%.
It is the third slide in a monthly review and nobody stops on it. Why would they? Thirty percent is a big number, it is pointed the right way, and the person presenting it did the work.
Two months later the company has hired two more people onto that team and shifted the quarter's roadmap toward the thing that produced the 30%. Revenue has not moved.
Here is what nobody asked in the room, and none of it is a hard question:
Thirty percent of what — a rise from 10% to 13%, or from 10% to 40%? Activation meaning what, exactly; who chose that definition and when? Compared to what — last month, last quarter, or a comparable company at this stage? And thirty percent of which users: everyone who signed up, or everyone who finished signing up, which is a smaller and much friendlier group.
The number is almost certainly true. Somebody computed it correctly. The problem is that a true number can be uninterpretable, and an uninterpretable number in a room full of busy people gets interpreted anyway.
So: what do you have to know about a number before it can support a decision?
Core concepts
Compared to what
"Our month-two retention is 40%."
Sit with that for a second and notice you cannot tell whether it is good. Not because you lack context about the company — because the sentence is structurally incapable of being good or bad. Forty percent month-two retention is excellent for a consumer mobile app, unremarkable for a B2B tool people use for work, and catastrophic for enterprise software with an annual contract.
A metric is uninterpretable without its reference class: the set of comparable situations you judge it against. Establishing the class is not context-setting that happens after the analysis. It is the analysis.
How to use it. Before arguing about what to do about a number, establish what the number should be. Four questions, in order:
- Compared to what? Which products, at what stage, in what category?
- At what stage? A seed-stage number and a Series-C number of the same size mean different things.
- In what category? The category determines the whole distribution.
- Measured how? This one gets its own section below, because it is where most comparisons quietly break.
How to spot the gap. Two failure modes are equally common and they look opposite: celebrating a mediocre number because it went up, and panicking about a perfectly normal number because it sounds low. Both come from the same place — arguing about the change without knowing the level.
One thing that has changed. Ask a model for a benchmark and you will get one, fluently, whether or not one exists. A fabricated benchmark is worse than no benchmark, because no benchmark leaves you searching and a fabricated one terminates the inquiry. A benchmark needs a citable source with a stated methodology — and if the source's definition differs from yours, you do not have a comparison, you have two numbers.
The general machinery here — base rates, the outside view, why the vivid specifics of your case should move you less than you want — belongs to critical thinking as a discipline and is treated properly in that domain's course. What matters for product work is the operator-facing habit: find out what good looks like before the meeting, not during it.
Measured how
Two teams at the same company report engagement. One says it is up, one says it is flat. Both are honest and both are right, and the argument that follows will last forty minutes and resolve nothing.
Between a concept — engagement, productivity, activation, quality — and a number lies a definition somebody chose. That choice, not the concept, is what the number reports.
How to use it. Ask measured how before asking is it good. Two failure modes to watch for:
- The metric doesn't capture the concept. Session count as "engagement" also rises when the product becomes confusing and people have to come back.
- Two compared numbers use different definitions. This is the single most common way a comparison silently inverts. It requires nobody to be careless — just two people to have made reasonable choices on different days.
And where a metric is also a target, expect the definition to have been optimised against. Not fraudulently; by ordinary human attention flowing to the thing being measured.
How to spot it. Ask who wrote the definition and when. If the answer takes more than a minute to find, the number has been travelling on trust.
One thing that has changed. This is acute for anything measuring AI. "Faster" in an internal AI-productivity metric almost always means output, not outcome — and you already know from chapter one that those are different things.
The shape of the curve
Signups are up 25% month over month. The chart is beautiful. Everyone can see it from the back of the room.
Now ask for the same picture drawn differently: take everyone who joined in January and plot what fraction were still active in February, March, April. Then do the same for February's group, and March's. That is a cohort — a group bucketed by when they joined — and its curve tells you something the signup chart cannot.
The clearest evidence that a product has found its market is a retention curve that flattens. A cohort whose usage stops decaying and settles at some level means a group of people found durable value. Acquisition growth cannot tell you this, because acquisition can be bought.
How to use it. Look for the curve, not the total. Rising signups with a retention curve that decays toward zero describes a leaking bucket being filled faster — and the faster you fill it, the more it costs. Also check what "retained" means in this specific product; the definition is where this metric is most often quietly gamed, which is the previous section arriving again in a new outfit.
How to spot it. Aggregate monthly actives, with no cohort chart anywhere. That single number hides decay behind acquisition and can rise for a year while the product gets worse.
One thing that has changed. Retention is harder to read for AI-native products, where novelty drives a large early usage spike that decays regardless of fit. Distinguishing habit from curiosity needs a longer window than teams want to wait, and a short-window dashboard will flatter almost any AI feature launch.
Back to the 30%
Run the three concepts over that slide and it dissolves.
Measured how: activation is defined by someone, somewhere, and if the definition changed in the last quarter the comparison is void. Compared to what: 30% is a change, not a level, and nobody in the room knows the level. What shape: even a real activation improvement tells you nothing about whether those users stayed — which is the question the hiring decision actually turned on.
None of that requires statistics. It requires refusing to accept a number until it arrives with the three things that make it mean something.
Worked examples
The number that survived every question but one
Ines is a founder reviewing a proposal to double spend on a paid acquisition channel. The case: cost per signup on this channel is £14, against £31 blended. The channel is more than twice as efficient. The recommendation is to move budget.
She runs the questions.
Compared to what? £14 is genuinely good for the category; she checks a practitioner benchmark and finds comparable products at £20–40. It survives.
Measured how? Cost per signup means marketing spend divided by accounts created, and both are pulled from systems she trusts. It survives.
At what stage? The channel has been running eleven weeks, past the noisy start. It survives.
What shape? Nobody has plotted it. She asks for cohort retention split by channel, which takes a day.
The £14 cohort retains at less than half the rate of everyone else at week four. Cheap signups, and the cheapness is the tell: the channel is efficient at acquiring people for whom the product is a marginal fit, which is exactly what you would expect a cheaper channel to be.
Corrected for retention, the effective cost per retained user on that channel is worse than blended. The recommendation inverts.
What is worth noticing: the proposal was not sloppy. It survived three of four questions and the number was real throughout. The failure was choosing a denominator — signups — that stops measuring one step before the thing anyone cared about.
The comparison that inverted
Two product teams present quarterly retention. Team A reports 62%. Team B reports 47%. The obvious reading is that A is doing better, and in the following quarter B loses a headcount slot to A.
Six weeks later someone chasing an unrelated bug finds the definitions.
Team A counts a user as retained if they logged in at any point during the month. Team B counts a user as retained if they completed a core action during the month. Neither definition is wrong. They were written by different people, eighteen months apart, for different dashboards.
Recomputed on Team B's stricter definition, Team A is at 39%.
Nobody lied. Nobody was even careless — each definition was reasonable in its original context. The failure was in the reading: two numbers were placed side by side, which implies a comparison, and nobody checked whether a comparison existed.
The habit that would have caught it costs ten seconds: when two numbers of the same name appear in one document, ask whether they were computed the same way, before reasoning about the difference between them.
The dashboard that was rising and wrong
A B2C app has monthly actives up and to the right for five straight quarters. The board deck leads with it. Growth spend rises each quarter on the strength of it.
Plot cohorts and the picture reverses. Every cohort decays to near zero by month four. It has done so from the beginning. The aggregate rises because each new cohort is larger than the decay of all the previous ones — which is arithmetic, not product/market fit.
The company is not growing. It is paying, every month, for a fresh set of people to briefly use its product and leave, and the number that describes this is going up.
The reason this pattern is so persistent is worth naming, because it is not stupidity. Acquisition responds to effort and money on a weekly timescale. Retention responds to product truth on a monthly one. The fast-moving number captures attention, the aggregate genuinely is going up, and the slow number requires someone to go and build a chart nobody asked for.
Scaling acquisition into a decaying curve is the most reliably expensive mistake in this domain. It is expensive precisely because it feels like success while it happens.
What great operators do
Establish the benchmark before arguing about the number. Without a reference class, metric debates become debates about optimism. Five minutes finding what good looks like reframes the whole discussion. Tell: the artifact says "comparable products at this stage see X, sourced from Y," and states the measurement definition. Absent: "retention improved 12%" with no base and no definition.
Watch the retention curve's shape, not its level. A flattening curve means some group found durable value; a decaying one means the product is a leaking bucket regardless of how good this month's number looks. Tell: cohort curves plotted over time. Absent: aggregate monthly actives, which hides decay behind acquisition.
Common failure patterns
Naked metric
What it looks like. A number offered as evidence with no reference class, no baseline, no definition, and no denominator. "Activation improved 30%." "Retention is 40%." "NPS is 42."
Why smart people do it. The number is real and the improvement may well be real. Supplying the context is extra work that can only make the number look worse, and there is rarely anyone in the room whose job is to ask for it.
The correction. Demand four things: compared to what, at what stage, in what category, measured how. Most naked metrics do not survive the fourth question.
Leaking bucket growth
What it looks like. Rising signups and aggregate actives presented as product/market fit, with no cohort retention curve anywhere. Growth spend is being increased on the strength of it.
Why smart people do it. Acquisition responds to money on a weekly timescale; retention responds to product truth on a monthly one. The fast number captures attention, and the aggregate genuinely is going up — nobody has to be fooled for this to work.
The correction. Plot cohorts and look for flattening before scaling spend. If the curve has not flattened, additional acquisition spend is buying a larger version of the same leak.
Make the call
Two scenarios, and they test different halves of this chapter.
Two roadmap bets, one team turns entirely on reading a cohort table. You get two options, both arguable, and the table is the thing that decides between them — but only if you read the shape rather than the headline. The common failure is arguing the strategy well and never opening the numbers.
Traffic but no signups is the optional second rep, from the Critical Thinking track. It hands you a diagnosis built on metrics and asks what is wrong with it, which is this chapter's habit pointed at someone else's reasoning rather than at a dashboard.
Honest difficulty note. The first is a 3 — the hardest in the library right now — and it is hard in a specific way: a confident, well-argued answer built on the headline number will read as strong right up until the cohort table contradicts it. The second is a 2 and is the gentler entry point of the two.