Here’s a number that feels like it means something: “42% of new users activate in their first week.” It’s honest — nobody cooked it — and it’s close to useless, because it’s an average, and the average may be describing a person who doesn’t exist. (That figure is illustrative, not a real metric.)
An average is a good summary when your data piles up around a center — one hump, most people near the middle, a few stragglers on each side. Height works like that. Report the mean and you’ve told the story. But user behavior rarely piles up around a center. It splits. A large group gets the product fast and loves it; another large group bounces in the first thirty seconds; and “42%” is just the seesaw balance point between them — a value that lands in the empty valley between your two real groups, describing neither one.
This matters because you steer by these numbers, and steering by the average optimizes the mushy middle — the exact place your real users aren’t. Nudge the mean up a point and you might be dragging bouncers into lukewarm engagement while learning nothing about why the thriving group thrives. The group you most want to clone is invisible inside the summary.
Splits aren’t the only shape that breaks the mean, either. Usage data is usually heavy-tailed: a small band of power users doing ten or a hundred times what everyone else does, hauling the average up toward a level the typical user never touches. Practitioners sometimes talk about “smile-shaped” curves for exactly this reason — a pile of people at zero, a pile at the top, almost nobody at the average. “Sessions per user: 6” can be true of a product where the most common number of sessions is one.
⁂
There’s a sharper, meaner cousin of this, where the average doesn’t just blur the groups — it points the wrong way. That’s Simpson’s paradox. The textbook real case is UC Berkeley’s 1973 graduate admissions: in aggregate the numbers looked biased against women, yet within almost every individual department women were admitted at rates equal to or higher than men. Women had simply applied in larger numbers to the most competitive departments, and the aggregate reversed the truth of nearly every part. The same reversal hides quietly in product data — an overall retention dip that is actually improvement in every segment, masked by a shift in which segment is growing.
So the move is small and changes roadmaps: stop asking “what’s our activation rate” and start asking “what does the distribution look like, and who’s in each hump.” Draw the histogram, not the mean. Split the number by cohort, or by the one action your power users take that everyone else doesn’t. If it’s one clean hump, the average earned its keep. If it’s two, you don’t have a 42%-activating product — you have two products wearing one metric, and your job is to find the path the winning group took and build the whole thing to funnel people onto it.
Liked this? Get the next one in Working Theory.
Going weekly in August (it's in beta now). One genuinely interesting read on building, the brain, and the science most people missed.