TheProduct Playbook

Step 3 — Hypothesis

In Step 2 you decided the problem was worth solving. Now you turn that into something you can actually test. A belief, written down, with the measures that tell you whether you're right.

That last part is where most teams fall apart. They have a gut feeling it'll work, and they call that a hypothesis. It isn't. A hypothesis you can't measure is just a hope with better posture.

A hypothesis has a shape

There's a difference between an assumption and a hypothesis, and the difference matters. An assumption is a belief you haven't shaped into anything you could test. A hypothesis always breaks down into the same form: if you do something for a particular user, they will respond a certain way. Everything else is just an assumption wearing a lab coat.

So write it in that shape, on purpose:

If I/we [do something] for [user], they will [respond a certain way].

That's it. The format forces you to name the user, the change you're making, and the response you expect. If you can't fill in all three, you don't understand the opportunity well enough to test it yet. Go back and define the problem before you try to test a belief about it.

Run our late-paying customer through it. By now you suspect the lateness isn't unwillingness to pay, it's friction at the moment of paying. The invoice lands, paying it is a hassle, so it sits. Write that as a hypothesis:

If we give customers a one-click payment link in the invoice email, they will pay faster, because the friction at pay-time is what's slowing them down, not an unwillingness to pay.

Now you have something with a shape: a user, a change, an expected response, and the belief about why. That because isn't one of the three blanks, but add it anyway: it's the belief your measures will actually test. And now you have a thing you can prove wrong, which is exactly what you want.

You can't measure it? Then it's not a hypothesis yet

A belief with no measure attached is just an opinion you've decided to like.

So before you go any further, name the metrics that would tell you the hypothesis held up. And this is not something you do alone at your desk. Work with a multi-functional team to develop the measures you'll use to validate it. Sales, support, finance, engineering. This process should not be done in a silo, because the person closest to the problem usually knows which number actually means something and which one just looks good.

For the late-payment hypothesis, the measures might be:

  • Current days-to-payment, measured before anything changes. That's your baseline.
  • A quantitative and qualitative survey of how customers experience paying us today, so you know the why behind the lateness, not just the size of it.
  • Days-to-payment after the one-click link goes live, measured against that baseline and against a group that didn't get the link.

Notice the first one is a baseline. You can't tell whether a number moved if you never established what normal was. I learned that working with a large beverage company on customer churn (the rate at which customers leave or cancel on you). It took six or seven months just to set a baseline, because people obviously drink more around certain holidays and times of year, and we had no idea what normal churn even looked like until we'd accounted for all of that. We had to create the number first, and only then could we look for ways to improve it. Skip the baseline and the seasonality alone will lie to you.

Vanity metrics: the number that feels good and proves nothing

Here's where teams reliably fool themselves. They pick a metric that goes up and to the right, watch it climb, and feel great. The problem is the metric doesn't actually prove the thing they care about.

Vanity metrics are ones that make you feel good on the inside but don't really prove the value of your experiment or company.

I once worked with a team whose training scored a perfect five out of five. Every customer who went through it rated it five stars. By that one number, the training was excellent. Except it wasn't. Out in the field, where customers were actually trying to do the thing the training was supposed to teach them, they kept hitting problems. The training everyone loved was, in reality, not working.

So what happened? Customers rated it five stars right after the training, before they'd ever applied any of it. They had three days getting hyped up and doing fun stuff, and at the end someone asked "did you like the training?" Of course they said yes. The rating measured the experience of the training, not the value of it. The real value, or the lack of it, only showed up days later when they got out there and tried to use what they'd been taught. Five stars told us everyone had a nice time. It told us nothing about whether the training worked.

That's a vanity metric. It felt like proof and it was noise.

Here's how you spot one. A metric is probably vanity if it does any of these:

  • It counts activity, not outcomes.
  • It's a crude proxy for what you actually care about.
  • It skips nuance and context.
  • It's often misleading.
  • It doesn't help you improve your product or business in any meaningful way.

You've probably seen versions of these. "We had 100,000 app downloads." Cool. Did we make any sales? No? Then who cares. "We shipped the API" (the connection other companies' software plugs into). Okay, how many people use it? Nobody? Then you shipped a useless piece of junk and threw a party for it. Downloads, page views, registered users, an API nobody calls. All of them climb, and none of them, on their own, tell you the business got better.

And the worst version isn't tracking a vanity metric at all. It's tracking a real one and doing nothing with it. A real metric you watch but never act on isn't vanity, it's worse: it's a number that could have told you something, and you wasted it. Act on the number and the trade is simple:

I'd rather feel like shit over a KPI than feel great over a vanity metric. At least the bad feeling comes with the truth attached.

One number doesn't tell the full story

There's a deeper trap underneath all of this, the one I'd put on a sticky note: one number doesn't tell you the full story. A big tech company I worked at drilled this into us. You don't just need to know the number, you need to know whether the number is good. Take cost to acquire a customer, the money you spend to win one new customer (people shorten it to CAC). Pick a number: a $60 CAC is awful for a $1 product and amazing for a car. Same with growth. "We grew 30% year over year" sounds great until a competitor grew 40%. A number on its own can't tell you whether it's good news or bad. A number with its context can.

That's exactly what the next test is built to catch.

The metric test: three questions

Before you commit to any measure, run it through three questions. If it can't answer all three, throw it out.

What business decision can you make with this metric? If the number moving wouldn't change a single thing you do, it's not a metric, it's a vanity stat. A real measure points at a decision. "We grew 30% year over year" can't point at one until you know the competitor grew 40%. A number without its context dies here too.

Can you intentionally reproduce the result? If you can't explain how you'd make the number move again on purpose, you didn't learn anything. You got lucky, or you got noise.

Is the data a real reflection of the truth? This is the five-star training question. Does the number measure the thing you actually care about, or does it measure something adjacent that happens to feel like it? A survey taken before anyone applied the training reflected the mood in the room, not the truth about the training.

Run all three on every measure of success before you trust it. Most vanity metrics die at question one.

Make it provable: the control group

The cleanest way to know whether your change caused the result is to compare it against the same situation without the change. That comparison group, the one that doesn't get the change, is a control group, and it's what lets you say your change caused the result instead of guessing. It's the difference between "the number went up" and "the number went up because of what we did."

Here's the mechanic, using a teaching example of a hiring manager named Jill. (Jill's made up, and so are her numbers, the method isn't.) Jill has a hypothesis: if hiring managers schedule interviews over text instead of email, candidates will respond quicker and interviews will get booked faster. So she doesn't just switch everyone to text and hope. She splits candidates into two groups. One schedules over text, the other keeps using email. Same role, same process, two channels. When the text group books interviews say 42% faster, she knows it was the channel, because the only thing different between the two groups was the channel.

That's the whole reason a control group earns its keep. Without it, Jill changes the channel, hiring speeds up, and she can't tell you whether it was the text messages or just a less busy hiring month. With it, she can.

For the late-payment test, the control group is obvious: a set of customers who keep getting the normal invoice while another set gets the one-click link. If days-to-payment drops for the link group and holds steady for the control, you've got your answer. If both drop, something else changed and your link gets no credit for it.

I'll show you how to actually run these experiments cheaply in the next two steps. The point here is to design them so the result can be trusted: a hypothesis with a clear shape, real measures that survive the three questions, and a control to prove cause instead of coincidence.

Then, and only then

Now you have a testable belief. Not a hope. A specific claim about a specific user, with the measures that will tell you, honestly, whether it held up, and a way to prove your change caused the result and not the calendar.

That's Step 3. Next you decide how you'll test it: the concepts you'll put in front of customers, and the cheapest way to learn whether you're right.