5  Executing Demand Experiments

Getting the right people to answer, and knowing what their answers are worth

A good design is necessary and it is not sufficient.

Between a well-built survey and usable evidence lies execution, which is where plans meet the world and where sound experiments quietly fail. Who actually sees the thing, who bothers to answer, and what they think you asked all shape what comes back, and none of those is settled by the design document.

The risks are rarely obvious, because bad execution does not usually produce chaos. It produces data that is clean, coherent, and pointing the wrong way. Noise announces itself. This does not.

Four things can go wrong between sending a survey and believing its answers. This chapter takes them in the order they happen.

Who Had a Chance to Answer

Start with the question that comes before all the others: who was actually given an opportunity to respond?

Your population is everyone whose decision the venture depends on. Your sample is the handful who answered. Between them sits the sampling frame: the subset who could have answered, because the survey reached them at all. The frame is almost never a small copy of the population.

Sampling frame — the people your survey could actually reach. Neither the population you care about nor the sample who answered.

Two things are easy to confuse here. Eligibility is not representativeness: everyone in your frame might genuinely belong to your customer group and still, collectively, look nothing like it. And size is not composition. Four hundred responses from people who all found you the same way is a large sample of a narrow slice, and the largeness does nothing to fix the narrowness. Doubling a biased sample doubles the bias with more decimal places.

Composition is what matters for demand specifically, because a demand curve is built from the spread of willingness to pay across people. If your frame systematically omits the price-sensitive, the curve does not come out slightly wrong. It comes out the wrong shape, too high and too flat, and the flatness is the dangerous part: it tells you that raising the price costs you very little.

Convenience shapes frames more than anyone likes to admit. You post where you have standing, ask people who will take your call, and reach through networks you already have. That is not a sin and it is often the only way to begin. It does mean the frame carries the shape of your own reach, and your reach is not neutral with respect to who values what you are selling.

So the protective rule is simple. When the frame is narrow, keep the conclusions narrow with it. Evidence gathered from a distinctive subset describes that subset. That is a real finding and a smaller one than you wanted, and saying so plainly costs you nothing you actually had.

Getting People to Answer

The last section asked who had a chance to respond. The next one asks who took it. Between those two questions sits work that nobody assigns you, and that this book has so far left out.

Your Sampling Frame Is Something You Built

It is easy to read the previous section as describing a frame you can go and inspect: where the survey was posted, who it was sent to, who happened to see it. That reading is too passive. Nobody hands you a frame. You build one, person by person, and what you end up with is the residue of where you looked, who agreed to help, and where you eventually stopped trying.

This is why two people studying the same market can produce evidence that cannot be compared. One posted a link in a general forum and took whoever wandered in. The other spent a week persuading a trade association to forward it to their members. Both have forty responses. Only one of them has forty responses from the people the decision is about.

Response Rate Is a Leading Indicator of Acquisition Cost

Most people treat low response as an obstacle standing between them and the real work.

Do the arithmetic before you accept that. Suppose it took two hundred approaches to collect forty answers to a survey that costs a respondent three minutes and no money whatsoever. You will eventually need people to do something considerably harder than that. They will have to hand you money.

Acquisition runs through the same channel you just used, against much higher friction. What it costs shows up in the profit equation directly: in variable cost if you pay for each customer acquired, in fixed cost if you build a channel and staff it. A channel that cannot deliver free survey responses cheaply is unlikely to deliver paying customers cheaply.

So the recruiting you are finding tedious is not standing between you and your measurement of \(\mathsf{q}\). It is your first measurement of \(\mathsf{c}\) and \(\mathsf{f}\).

Take that literally rather than as consolation. Record how many approaches each response cost you, and record it by channel. Those numbers are evidence, and they are close to the only evidence about acquisition cost available to you before you have something to sell.

Track the Funnel, Not the Count

The number that matters is not how many responses you have. It is where you lost everyone else.

Count people at four points: approached, started, passed the screener, finished. Each gap says something different, and each one is a finding rather than a failure.

Lost before starting. You cannot reach this group through this channel. That is an access finding, and it is often worth more than what the survey would have told you.

Screened out in volume. The segment is rarer in this population than you assumed. Notice what you now have: an empirical estimate of how common your customer actually is, which is precisely the number market sizing asserts and never measures.

Started, then abandoned at the price questions. Something is unclear, in the instrument or in the product description, or the questions ask more than a stranger will give. This is fixable, and worth fixing before you gather anything further.

Finished, with willingness to pay clustered near zero. You have your answer. It arrived cheaply and early, which was the entire point of asking.

Everyone who failed to finish told you something. A raw count throws all of it away.

Where the People Already Are

The common failure is to broadcast and hope the right people identify themselves.

Broadcasting works badly for a structural reason. Your target segment is by construction a minority of any general audience, so a general channel delivers mostly people you will screen out, at full effort per person. You exhaust yourself assembling a sample you then discard.

Somebody has already done the aggregating for you. Your people gather somewhere: a club, a trade association, a subreddit, a Discord, a conference, a waiting room, a shop that serves them, a queue they stand in. Find the gathering rather than the individuals, then ask whoever runs it.

Borrowed credibility is the largest single lever available to you. A link from a stranger converts close to nothing. The same link forwarded by a moderator, an organizer, an association’s staff, or a shop owner converts many times better, because the trust already exists and you are borrowing it instead of building it. A week spent earning one introduction usually beats a week spent posting.

Ask in person, deliver digitally. A face-to-face request followed by I’ll send you the link outperforms a purely digital approach by a wide margin, because the commitment becomes social rather than transactional. This is the move most people avoid and the one that works.

Be specific about why you are asking. I’m a student working on a project requests charity. I’m trying to work out whether this particular thing, for people like you, is worth building, and I’ll send you what I find offers an exchange.

Length is the biggest thing you control. Every question costs completions, and the price questions you actually need are already demanding. Cut to the screener, the willingness-to-pay block, and only those characteristics you will genuinely segment on. Everything else is a tax paid in lost responses.

What to Offer

Incentives are where a demand survey differs from every other kind of survey.

An ordinary survey can pay its respondents without much harm. This one cannot, for a specific reason. You are asking people what they would pay. Paying them to answer makes money salient at exactly the wrong moment, and it recruits people who answer surveys for money, a group defined by a trait that has nothing to do with wanting your product.

Offering the product itself is worse, which is counterintuitive enough to say plainly. Cash selects for people who want cash, roughly unrelated to your question. Your product selects for people who want your product, which is the very thing you are trying to measure. You would be recruiting on your own dependent variable. And giving it away announces that its price is zero, moments before you ask what someone would pay for it. A discount carries the same defect in milder form, and adds one of its own by naming a price before you ask for theirs.

What remains is workable. Early access at the eventual price concedes nothing on price and selects only mildly on interest. An adjacent item, raffled (an accessory, or something else from the same category) is the best incentive most ventures can offer, because it biases toward category interest, which overlaps your screener rather than your outcome. The findings themselves, sent back to everyone who answered, remain the cleanest offer available and cost you nothing but the sending.

One further possibility belongs outside this list, because it is not an incentive at all.

If you are tempted to give people the product in exchange for answering, consider selling it to them instead. A preorder, a deposit, or a pledge at a stated price is not a stated preference. It is a purchase, and it is stronger evidence than any survey can produce. It does not replace the survey, which estimates a whole curve where a preorder gives you a single point. It validates one, at a real price, with real money, which is exactly the one-point check the estimated curve should be held to.

When You Cannot Reach Them At All

Sometimes the honest outcome of recruitment is that it does not work.

You tried several channels, spent real effort, and could not assemble enough of the right people to say anything. The temptation then is to widen the definition until responses arrive, which converts a clean negative result into a misleading positive one.

Record what happened instead. A group you cannot reach is not a market you can serve. Whatever they might have paid, you have no way to put an offer in front of them, and no way to do it repeatedly at a cost you could carry.

That is not a failed experiment. It is an answer to a question the profit equation asks and market sizing never does.1

Who Chose to Answer

Of the people who could have responded, some did. That was a decision, and decisions are not coin flips.

The people who answer are more interested in the problem, more opinionated about the solution, more inclined to be helpful, more optimistic, or more personally connected to you. None of that makes them dishonest. It makes them selected, and selected in a direction you can predict before you look at the data.

The direction is consistent enough to state. Early demand evidence tends to overstate appeal, overstate willingness to pay, understate price sensitivity, and underrepresent indifference. Which is why early results so often look encouraging and so often fail to survive contact with a wider audience. Your respondents were not lying. They were not typical.

The mistake that does the most damage is treating silence as missing data. It is not missing; it is an answer you did not record. People who ignored your survey may be indifferent, may be busy, may have found the offering irrelevant, or may not have cared enough to click. And here is the part worth sitting with. Your marginal customer is the one deciding right at your price, and it is that person, rather than your enthusiasts, who determines whether demand is steep or shallow, robust or fragile. The people who did not answer are usually closer to your marginal customer than the people who did. Your dataset is systematically thinnest exactly where the answer is decided.

Marginal customer — the person deciding right at your price, whose choice settles whether demand there is steep or shallow.

Selection bias also compounds rather than adds. It operates inside whatever frame you already have, so a narrow frame full of engaged people gets narrower and more engaged. A well-designed survey can still produce confident nonsense if this is ignored.

None of which means surveys are useless or early demand cannot be learned. It means the evidence is conditional rather than definitive, and it is strongest when it lines up with what you would expect on other grounds, survives a follow-up with fresh people, and stays stable across different samples and framings.

What They Thought You Were Asking

Even with the right people answering, one risk remains: they may have answered a different question from the one you wrote.

A survey question is not transmitted. It is interpreted. Every respondent brings assumptions about what the product actually is, what one unit means, how often the decision comes up, what constraints apply, and what situation they should be imagining. When those assumptions differ from yours, the answers can be internally consistent, carefully considered, and about something else.

Two of these you have already handled. Ambiguous units and drifting time frames are the most common sources of interpretation error, which is precisely why the previous chapter spent so long fixing them before any question got written. If either is loose, this is where the looseness turns into data.

The third is harder to design out. Respondents quietly substitute a context of their own: a different version of the product, a different set of alternatives, a different budget, a different level of urgency. They answer sincerely about the thing they pictured. The more abstract your description, the more room there is for the picture to drift, which is the practical argument for making the scenario concrete rather than aspirational.

What makes measurement error dangerous is that it does not look like error. Responses come back smooth and well behaved and simply misaligned, and analysis then processes them faithfully. Precision is applied to the wrong quantity and the output is confident and wrong.

You can watch for the symptoms. Unexpected dispersion in answers that should cluster, quantities or prices that are not plausible for the unit you defined, confusion visible in the open-ended responses, and difficulty explaining a result in plain language to somebody who was not involved. That last one is the most reliable test there is, and it costs nothing.

When these appear, the response is redesign rather than adjustment. There is no statistical repair for a question that was understood differently by different people, and attempting one converts a fixable problem into a permanent one.

Why Bad Data Is Worse Than None

Demand experiments rarely fail loudly. They produce results that look reasonable and quantitative and actionable while pointing somewhere you should not go.

Consider what you have when you have nothing. You know you are guessing. That knowledge keeps decisions provisional, keeps you alert, and keeps you willing to change your mind when something surprising turns up. Ignorance, when it is visible, is a functioning safety mechanism.

Bad data disables it. Numbers appear, charts get drawn, estimates get computed, and the language of evidence enters the room. The uncertainty has not gone anywhere; it has just become invisible. Worse, decisions justified with numbers feel responsible, which makes them far harder to unwind later when the picture changes.

Analysis will not save you here. A model assumes its inputs mean what you think they mean. It has no way to ask whether you sampled the right people, whether respondents read the question consistently, whether the prices and quantities refer to the same unit and period, or whether anyone found the scenario believable. Violate those assumptions and models do not fail gracefully. They amplify the error with precision, and what comes out is not noise but confident misdirection.

Entrepreneurs are exposed to this more than most. The decisions feel urgent, so there is pressure to have something rather than wait for something better. Early evidence comes from small samples, where design and execution errors do their most concentrated damage. And once a price is set or an investment is made, revisiting the evidence underneath it becomes personally and organizationally awkward. Bad data does not mislead one decision. It sets a direction and then makes the direction expensive to question.

So a rule worth carrying:

If the data would make you more confident than you should be, do not use it yet.

This does not ask for perfect evidence, which does not exist at this stage. It asks for evidence whose limits you can state. Sometimes the disciplined response to a weak experiment is not correction but restraint, and declining to act on fragile results is itself a decision you can defend.

None of this chapter was really about technique. It was about protecting your judgment from your own instruments, so that whatever confidence you end up with was earned and whatever uncertainty remains is still visible to you.

Putting It to Work

Ask yourself — who is missing, and which way does it push?

Before you look at a single number, write two lists. Who was most likely to answer this, and who was most likely to ignore it.

Then name the direction of the error rather than pretending there is not one. If your respondents skew toward people already sympathetic to the idea, your curve sits too high and too flat, which means appeal is overstated and price sensitivity is understated. A curve you know is optimistic remains useful, because it can still tell you when the answer is clearly no.

Then the harder move. Find three people who did not respond and ask them why. They sit closer to your marginal customer than anybody in your dataset, and the marginal customer is the one who settles how steep your demand really is.

The move: Name the direction of your error before you look at the number. Evidence whose limits you can state beats evidence you simply trust.

You now have a pile of individual answers and a reasonable sense of what they are worth. Turning them into a demand curve is a single transformation, and it is more intuitive than its reputation suggests.


  1. Expeditionary Innovation makes this test the explicit gate of its first phase: see Test Your Access. The act is the same and the purpose differs. There you are checking whether sustained contact is possible before committing months to a community. Here you need that contact to produce a demand curve. A failure in either place is the same finding.↩︎