4  Designing Experiments to Learn Demand

Asking a question whose answers can become a demand curve

This chapter is about getting the right evidence rather than doing the mathematics.

Design the experiment well and the analysis that follows is close to mechanical. Design it badly and the analysis will run perfectly and produce nonsense, which is the more dangerous outcome, because nonsense that has been through a model comes out looking like a finding.

So the work here is judgment rather than statistics. Before asking anyone what they would pay, you have to know what question you are answering and what kind of demand that question implies.

From a Decision to a Question

Entrepreneurs start with a decision, not a dataset. Should I build this? What should I charge? Is this worth another six months? Those are the right questions to be asking and none of them can be put to a customer.

Translating them into something answerable is a step that is easy to skip and expensive to skip. It takes four commitments.

Who is the customer. Whose decision are you trying to learn about, stated precisely enough that someone else could sort a list of people into in and out. A survey of anyone who might like this produces evidence about nobody in particular.

What counts as one unit, and over what period. Both were settled in the previous chapter, and both are load-bearing here rather than administrative. Every question you write from this point forward refers to them, and if either drifts mid-survey the responses stop being comparable to each other.

What evidence would change the decision. If the answer came back higher than you expect, what would you do? Lower? If both answers lead to the same next move, you are about to spend weeks confirming something that cannot matter.

These four are not preliminaries to get past. They define what demand means for the decision in front of you, and a survey written without them collects numbers that cannot be interpreted afterward at any level of statistical care.

Is Quantity Fixed, or Is It the Unknown?

Once the unit and period are fixed, a fork appears that changes everything downstream.

In some decisions quantity is settled by the shape of the thing. A customer subscribes or does not. They buy the license or they do not. They pick one plan. Within the period, if they buy at all, they buy one, and the only uncertainty is who buys at which prices. Call this yes/no demand.

In others, quantity is itself part of what you need to learn. Meals per week, rides per month, replacement filters per year. Customers differ not only in whether they buy but in how much, and the difference between a light user and a heavy one may be the difference between a business and a hobby. Call this how-many demand.

Yes/no and how-many demand — whether a customer can buy at most one unit in the period, or several. The fork decides which questions you ask and what the answers can support.

The fork matters because the two require different questions, produce differently shaped data, and support different conclusions. Designing a yes/no experiment for a how-many decision does not make your answer less precise. It changes what your answer is about, and it does so invisibly, because the data will look fine.

So before writing a single question: is quantity fixed at one, or is quantity the thing I am trying to learn?

For the Curious — What economists call these

The plain names used here have precise counterparts. A customer who buys at most one unit in the period is said to have unit demand, a term from auction and mechanism design; a customer who may buy several has multi-unit demand. Empirical economists cut the same distinction differently and speak of the extensive margin, whether someone participates at all, against the intensive margin, how much they buy given that they do. Yes/no demand varies only on the extensive margin. How-many demand varies on both, which is exactly why it takes more questions to pin down.

It is worth seeing why the more familiar durable-versus-nondurable split does not do this job, since it is the first place most people reach. Durability is a property of the good. This distinction is a property of the purchase decision inside the period you defined, which means the same good can land on either side depending on a choice you made earlier. A car is yes/no demand over a year and how-many demand over twenty. A sandwich is how-many demand over a month and yes/no over a single lunch break.

That is not a defect in the categories. It is a consequence of the fact that quantity has no meaning until a period is fixed, and it is one more reason to settle the unit and the period before anything else.

What Willingness to Pay Actually Is

Willingness to pay is the most misunderstood instrument in early demand work, mostly because the phrase sounds like it means the price to charge.

Willingness to pay — the highest price at which somebody would still buy, given a stated unit and period. A boundary on their behavior, never a recommendation about yours.

It does not. A person’s willingness to pay is a boundary: the highest price at which they would still choose to buy, given a specific unit in a specific period. Above it they do not buy. Below it they might. That is the whole of what the number claims.

Read as a boundary rather than a recommendation, it becomes useful precisely because it is modest. It rules out prices that cannot work. It identifies ranges where demand plausibly exists. It converts a vague sense of people would probably pay something into a structure you can reason about, which is a smaller promise than accuracy and a far more reliable one.

It is also not a stable property of a person. The same customer will report different values depending on what exactly they are buying, how often the decision repeats, what else is available, and what constraints they are under that week. This is why context is not decoration around the question. It is part of the measurement.

Ask for a Maximum, Not a Verdict

There are two shapes a money question can take, and confusing them is the most common way a demand survey fails while appearing to succeed.

Would you pay $40 for this?

A yes here tells you that this person’s boundary sits somewhere above forty dollars. It does not tell you where. Ask a hundred people the same thing and you have a hundred inequalities all pointing the same direction, from which a curve can be recovered only by varying the price across respondents and assuming a shape for the distribution. That is a real method, and it needs a much larger sample than you are likely to have.

What is the most you would pay per unit, per period?

This returns the boundary itself: the price above which that person buys nothing at all. One respondent, one point. Forty respondents, forty points, and the curve is built by counting rather than by fitting an assumption.

Both questions are legitimate and they belong at opposite ends of the work. Eliciting a maximum is how you build a curve. Naming a price and asking for a yes is how you test one, later, after the curve exists and the arithmetic has proposed a price worth testing. Reaching for the test form early feels efficient and quietly costs you the thing you came for.

Avoid, in both cases, asking what a fair or reasonable price would be. Those questions collect a respondent’s theory of pricing rather than their own boundary, and the two have very little to do with each other.

Asking for Yes/No Demand

When the decision is binary, one question carries the weight.

Ask for the maximum price at which the respondent would still choose to buy one unit in the defined period. Make sure they know exactly what one unit is. Treat what comes back as a boundary rather than a promise.

That is all the elicitation requires, and it is enough, because each answer places one person on the price axis. Sort those boundaries and you can read off, at any price you like, how many people are still standing.

Asking for How-Many Demand

When quantity varies, one number per respondent is not enough. You need their whole relationship between price and quantity, and there are two practical ways to get at it.

Naming a series of prices

Show respondents a short list of prices and ask how many units they would consume at each. At $12, how many per month? At $15? At $18?

This is intuitive and it mirrors how people meet prices in the world. It also returns quantity at specific prices directly, with no inference required, which is reassuring.

The cost is fatigue. You are asking one question repeatedly with a number changed, and attention falls off a cliff after the first few. Keep the list short, four to eight prices at most, and accept that later answers are worth less than earlier ones. The repetition also makes the exchange feel like haggling, which tends to pull answers downward.

Anchoring at both ends

The alternative asks three questions that each require a different kind of thinking.

  1. If this were permanently free, how many would you use per period?
  2. What is the most you would pay per unit?
  3. At that price, how many would you use per period?

The first strips price out completely and asks about need and usage, which sets a collaborative frame and gives you a ceiling on consumption. The second is the boundary. The third finds the other end of the same person’s demand.

Two points and a cut-off define that respondent’s individual demand: quantity falls from the free-consumption level down to the level at their maximum, and above their maximum they are gone entirely. Do that for everyone, add the individual curves together at each price, and you have a market demand curve assembled out of the respondents’ own numbers.

For the Curious — Why “per unit” and not “for one unit”

The second question asks for the most someone would pay per unit, and the wording is deliberate. The natural alternative, the most you would pay for one unit, reads as the reservation price of a single item, and if that is what it meant then the answer to the third question would always be one: at a price exactly equal to what the first unit is worth to you, you take exactly that one. The third question would be redundant.

It is not redundant, and the collected survey data says so. Respondents routinely report that they would still take several units at their maximum price. Two things account for it.

The period caps consumption before value runs out. Somebody who would happily buy five lunches a week is not revealing that the sixth is worth less to them; they are revealing that there are only five working days. When the binding constraint is the calendar rather than diminishing value, quantity sits flat right up to the price ceiling and then stops.

Units are often inputs to a goal rather than ends in themselves. Someone trying to reach a hundred grams of protein a day does not value a twenty-gram serving at a fifth of the goal. They need five, and the fifth is worth as much as the first, because four servings do not get them there. Where usage runs in fixed proportions like this, marginal value does not decline smoothly across the units at all.

So \(\mathsf{P_{max}}\) is a ceiling on the price per unit, not the value of one particular unit. At or below it they buy however many the period and the goal call for. Above it they buy none.

That last sentence is what makes the individual demand curve drop to zero rather than tail off. It looks abrupt on a chart, and it is the economically correct shape: a per-unit ceiling exceeded does not make someone buy fewer, it makes them stop.

Two features of this deserve notice. The prices at which the market curve bends are the respondents’ own maximum willingness to pay, so you never had to guess which prices to test. And the yes/no case turns out to be this same procedure with the quantities set to one: each person contributes a single unit below their boundary and nothing above it. The two methods are one method. Yes/no is what the anchor approach becomes when nobody can buy more than one.

The tradeoff is that quantities at intermediate prices are inferred rather than observed. Three questions instead of eight buys you engagement and a less transactional frame, and asks you to be comfortable treating the middle of each respondent’s curve as an estimate. Which cost you would rather pay depends on your respondents, and the honest answer is that both approaches work when the rest of the design is sound.

A Structure That Keeps the Data Honest

Respondents are not adversaries and they are not measuring instruments either. They arrive with incentives, assumptions, and a strong instinct to be socially appropriate, and if the design ignores those forces the data will look precise while being wrong in a consistent direction.

Good design does not remove that. It redirects it.

The most damaging framing is the one that arrives by default: show someone a product and immediately ask what they would pay, and you have staged a negotiation. Some respondents shade downward, because understating costs them nothing and might get them a better price later. Others shade upward, because they like you, or the idea, and want to be encouraging. Neither is dishonest. Both are ruinous, and they do not cancel.

The fix is to separate evaluation from valuation, and to do them in that order.

  1. Screen. Establish that this person belongs to the group whose decision you are studying, before they answer anything else.
  2. Set the context. Explain what is being considered and why an honest answer matters. Respondents who understand the thing may not get built answer differently from respondents who assume it is already a product.
  3. Make the decision real. Describe the problem and the proposed solution well enough that they can picture meeting it. This is not marketing copy; it is giving them enough to reason with.
  4. Ask about appeal, without money. What do they think of it, what do they like, what worries them, what would they change. This lets enthusiasm and doubt come out somewhere other than the price.
  5. Then ask about money. Willingness to pay arrives as one more way of expressing a judgment they have already been forming, rather than as the opening move in a haggle.
  6. Collect characteristics last. Whatever you need to interpret variation and sharpen the customer definition, once the valuable answers are safely recorded.

By the time price comes up, the respondent has already said what they think in words and ratings. The money question becomes a consistency check on a judgment in progress, which is a very different cognitive task from being asked to name a number cold.

Hypothetical Bias

Ask people about a choice they have not made and their answers will carry more than preference. Imagination, intention, mood, and a wish to be helpful all come along.

This is not dishonesty. It is what answering a hypothetical question consists of.

Willingness to pay is especially exposed, because it compresses several judgments into one number. The respondent has to picture the product, the situation they would use it in, what else they might buy instead, what money is tight that month, and what they will care about later. Small changes in framing move all of that, which is what makes the instrument powerful and fragile at once.

The error is not using stated willingness to pay. The error is reading it as a promise instead of as evidence with known limits. Every method of learning demand before revenue exists leans on somebody’s imagination somewhere; a survey has the virtue of making that reliance visible, where a projection hides it inside a spreadsheet.

Stated demand is not false. It is incomplete.

Imperfect Evidence, Used Responsibly

Early demand evidence is nearly always imperfect, and this is structural rather than a failure of care.

Convenience samples have a legitimate job. Entrepreneurs are warned off them for good reason, since a sample of whoever was nearby is rarely representative and treating it as though it were causes real damage. But refusing them entirely gives up something valuable. Run your instrument past a handful of convenient people first and you find out whether the questions are understood as intended, whether the unit is clear, whether the scenario is plausible, and where people hesitate. Used that way a convenience sample is not evidence about demand at all. It is evidence about whether your instrument works. The mistake is not starting there; it is stopping there.

The biases run in both directions. Some respondents understate, guarding against a future price or avoiding the look of eagerness. Others overstate, out of sympathy or because the idea is pleasant in the abstract and the consequences are not real. These do not cancel neatly, and pretending they do is wishful. What they usually produce at the sample level is not fantasy but noisy direction: a curve whose shape and rough position mean something and whose exact values do not.

Which is why an early demand curve is a provisional map. It rules things out, points at promising regions, and tells you where to look next. It does not yet justify a commitment you cannot reverse.

Validation closes the loop, and this is where a named price finally earns its place. Return to a fresh group and ask whether they would buy at $X, where X came from the curve rather than from your hopes. The answers should line up roughly with what the curve predicted. If they do, your confidence rises, not because the curve is correct but because the evidence is coherent. If they do not, you have learned that something in the framing or the sampling distorted the earlier work, which is worth knowing before you act on it. The strongest version of this test is not a question at all but a preorder, since a person who pays has settled the matter.

Before you field an instrument

An instrument goes out once. Whatever it fails to ask, you will not learn without asking a second group, and there is rarely a second group. Read your draft as a respondent and check three things.

  • Where would you guess? Find every question you could not answer honestly about your own life, and either fix it or accept that its answers are noise.
  • Where would you answer to please? Anything that reads as selling invites agreement. The price question is the one that suffers, and it is the one you came for.
  • Does a number come back that you can count? A rating is not demand and neither is interest. If no question yields a price and a quantity, the instrument is a conversation rather than an experiment.

If you cannot answer all three, you are about to spend your one contact with these people learning something else.

Putting It to Work

Ask yourself — could someone else run this?

Write your demand question out in full, exactly as a respondent would meet it, with the unit and the period in the sentence rather than assumed.

Then check it against the commitments this chapter asked for. Could a stranger sort people into your customer group and everyone else, using only what you wrote down? Is quantity fixed at one, or is it the thing you are trying to learn, and does your question match your answer? Does the question ask for a maximum, or does it name a price and ask for a verdict?

Now the one people skip. Write down what you would do if the answers came back strong, and what you would do if they came back weak. If the two are the same, stop. You have designed a survey that cannot change anything, and the cheapest moment to discover that is before you send it.

The move: Ask for the most someone would pay, not whether they would pay a number you chose. The first builds a curve. The second tests one you do not have yet.

A well-designed instrument still returns a pile of individual answers rather than a demand curve. Turning one into the other is a specific transformation, and it is more intuitive than it sounds once you see what it is doing.