011 min
Sampling at a glance
- What it is: Studying some members of a group and using the result to describe all of them, because asking everyone costs too much.
- Origin: Laplace estimated the population of France from a sample in 1786, and the 1936 Literary Digest poll became the standard failure case.
- Confused with: Sample size. Size is how many you picked. Sampling is how you picked them.
- Watch for: A very large sample chosen badly is still wrong. The Literary Digest received almost 2.4 million replies and still predicted the wrong winner.
021 min
Why sampling matters
A team that reads a survey, a usability study or a dashboard is trusting the sampling behind it, usually without seeing it. The decision at risk is simple: do these results describe the people we plan to serve, or only the people who were easy to reach?
Getting this wrong has a direct cost. A product team that surveys only the customers who answer its email will hear from its most engaged users. It may then cut a feature that most customers rely on and never hear a complaint, because those customers were not in the sample. A research lead who says "we talked to 12 people, so this is solid" is making a claim about sampling, not about the number 12.
The same question applies to an A/B test, a satisfaction score, an NPS survey and a market-size estimate. In each one, someone chose who counts. This article explains how that choice is made, how it fails, and what to ask before you act on a result.
032 min
What sampling actually involves
Four terms carry most of the meaning. They are easy to mix up, so each is defined here once.
- Population: every person or item you want to describe. It must be defined. "Our users" could mean everyone who ever signed up, or only people active in the last 30 days.
- Sampling frame: the list you actually draw from. In an opinion poll, possible frames include an electoral register and a telephone directory (Wikipedia, "Sampling (statistics)"). Anyone missing from the frame has no chance of being picked.
- Sample: the people or items chosen from the frame.
- Estimate: the number you calculate from the sample, such as "62% prefer the new layout". It is a guess at the true population value, not the value itself.
A result can go wrong at each step. If the frame leaves people out, the sample cannot include them. If chosen people do not reply, the sample shifts toward those who do. If the estimate is calculated badly, the number is off even when the sample was good. This is why a single question, "how many people did you ask?", cannot tell you whether a result is trustworthy.
Two kinds of error
Random variation is the first. Two fair samples from the same population will give slightly different numbers. Pew Research Center describes the margin of sampling error as how close a survey result can reasonably be expected to fall to the true population value. It shrinks as the sample grows.
Systematic error, called bias, is the second. It pushes every result in the same direction, and a bigger sample does not fix it. Most real sampling failures are bias, not random variation.
041 min
Where sampling comes from
Choosing by lots is an old idea. According to the Wikipedia article on statistical sampling, Pierre Simon Laplace estimated the population of France from a sample in 1786 and computed probabilistic estimates of the error. Alexander Ivanovich Chuprov introduced sample surveys to Imperial Russia in the 1870s.
The best-known failure in the history of the method is the 1936 US presidential poll described below.
052 min
Probability sampling methods
Probability sampling has two traits in common, as Wikipedia puts it: "Every element has a known nonzero probability of being sampled and involves random selection at some point." Because the chances are known, the margin of error can be calculated. The common methods differ in how they organize the draw.
| Method | How it works | Use it when | Main risk |
|---|---|---|---|
| Simple random | Every unit has the same chance, like drawing names from a full list | You have a complete frame and a modest population | Needs the full list; can miss small subgroups by chance |
| Systematic | Pick a random start, then take every kth unit | The list is long and ordered | A pattern in the list that matches k |
| Stratified | Split the population into groups, then sample inside each | You need reliable numbers for each subgroup | Needs to know group membership before sampling |
| Cluster | Select whole groups, such as offices or schools, then study all of each | Reaching individuals is costly, but groups are easy to list | Members of one cluster resemble each other, so results are less precise |
An illustrative case: a PM wants to compare a subscription app's web and mobile users. The app has 90,000 monthly web users and 10,000 monthly mobile users. A simple random sample of 1,000 would contain about 100 mobile users, which may be too few to read. A stratified sample could take 500 from each group and then weight the results back to the real 90/10 split.
061 min
Non-probability sampling methods
In non-probability sampling, some members of the population may have no chance of being selected, or the chance cannot be calculated. Wikipedia lists the main types.
- Convenience (accidental) sampling: take whoever is close at hand. The researcher cannot scientifically generalize to the whole population from it.
- Voluntary sampling: people choose to take part, for example by following a link in an advertisement. This is the setting for voluntary bias, where people who opt in differ from people who do not.
- Quota sampling: divide the population into groups, then choose people from each by judgment until a quota is filled. Wikipedia notes that interviewers may be tempted to choose people who look most helpful.
- Snowball sampling: start with a few people and ask them to recruit others. It suits hidden populations that cannot be listed.
Non-probability methods are not wrong by default. They are fast and cheap, and some research questions have no list to draw from. The cost is that the margin of error from a probability sample no longer applies, so the reader must judge how far the result can be trusted.
072 min
A worked example: how many people do you need?
For a yes/no question in a simple random sample, the number of responses needed depends on three inputs.
n = z² × p × (1 − p) ÷ e²
Here z is 1.96 for a 95% confidence level, p is the expected share (0.5 is the most cautious choice), and e is the margin of error you can accept, written as a fraction.
With e = 0.03 (plus or minus 3 percentage points): n = 1.96² × 0.25 ÷ 0.03² = 0.9604 ÷ 0.0009, which is about 1,067. Pew Research Center reports the same figure: a simple random sample of 1,067 cases has a margin of error of plus or minus 3 percentage points for estimates of overall support for individual candidates.
The second case changes one input, the margin of error, and leaves the rest.
| Accepted margin of error | Responses needed (95% confidence, p = 0.5) |
|---|---|
| plus or minus 5 points | about 384 |
| plus or minus 3 points | about 1,067 |
| plus or minus 2 points | about 2,401 |
Cutting the margin from 3 points to 2 more than doubles the responses needed. Precision gets expensive quickly. That trade-off, not any fixed number, is why a team sizing a survey should first decide how much error it can accept.
This calculation assumes a probability sample with a high response rate. It says nothing about bias. A sample of 2,401 volunteers has a precise-looking margin of error and may still miss the true value by far more.
081 min
A named case: the 1936 Literary Digest poll
The Literary Digest, a US magazine, mailed 10 million ballots before the 1936 presidential election. According to Lohr and Brick (2017), the mailing list was drawn from telephone books, club rosters, city directories, lists of registered voters, and mail-order and occupational data. Almost 2.4 million ballots came back, a response rate of 24%.
The poll predicted that Republican Alfred Landon would win with 54 percent of the popular vote. In the election, Franklin Roosevelt won more than 60 percent of the popular vote and 523 electoral votes, carrying every state except Maine and Vermont.
The failure is a clean example because the size was not the problem. Researchers still disagree on the cause, and two explanations are on record. Gallup (1938) blamed the sampling frame, which was built largely from lists of telephone and automobile owners and so overrepresented the well-to-do. Bryson (1976) argued that the frame did not explain the errors and blamed nonresponse bias, meaning that the people who replied differed from those who did not.
Lohr and Brick added a third finding. If the poll's information about 1932 votes had been used to weight the results, it would have predicted a majority of electoral votes for Roosevelt, but the bias in the estimates would still have been very large. Both a flawed frame and nonresponse are sampling problems, and neither is fixed by a larger sample.
091 min
A second case: online panels that opt in
The 1936 poll used mail. A modern version of the same question is the online panel, where people sign up to answer surveys. Pew Research Center noted in 2016 that there is no comprehensive sampling frame for the internet, no way to draw a national sample in which almost everyone has a chance of being selected.
Pew tested nine online non-probability samples from eight vendors with the same 56-item questionnaire and compared them with 20 government benchmarks. The variable that differs from the 1936 case is the way people reach the sample: a signup list and statistical weighting, not a printed directory.
Two results are worth noting. First, all of the samples evaluated included more politically and civically engaged individuals than benchmark sources indicate should be present. Second, the average estimated bias on benchmarked items was more than 10 percentage points for both Hispanic and Black adults. Samples with more elaborate sampling and weighting procedures and longer field periods generally gave more accurate results, though Pew calls that conclusion preliminary because only nine samples were tested.
101 min
How sampling shows up in product, design, and engineering work
Most product teams make sampling decisions without calling them that. These are the places it appears.
- Surveys and NPS. A PM emails a satisfaction survey to all users and gets a 4% reply. The result describes the 4%. Before acting, the PM should compare responders with the user base on plan type, tenure and region.
- Usability tests. A designer recruits five participants from a panel of people who like testing new products. The sample covers a narrow slice of the real audience, whatever its size.
- A/B tests. Random assignment of users to variants is sampling inside a product. It gives a result about the users who visited during the test window. A test run only in the first week of the month may miss users with different billing habits. See A/B testing for the full method.
- Logs and analytics. An engineer keeps 1% of events to save storage. If the 1% is chosen by time, a traffic spike may be undercounted.
In each case the decision that changes is concrete: who you can say a result applies to. A result from a sample is a statement about the frame, and only an argument links the frame to the population you care about.
111 min
Sampling in qualitative research
Not every study aims at a population number. In qualitative work, such as interviews and usability tests, the goal is to find problems and reasons, not to measure shares. These studies usually use purposive sampling, meaning the researcher picks participants because they have the experience under study.
Nielsen Norman Group advises testing 5 users in a qualitative usability study. Jakob Nielsen's argument is about return on investment: testing costs increase with each added participant, while the number of findings quickly reaches diminishing returns. The same article says to test about 5 users per distinct group when the groups differ.
Five participants cannot tell you that "40% of users struggle". They can show that a problem exists and how it happens. Claiming a percentage from a purposive sample is a common error.
121 min
Using sampling well: a short checklist
Before trusting or running a study, ask these questions in order.
Who is the population?
Write it in one sentence, with a time window.
What is the frame?
Name the list. Say who is missing from it.
How were people chosen?
Chance, convenience, volunteering, or judgment?
Who answered?
Compare responders with the population on any attribute you know.
What error can you accept?
Choose it before collecting data, then size the sample from it.
Does the claim match the method?
A volunteer sample supports "among people who chose to respond". It does not support "our users think".
If the answers to 2 to 4 are weak, say so in the report. A stakeholder can handle "this describes engaged customers" far better than a number that looks general and is not.
131 min
Where sampling goes wrong
Three failure modes cover most cases.
- Coverage error. The frame leaves out part of the population. Tell: ask who could never have been selected. A telephone directory misses people without telephones.
- Nonresponse bias. People who reply differ from people who do not. Tell: a low response rate, especially when the topic is one that some people care about more.
- Self-selection. Participants chose themselves. Tell: the invitation was open, such as a link or an in-app prompt. This overlaps with voluntary bias.
Two related traps sit outside the sample itself. A sample can show that two things move together, but not why, which is the subject of correlation and causation. And a researcher who expects a result may recruit people likely to confirm it, which is confirmation bias applied at the recruiting step.
One common claim needs correcting: that a sample must be a fixed percentage of the population to be valid. The calculation above contains no population size. For a probability sample, precision depends on how many responses you obtain and on how the sample was drawn, not on the share of the population it represents.
141 min
Sampling vs. nearby concepts
Sampling vs. sample size. The deciding fact is that sample size describes quantity and sampling describes selection. A larger sample reduces random variation. It does not remove bias, as the 1936 poll showed.
Sampling vs. a census. A census tries to include every member of the population. It has no sampling error, but it costs more and still suffers from nonresponse.
Sampling bias vs. selection bias. The Wikipedia article on sampling bias calls sampling bias usually a subtype of selection bias. It adds that a distinction, not universally accepted, is that sampling bias undermines external validity, the ability to generalize to the whole population, while selection bias mainly affects the comparison inside the study itself.
151 min
Where the evidence is contested
Two disagreements affect practice.
What went wrong in 1936. As described above, Gallup (1938) blamed the frame and Bryson (1976) blamed nonresponse. Lohr and Brick, writing in 2017, found that weighting could have changed the predicted winner while the estimates stayed badly biased. Until the question is settled, a careful reader treats the poll as evidence for both risks, not for one.
Whether non-probability samples can be trusted. Pew's 2016 report says that for roughly 15 years independent studies suggested the answer was generally "no" for accurate population estimates. It adds that researchers and vendors have since developed techniques to improve representativeness, and that some recent case studies suggest accurate estimates may be possible without a probability sample. Pew's own test found large variation between vendors and large errors for some subgroups.
In practice, use probability samples when a number will drive a large decision, and use non-probability samples for exploration with the limits stated.
?8 questions
Questions people ask
What is sampling?
What is the difference between probability and non-probability sampling?
What is sampling bias?
How big does a sample need to be?
Does the sample need to be a fixed percentage of the population?
What is a good example of sampling going wrong?
How do you avoid sampling bias?
Is 5 users enough for a usability test?
§6 sources
Sources
Lohr, S. L. and Brick, J. M. (2017). Roosevelt Predicted to Win: Revisiting the 1936 Literary Digest Poll. Statistics, Politics and Policy 8(1), 65-84. DOI 10.1515/spp-2016-0006.
Mercer, A. (2016). 5 key things to know about the margin of error in election polls. Pew Research Center.
Kennedy, C., Mercer, A., Keeter, S., Hatley, N., McGeeney, K. and Gimenez, A. (2016). Evaluating Online Nonprobability Surveys. Pew Research Center.
Wikipedia. Sampling (statistics). (statistics)
Show all 6 sourcesShow fewer sources
Wikipedia. Sampling bias.
Nielsen, J. (2012). How Many Test Users in a Usability Study? Nielsen Norman Group.
