Survivorship Bias

A sampling error that comes from studying only the cases that survived a selection process and ignoring those that dropped out, so results look better than the full group.

11 min read

Β· Also in

By Ravi SuranaUpdated 6 sources

Quick answer

~20 sec

Survivorship bias is a sampling error that happens when you study only the cases that passed a selection process and ignore the ones that dropped out. Survivors are not typical, so conclusions drawn from them look better than reality. The standard case is Abraham Wald's analysis of damage on World War II aircraft that returned from missions.

011 min

Why survivorship bias matters

A product team reads its feedback survey and sees a score of 4.6 out of 5. The survey went to current customers. The customers who left were never asked. The team concludes that people like the product, and it spends the next quarter on small polish instead of on the problem that made people leave.

This is the decision survivorship bias damages: trusting a number that was computed after the selection had already happened. The cost is not random noise. The error always points the same way, toward success, because the cases that failed are the ones that left the data.

The same structure appears in three places this audience meets often: reading success stories, reading dashboards built on active users, and evaluating any group where weak members drop out over time. In each, the sample you can see is not the group you care about.

The bias is hard to notice for one reason. A missing case leaves no trace in the data. You cannot see a gap you were never shown.

022 min

How survivorship bias distorts a sample

In sampling terms, this is a selection effect: the process that decides who is in the data is linked to the thing you measure. Survivorship bias is the case where that process is survival. Something is lost along the way, and what is lost is related to the outcome.

Two conditions must both hold.

  • A selection step removes some cases before you measure. Examples are an aircraft that does not return, a company that closes, a customer who cancels, a fund that is shut down.
  • The cases removed differ from the cases kept in the quantity you want to study.

If the removed cases were a random draw, the remaining sample would still represent the whole group. The error appears only when removal is linked to the outcome. A fund is usually shut down because it did badly. An aircraft is usually lost because of a serious hit. A customer usually cancels because of a problem.

The result is a sample that looks better than the group it came from, or that shows a pattern the whole group does not have.

Group you want to studySelection stepWhat you can seeWhat the survivors hide
Aircraft sent on missionsAircraft hit in critical places do not returnDamage on returning aircraftDamage on lost aircraft
Mutual funds over 15 yearsWeak funds close or mergeFunds still operatingThe funds that closed
Customers who signed upUnhappy customers cancelFeedback from current customersReasons for leaving
Startups in a case studyFailed startups are not written aboutSuccess storiesThe same habits in failed startups

The fix is also the same in every row. Find a way to account for the cases that left, or avoid generalising beyond the survivors.

032 min

Wald's aircraft work and survivorship bias

Abraham Wald was a mathematician at Columbia University. During World War II he worked in the Statistical Research Group, a team of statisticians that answered questions for the US military. Marc Mangel and Francisco Samaniego described his aircraft work in the Journal of the American Statistical Association in 1984, at about the time the original memoranda were first made public through the Center for Naval Analyses.

Wald's work on this problem is a set of eight memoranda, and five of them deal with one question: how likely is an aircraft to survive, given that it has already been hit. The data he was given covered only aircraft that returned: how many were sent, how many came back, and how many hits each returning aircraft had taken.

The key idea in the work is to treat the aircraft that did not return as a quantity to estimate, not as cases to ignore. Wald could not observe the lost aircraft. He used the observed distribution of hits among the survivors, and stated assumptions, to estimate how many of the lost aircraft had taken each number of hits.

Wald's first simplifying assumption was that aircraft are lost because of enemy fire, not because of mechanical failure. This is typical of how he worked. He wrote his assumptions down, because the answer depends on them.

Casselman describes the aircraft damage problem as an example of what is known as survivorship bias. The memoranda are the source people cite when they explain it.

042 min

Worked example: Wald's 400 aircraft

Casselman, writing for the American Mathematical Society, reproduces the numbers that Mangel and Samaniego used to explain Wald's method. I use them here as an illustration of the method.

Suppose 400 aircraft fly a mission. 380 return and 20 do not. The 380 survivors took the following numbers of hits.

Hits takenAircraft that returnedShare of the 400
032080.0%
1328.0%
2205.0%
341.0%
420.5%
520.5%

A reader who looks only at the survivors computes the average: (32 x 1 + 20 x 2 + 4 x 3 + 2 x 4 + 2 x 5) / 380 = 102 / 380, or about 0.27 hits per returning aircraft. That looks mild. It is also incomplete, because it ignores the 20 aircraft that were lost.

Wald's question was different: given that an aircraft has been hit, what is the chance it survives? To answer it, he let q be the chance of surviving one hit. Casselman shows that if every hit is assumed to have the same chance of being fatal, the numbers above give a value of q near 0.85. In other words, about 15% of hits are fatal.

Casselman notes that this assumption amounts to saying that a hit does not weaken an aircraft, which does not seem to be far off the truth.

Compare the two numbers. The loss rate of 5% is a figure per aircraft. Wald's figure is per hit: each hit carries about a 15% chance of loss. The loss rate per aircraft is lower only because 80% of the aircraft took no hit at all. The second number is the one that helps a planner decide how much protection a part is worth, and it can only be reached by reasoning about the aircraft that did not return.

Change one variable to see why this matters. Suppose 40 aircraft were lost instead of 20, with the same hit counts among the 380 survivors. The survivors look identical in both cases. But the second case involves more fatal hits, so the chance of surviving a hit must be lower. The survivors alone cannot tell the two situations apart.

052 min

A second case: mutual fund performance

The first case concerns physical damage. This one concerns money, and it differs in one variable: the selection step is a decision by a fund company, not enemy fire.

A mutual fund that performs badly is often closed or merged into another fund. Databases that list only funds still operating therefore describe a group of winners. Elton, Gruber and Blake studied this in a paper published in 1996. They tracked every fund that existed at the end of 1976 and included returns from funds that merged. Elton, Gruber and Blake wrote in 1996 that funds that disappear tend to do so because of poor performance.

Carhart and colleagues found that funds in their sample disappear mainly because of multi-year poor performance. So the removal is linked to the outcome, which is the condition that produces the bias.

How large is the error? Carhart, Carpenter, Lynch and Musto measured the survivor bias in annual return at 17 basis points for one-year samples, 43 basis points for five-year samples, and about one percent for data sets longer than fifteen years. A longer sample gives more time for weak funds to be removed, so the survivors look better the longer you follow them.

Here is an illustrative calculation with round numbers. A group starts with 100 funds. After ten years, 20 have closed after poor results, with an average annual return of 3%. The 80 survivors average 8%. The average of the whole group is (80 x 8% + 20 x 3%) / 100 = 7.0%. A study of survivors reports 8.0%, which is one percentage point too high.

The effect can also create a false pattern. Brown, Goetzmann, Ibbotson and Ross showed in 1992 that a sample truncated by survivorship can make past mutual fund performance look like a predictor of future performance. The whole group may not have that pattern.

What the contrast shows. In the aircraft case the lost items were physically gone. In the fund case the lost items were removed from the database by the industry itself. Both cases need the same check: ask what left the sample, and why.

062 min

How survivorship bias shows up in product, design, and engineering work

Surveys and feedback. A survey sent to active users cannot report why users left. This is the opening error again, with a concrete fix. An illustrative case: a product manager reads a feedback score from current users and decides the onboarding flow is fine. Exit interviews with people who cancelled in the first week could show a different picture. The decision this changes is where the next research round goes.

Usage dashboards (illustrative). A chart of average sessions per user over six months looks like it is improving. The users with few sessions stopped using the product and left the denominator. The remaining users are the committed ones, so the average rises with no improvement in the product. Check the metric on a fixed cohort that you follow from sign-up.

Experiments. An A/B test that analyses only users who finished onboarding can produce a false winner. If a new variant makes the flow harder, weaker users leave in that variant, and the users who finish look stronger. Analyse everyone who was assigned to a variant, including those who left.

Success stories. A founder studies ten successful companies and finds they all moved fast and raised money early. Companies that failed after doing the same things are not in the sample, so the study cannot tell whether those habits help. This is a correlation without causation problem built on missing cases.

Engineering. A team reviews only the incidents that were written up. Smaller failures that were never reported are missing. A reliability review that uses only reported incidents misjudges how often systems fail.

In a related case, people who choose to answer a survey are not a random sample either, which is the topic of voluntary bias. The two errors often appear together: first some people leave, then only some of the rest respond.

072 min

How to check for survivorship bias

Ask these questions before you trust a result.

  1. What was the starting group?

    Count everyone who entered the process, not only those at the end. If you cannot find the starting count, treat the result as unreliable.

  2. What removed cases from the group, and when?

    Name the step: a cancellation, a closure, a failed flight, a filter in your own query.

  3. Is the removal linked to the outcome?

    If cases leave because of the thing you measure, expect a bias. If they leave for unrelated reasons, the risk is lower.

  4. Can you get the missing cases?

    Look for exit data, closed-fund records, cancellation reasons, or logs of failed requests.

  5. If not, can you estimate their effect?

    Wald did this by stating assumptions and computing what the lost cases most likely looked like. Write your assumptions next to the result.

  6. Follow a fixed starting group.

    Define a cohort at the start and keep every member in the analysis, including those who left.

Do not ask "what do the winners have in common?" Ask "what did the group look like at the start, and who is no longer in it?"

Two practical habits help. When you read a list of successes, ask for the list of attempts it was taken from. When you build a metric, write down who can leave the denominator and what happens to the number when they do.

081 min

Where this goes wrong

Treating the missing cases as random. The bias exists only when removal is linked to the outcome. Check this before applying the correction. Wald's method assumes aircraft were lost to enemy fire. Losses from mechanical failure would not fit that assumption.

Counting only the visible damage. The popular story says to put armour where returning aircraft show no damage. The documented work is narrower: it estimates the chance of surviving each hit.

Fixing the problem by adding more survivors. A larger sample of survivors has the same bias. Ten thousand surviving funds are no better than a thousand if closed funds are absent.

Correcting too far. Not every dropout is informative. Some users leave because they moved away or changed jobs. Assuming every leaver was unhappy replaces one bias with another.

Assuming the bias is a fixed size. In the fund data it grows with the length of the sample. Carhart and colleagues also note that it is possible to build examples where it does not grow with sample length. Measure it for your own data when you can.

Believing a well-known story without a source. The aircraft story is the best known example of this bias, and the details people repeat are often not in the record. The next sections set out what is documented.

091 min

Survivorship bias vs. nearby concepts

Compared withThe fact that separates them
Not in the library yetSelection biasSelection bias is the broad family: any process that makes the sample differ from the group. Survivorship bias is the case where the process is survival over time.
PsychologyConfirmation biasConfirmation bias is how a person reads evidence. Survivorship bias is which evidence exists. A careful reader with no preference can still be misled by it.
PsychologyHindsight biasHindsight bias is believing after the fact that an outcome was predictable. Survivorship bias is missing data about outcomes that did not happen.
PsychologyNarrative fallacyNarrative fallacy is building a neat cause-and-effect story from events. Survivorship bias supplies one-sided events for that story to be built from.

The deciding question: did the cases that failed ever have a chance to enter the data? If they did not, survivorship bias is likely to be present.

102 min

Where the evidence is contested

The first contested point is the aircraft story itself. Bill Casselman, writing for the American Mathematical Society in 2016, argues that the popular version is mostly a reconstruction. Almost everything repeated online about what Wald told the military is a reconstruction, and the documentary record is thin. He writes that there is extremely little source material for what Wald said about aircraft damage. What survives is Wald's own memoranda and two short, vague mentions in W. Allen Wallis's memoir of the Statistical Research Group.

The memoranda are technical. The memoranda say nothing about what the military should do to improve the aircraft. The statement that Wald told the military to armour the engines, the unmarked parts, therefore is not supported by the documents Casselman could find. The statistical idea is real and well documented. The dramatic scene around it is not.

Casselman also gives a fair reading of why people like the story: it is a clear example of a real problem. His objection is to the added detail, not to the principle.

The second contested point is how large and how stable the bias is in finance. Carhart and colleagues report an attrition rate of 3.6 percent of funds in an average year in their sample, while Elton, Gruber and Blake found 2.3 percent. The two groups used different samples. Carhart and colleagues point out that Elton, Gruber and Blake studied a single cohort of funds, so each later year required funds to have survived.

Where this leaves a practitioner: the principle is not in doubt, and the size of the effect depends on the data. Use the principle to ask what is missing, use documented numbers only with their source and sample, and avoid repeating the aircraft anecdote as proven history.

?8 questions

Questions people ask

What is survivorship bias?

Survivorship bias is a sampling error where you study only the cases that passed a selection process and ignore the ones that dropped out. Because survivors differ from the full group, conclusions drawn from them are usually too optimistic.

What is an example of survivorship bias?

Abraham Wald's work on World War II aircraft is the standard case. Damage was recorded only on aircraft that returned, so the data hid the aircraft that did not. Wald estimated survival chances per hit instead of trusting the visible damage.

What is survivorship bias in investing?

It is the error of measuring fund performance using only funds that still exist. Funds usually close after poor results, so survivors look better. One study measured the bias at 17 basis points a year for one-year samples and about one percent for samples longer than fifteen years.

How do you avoid survivorship bias?

Start from the full group that entered the process and keep every member in the analysis, including those who left. Then look for data on the leavers, or state assumptions about them and test how much the result changes.

What is the difference between survivorship bias and selection bias?

Selection bias covers any process that makes a sample differ from the group you care about. Survivorship bias is one case of it, where the process is survival over time, such as customers cancelling or funds closing.

Is the Abraham Wald aircraft story true?

The statistical work is documented in his memoranda, but many details told online are not. Casselman, writing in 2016, says almost everything beyond the memoranda and two short mentions in a memoir is plausible reconstruction.

How can survivorship bias affect a product team?

It can mislead a team that analyses only active users. Surveys, average-session charts and A/B tests that exclude users who left all show better results than the full group would.

Does a larger sample fix survivorship bias?

No. A larger sample of survivors has the same bias, because the cases that left are still missing. The fix is to recover or estimate those cases, not to collect more survivors.

Β§6 sources

Sources

  1. Mangel, M. and Samaniego, F. J. (1984). Abraham Wald's Work on Aircraft Survivability. Journal of the American Statistical Association 79(386), 259-267

  2. Casselman, B. (2016). The Legend of Abraham Wald. American Mathematical Society Feature Column

  3. Carhart, M. M., Carpenter, J. N., Lynch, A. W. and Musto, D. K. (2002). Mutual Fund Survivorship. Review of Financial Studies 15(5), 1439-1463 (working paper version, 2000)

  4. Brown, S. J., Goetzmann, W., Ibbotson, R. G. and Ross, S. A. (1992). Survivorship Bias in Performance Studies. Review of Financial Studies 5(4), 553-580

Show all 6 sources
  1. Elton, E. J., Gruber, M. J. and Blake, C. R. (1996). Survivorship Bias and Mutual Fund Performance. Review of Financial Studies 9(4), 1097-1120

  2. Wald's original memoranda were reissued in 1980 by the Center for Naval Analyses as A Method of Estimating Plane Vulnerability Based on Damage of Survivors. The copy online could not be fetched for this entry, so the method here is taken from the Mangel and Samaniego paper and the Casselman column that works through it. Full text of the 1992 and 1996 papers is behind a subscription, so those two are cited through their published abstracts.

Keep reading

More from Research

All of Research
All of Research