012 min
Where the planning fallacy shows up
Consider a staff engineer in sprint planning, asked how long a payment-retry feature will take. This scenario is illustrative. They think through the work: two days on the queue, one day on the retry policy, one day of tests. They say four days. They do not mention the three features they estimated last quarter, each of which took about twice as long as they said it would. The feature ships on day nine.
What to watch for is how the answer was built, not how large it is. The estimate is assembled forward, step by step, from a plan for this specific piece of work. Nothing in it refers to how similar work has gone before.
Now run the same moment again and change one input. The release date is four weeks away instead of one. The estimate goes up.
Buehler, Griffin and Ross measured exactly this in 1994. Students given a one-week deadline for a computer assignment predicted 3.9 days on average. Students given a two-week deadline predicted 7.9 days. The deadline also changed behaviour, not only prediction: the one-week group finished in 4.3 days, the two-week group in 9.1 days. Both groups still finished later than they had said they would.
That second result is the useful one. A distant deadline does not remove the error. It makes the error larger. The gap between what you say and what happens grows in proportion to the time available.
021 min
Where the planning fallacy comes from
Daniel Kahneman and Amos Tversky named the effect in a 1979 technical report for the United States Office of Naval Research, Intuitive Prediction: Biases and Corrective Procedures.
Their definition was narrow. The planning fallacy, they wrote, "is a consequence of the tendency to neglect distributional data, and to adopt what may be termed an 'internal approach' to prediction." Distributional data means the record of how a set of similar cases turned out.
They named writers and scientists as the example. Such people are, in the report's words, "notoriously prone to underestimate the time required to complete a project." That holds "even when they have considerable experience of past failures to live up to planned schedules."
The first experiments came fifteen years later. Roger Buehler, Dale Griffin and Michael Ross telephoned 37 psychology students at the University of Waterloo who were in the final semester of an honours thesis course. Each was asked to predict, as accurately as they could, the day they would hand the thesis in. The course coordinator recorded the real submission dates.
The students' average best estimate was 33.9 days. They took 55.5 days. Only 29.7% handed in by the date they had given as their most accurate prediction.
032 min
Why the planning fallacy happens
The cause is a choice of evidence, made without the person noticing that a choice existed.
Kahneman and Tversky separated two kinds of information. "Singular information, or case data, consists of evidence about the particular case under consideration. Distributional information, or base-rate data, consists of knowledge about the distribution of outcomes in similar situations." A plan is singular information. The record of your last ten projects is distributional information. People predict almost entirely from the first kind.
What people actually think about while estimating
Buehler and colleagues asked participants to write down their thoughts as they produced an estimate. In their fifth study, 92.7% of the people estimating their own completion time wrote about future plans. 2.4% mentioned a past success. Not one of them, 0%, mentioned a past problem. The past was not considered and rejected. It did not come up at all.
Why the past gets disqualified
People do remember being late. What they do with the memory is decide it does not apply here. Across these studies, participants described their own previous delays as caused by factors that were external, temporary, and specific to that one occasion: a supervisor was slow to comment, a machine failed, one particular week was unusually busy. Each explanation is often true. Taken together they make the last occasion look like a special case with nothing to say about this one. The same participants explained a colleague's lateness in more stable, personal terms.
This account rests on what people reported thinking, which is weaker evidence than behaviour. The authors said so themselves, writing that they "doubt that they provide a complete or fully accurate rendering of cognitive process."
041 min
Why the planning fallacy shrinks for observers
The error is specific to your own plans. In the fifth Buehler study, 123 undergraduates were each paired with a student from an earlier experiment. Each was asked to predict when that person would finish a computer assignment.
The people doing the work predicted 5.5 days on average. The observers predicted 8.5 days. The real average was 6.8 days.
The observers were not correct. They were incorrect in the other direction. The people doing the work were short by 1.3 days. The observers were long by 1.7 days. Measured as absolute error, ignoring direction, the two groups were level: 1.8 days against 2.7 days, a difference that did not reach statistical significance.
The written thoughts show why the direction flipped. Observers mentioned things that might go wrong about four times as often as the people doing the work did. One group was building a plan. The other group was judging whether a plan was credible. Those are different tasks, and they bring different evidence to mind.
The practical reading is narrow. Asking a colleague to estimate your work does not produce a correct number. It produces a number that is wrong in a direction you can anticipate. In this study the real completion time fell between the two figures, which is the argument for collecting both rather than trusting either.
052 min
The planning fallacy in large public projects
The effect is not confined to students and small tasks.
Bent Flyvbjerg built a database of large transport projects and measured how far each cost forecast sat from what was eventually spent, in constant prices. Rail projects were out by an average of 44.7%. Bridges and tunnels by 33.8%. Roads by 20.4%. Across the seventy years for which cost data exist, that accuracy did not improve (Project Management Journal, 2006).
A later dataset of 2,062 capital investment projects across eight types shows the same direction. Measured as actual cost divided by estimated cost, dams averaged 1.96, rail 1.40, tunnels and power plants 1.36, roads 1.24 (Project Management Journal, 2021).
Sydney Opera House
Buehler, Griffin and Ross open their 1994 paper with it. The 1957 estimate was completion in early 1963 at a cost of $7 million. A reduced version of the building opened in 1973 at $102 million.
Edinburgh, where somebody used the outside view first
In October 2004 the consultancy Ove Arup and Partners Scotland reviewed the business case for Edinburgh Tram Line 2. The promoter had put the base cost at £255 million and added £64 million, or 25%, for contingency and optimism, giving roughly £320 million in total.
Arup applied the UK Department for Transport's published uplifts for that class of project. They reported an 80th percentile cost of £400 million. That is the figure carrying an 80% chance of staying inside budget. Arup concluded that the promoter's estimate was optimistic.
Line 2 was never built. Line 1a was, and it is a separate line of the same programme. The Edinburgh Tram Inquiry reported in 2023 that Line 1a was approved with a £545 million budget, to open in summer 2011. It opened on 31 May 2014 at a cost of £776.7 million. The route was shorter than the one that had been approved.
062 min
How the planning fallacy shows up in software
Two situations, one for people building and one for people measuring.
A team estimates a migration. The inside view counts the services to move and multiplies by a per-service figure. The outside view asks what the last three service moves actually took, and starts from that. Both numbers are available in the same repository. Only the first is usually used.
An analyst plans an evaluation run for a model. They need a labelled test set, and they estimate from the labelling rate observed on a clean sample. The estimate leaves out what consumed most of the last two evaluation sets: arguing about what the labels mean. The labelling rate is singular information about this set. The two previous sets are the distribution.
The numbers, with their owners
Kjetil Moløkken-Østvold and Magne Jørgensen reviewed the survey literature on software estimation. Between 60% and 80% of projects overran on effort, schedule, or both, and the average effort overrun was 30% to 40% (Moløkken-Østvold, PhD thesis, University of Oslo, 2004).
That range matters because a far larger figure circulates more widely. The Standish Group's 1994 CHAOS report gave an average cost overrun of 189% for the projects it called challenged. Jørgensen and Moløkken-Østvold examined that figure and argued it is inconsistent with the results of every other cost accuracy survey and is probably far too high. Before quoting an overrun statistic to justify a schedule, check which of the two figures it is.
Where a low estimate is not an error
Sometimes the number is low on purpose. Flyvbjerg separates the planning fallacy from strategic misrepresentation. The planning fallacy is unintentional. Strategic misrepresentation is "the tendency to deliberately and systematically distort or misstate information for strategic purposes", and it is deliberate. A supplier quoting a delivery date it does not believe, in order to win a contract, is doing the second thing. Techniques that correct estimates do not address it. Holding forecasters accountable for their forecasts does.
072 min
How to guard against the planning fallacy
Awareness is not one of the techniques, and there is a direct test of that.
In Buehler's fourth study, one group was asked to remember and describe their past experience with similar assignments immediately before estimating. That group predicted 5.3 days against a control group's 5.5, and underestimated by almost exactly as much. Remembering, on its own, did nothing.
A second group described the same past experiences and then answered further questions that made them connect those experiences to the assignment in front of them. That group predicted 7.0 days and took 7.0 days. 60.0% finished inside their own estimate, against 29.3% of the control group.
The instruction that follows is to force the link, not the recall. A memory that is not attached to the current prediction has no effect on it.
Reference class forecasting
Flyvbjerg turned the same idea into a procedure an organisation can run:
Identify a class of past, similar projects, broad enough to be statistically meaningful and narrow enough to be truly comparable. Establish the distribution of outcomes for that class. Place the project in that distribution to get its most likely outcome.
The UK Treasury's 2003 Green Book told appraisers of large public projects to correct for this. Its wording was "explicit, empirically based adjustments to the estimates of a project's costs, benefits, and duration". In the summer of 2004 the Department for Transport and the Treasury adopted the method for large transport schemes. The American Planning Association endorsed it in April 2005.
Unpack before estimating, not after
Justin Kruger and Matt Evans asked people to estimate everyday multi-part tasks such as preparing food and formatting a document. Half listed the component steps before giving a number, half listed them afterwards. Listing first cut the size of the underestimate by more than half (Journal of Experimental Social Psychology, 2004).
One caveat belongs with all of this. In Buehler's study the corrected group was less biased but no more accurate in absolute terms. Their average absolute error was 1.9 days, against the control group's 1.8. Removing a bias that always runs in one direction is not the same as making an estimate precise.
082 min
Where the planning fallacy evidence is contested
The effect replicates well for long, familiar, multi-step work. Two lines of criticism narrow it.
The first disputes the explanation. Michael Roy, Nicholas Christenfeld and Craig McKenzie argued in Psychological Bulletin in 2005 that people are not misusing their memories of how long things took. Those memories are themselves too short. On this account the error sits in recall rather than in planning. Their evidence has three parts.
- The underestimation disappears when the task is new to the person.
- The same bias appears when people estimate how long something in the past took.
- Anything that changes memory for duration changes prediction in the same way.
Griffin and Buehler replied in the same journal that year, arguing that the two accounts cover different domains and that comparing them head to head is misleading. This is unresolved.
The second disputes the range. Torleif Halkjelsvik and Magne Jørgensen reviewed the literature on judgment-based time prediction in Psychological Bulletin in 2012. Underestimation was reported more often than overestimation in studies from engineering and management. In studies from psychology it was not. Underestimation is therefore not a general law of time prediction. It is what happens under particular conditions, and long project work is one of them.
Flyvbjerg, who relies on the concept, names the same limit on the laboratory evidence: it is "mostly from simple laboratory experiments with students," and how far it carries into real project planning is an open question.
While this stays unsettled, the practical position is firm for project work and weak elsewhere. Add an uplift to a schedule for work your team has done versions of before. Do not assume the same bias on a task nobody has attempted.
092 min
Planning fallacy vs. nearby concepts
Four concepts are close to it and often get used in its place.
| Concept | The axis that separates it |
|---|---|
| Optimism bias | Scope. It covers expecting better outcomes of every kind, including health and accidents. Flyvbjerg treats the planning fallacy as the part of it that concerns cost, schedule and benefit. |
| Uniqueness bias | Role. It is a cause, not a synonym: the tendency to see your project as more singular than it is, which is what makes the record of comparable projects look irrelevant. |
| Strategic misrepresentation | Intent. A low estimate given deliberately to win approval is not a cognitive error, and its cure is accountability rather than better forecasting. |
| Escalation of commitment | Timing. It governs whether you keep funding something because of what you already spent. The planning fallacy happens before the money is committed. |
Those four definitions are taken from Flyvbjerg's 2021 overview of behavioural biases in project management.
The pair worth keeping apart is the first and the third, because from outside a low estimate looks the same either way. The test is whether the forecaster believes their own number. Optimism and the planning fallacy are unreflected, which is Flyvbjerg's word for a distortion the person cannot see in themselves. When forecasters are surveyed about why their forecasts were wrong, they name scope changes, complexity, price changes and weather. They do not name their own optimism.
102 min
Common misunderstandings about the planning fallacy
"Just add a buffer." Buehler's first study also asked the thesis students for a worst case: what they would predict "if everything went as poorly as it possibly could." Their average worst case was 48.6 days. They took 55.5. Fewer than half of them, 48.7%, met even that date. Padding moved the figure without repairing it, and it did not improve accuracy at all. Average absolute error was 23.2 days for the worst case against 22.6 days for the plain best guess.
"So the estimate is worthless." No. In that same study, predicted and actual times correlated at r = .77. The students who said they would take longer did take longer. The ranking carries information. The level carries the bias. That split is why an uplift works at all, because an uplift corrects a level while leaving the ranking intact.
"It is a discipline problem." The people in the founding studies were finishing an honours thesis. The projects in Flyvbjerg's database were forecast by professionals, over seven decades, and the accuracy did not improve across them.
"The definition always covered cost." It does now. It did not originally. Kahneman and Tversky's 1979 term was about task completion times. Flyvbjerg calls the broadened version "the planning fallacy writ large", meaning underestimated cost, schedule and risk alongside overestimated benefit. Both meanings are still in circulation, so it is worth checking which one a given source is using.
?8 questions
Questions people ask
What causes the planning fallacy?
What is an example of the planning fallacy in software work?
Does adding a buffer fix the planning fallacy?
How do you avoid the planning fallacy?
What is the difference between the planning fallacy and optimism bias?
Is the planning fallacy the same as lying about a deadline?
Is the planning fallacy real, or has it failed to replicate?
Are other people better at estimating my work than I am?
§10 sources
Sources on the planning fallacy
Kahneman, D. and Tversky, A. (1979). Intuitive Prediction: Biases and Corrective Procedures. Technical report, Office of Naval Research.
Buehler, R., Griffin, D. and Ross, M. (1994). "Exploring the 'planning fallacy': Why people underestimate their task completion times." Journal of Personality and Social Psychology, 67(3), 366-381.
Flyvbjerg, B. (2006). "From Nobel Prize to Project Management: Getting Risks Right." Project Management Journal, 37(3), 5-15.
Flyvbjerg, B. (2021). "Top Ten Behavioral Biases in Project Management: An Overview." Project Management Journal, 52(6), 531-546.
Show all 10 sourcesShow fewer sources
Roy, M. M., Christenfeld, N. J. S. and McKenzie, C. R. M. (2005). "Underestimating the duration of future events: Memory incorrectly used or memory bias?" Psychological Bulletin, 131(5), 738-756.
Halkjelsvik, T. and Jørgensen, M. (2012). "From origami to software development: A review of studies on judgment-based predictions of performance time." Psychological Bulletin, 138(2), 238-271.
Moløkken-Østvold, K. J. (2004). Effort and Schedule Estimation of Software Development Projects. PhD thesis, University of Oslo.
Kruger, J. and Evans, M. (2004). "If you don't want to be late, enumerate: Unpacking reduces the planning fallacy." Journal of Experimental Social Psychology, 40(5), 586-598.
Edinburgh Tram Inquiry (2023). Report of the Edinburgh Tram Inquiry.
HM Treasury. The Green Book: appraisal and evaluation in central government.




