011 min
Heuristic Evaluation at a glance
- What it is: Evaluators independently check an interface against 10 usability principles, then pool results.
- Origin: Nielsen and Molich introduced the method in 1990; Nielsen refined the 10 heuristics in 1994.
- Evaluators needed: Three to five — a figure Nielsen derived from cost-benefit modelling of his own past projects.
- Backfires when: run alone. Experts and real users flag different problems, so it misses roughly half of what user testing finds.
021 min
How Heuristic Evaluation works
Heuristic evaluation asks a small group of evaluators to inspect an interface on their own, before anyone compares notes. Each evaluator goes through the interface at least twice: once to get a general sense of the flow, then a second pass checking individual screens and controls against ten usability principles — Nielsen's heuristics, covering things like whether the system shows its current status and whether a mistake can be undone.
When an evaluator finds something that breaks one of the ten principles, they write it down, name which principle it breaks, and give it a severity rating: how much it would slow someone down or stop them completing a task. This step happens alone and in writing, not in a group discussion, on purpose. Evaluating together first would let one evaluator's opinion set what everyone else looks for, a version of confirmation bias the method is built to avoid.
Only after every evaluator has finished independently does the group meet to combine their lists. No individual evaluator, no matter how experienced, catches every usability issue alone. Different evaluators tend to catch different problems, so the combined list from several evaluators includes more real problems than any one evaluator's list by itself.
031 min
Where Heuristic Evaluation comes from
Jakob Nielsen developed heuristic evaluation together with Rolf Molich in 1990, while both worked as usability consultants. Their original heuristics grew out of guidelines they had already been using in consulting work, checked against interfaces they had actually evaluated rather than written from theory alone.
Four years later, Nielsen rebuilt the list from a different kind of evidence. He collected 249 usability problems from 11 earlier projects, then measured how well each candidate heuristic explained the problems that had actually been found. That analysis produced nine heuristics. A tenth, Help and Documentation, was not part of that statistical result — Nielsen added it back afterward, and the ten heuristics have stayed the same since 1994.
042 min
The 10 heuristics behind Heuristic Evaluation
Nielsen's ten heuristics are broad rules an evaluator checks an interface against, not specific guidelines written for one particular screen. Each one names a property a good interface has, stated as something an evaluator can check for directly.
| # | Heuristic | What it checks |
|---|---|---|
| 1 | Visibility of System Status | Whether the interface tells people what is happening, quickly enough for them to act on it |
| 2 | Match Between the System and the Real World | Whether wording and order follow how people already think and talk, not internal jargon |
| 3 | User Control and Freedom | Whether people can back out of an action or undo it without a lengthy process |
| 4 | Consistency and Standards | Whether the same word, icon, or action means the same thing everywhere, matching Jakob's Law and the conventions of other products |
| 5 | Error Prevention | Whether a mistake is stopped before it happens, not just explained after it happens |
| 6 | Recognition Rather Than Recall | Whether the information someone needs is visible on screen, instead of something they have to recall |
| 7 | Flexibility and Efficiency of Use | Whether shortcuts exist for experienced users without slowing new users down |
| 8 | Aesthetic and Minimalist Design | Whether every visible element earns its place, instead of competing with what matters |
| 9 | Help Users Recognize, Diagnose, and Recover from Errors | Whether an error message states what went wrong in plain language, with a way to fix it |
| 10 | Help and Documentation | Whether help exists for the parts of the interface that still need explaining, and is easy to search |
The same ten checks apply to almost any interface. None of them names a specific screen or technology, so they work as well for a mobile app as for a spreadsheet or a physical kiosk.
052 min
How many evaluators a Heuristic Evaluation needs
Nielsen recommends three to five evaluators. That number comes from his own cost-benefit modelling of past projects, not from an outside study measuring the same question on different projects.
The case for more than one evaluator starts with how little a single evaluator finds alone. Averaged across six of Nielsen's own projects, single evaluators found only 35% of the usability problems in the interfaces they reviewed. Individual find rates in those six projects ranged from 19% to 51%, averaging 34%, on interfaces that held 16 to 50 total problems, averaging 33. Aggregating several independent evaluators fixes this, because different evaluators tend to catch different problems.
The case for stopping around five, rather than adding more, is about cost. Based on figures from his own projects, Nielsen estimated a heuristic evaluation's fixed cost at $3,700 to $4,800, plus $410 to $900 per additional evaluator. Applying those numbers to one sample project, a fourth evaluator brought the total cost to $6,400 and found usability problems Nielsen valued at $395,000. Past that point, each additional evaluator kept costing the same amount while finding progressively less that the others had not already found.
Both halves of this argument are Nielsen's own estimates from his own consulting work. Nobody has independently repeated the cost-benefit study on a different set of projects, so three to five is best read as Nielsen's own calculation of where the return stays high, not a number an outside study reproduced.
061 min
How Heuristic Evaluation shows up in practice
The clearest evidence for what a heuristic evaluation actually turns up comes from Nielsen's own record of running the method. Across six real interfaces in Nielsen's own case studies, heuristic evaluation identified a total of 59 major usability problems and 152 minor usability problems. Most of what shows up is minor: a label that reads unclear out of context, an icon two people would read two different ways, a confirmation step where none was needed. Problems serious enough to stop someone completing a task showed up too, just in smaller numbers, which is why every problem found gets a severity rating instead of being treated as equally urgent.
In a product team, a heuristic evaluation usually runs before a design ships and before a usability test is scheduled, not instead of one. A designer and an engineer check a new feature against the ten heuristics before it goes live, fix what an expert can catch on sight, and still set aside time to watch someone unfamiliar with the product try to use it.
071 min
How to run a Heuristic Evaluation
Define the scope.
Decide which flows or screens the evaluation covers. A full evaluation of every screen in a large product takes too long to stay useful.
Brief the evaluators.
Walk them through the ten heuristics and the interface's intended users and main tasks, without telling them what to look for.
Send evaluators through alone.
Each person goes through the interface at least twice: once for a general feel, once checking specific elements against each heuristic.
Record problems as they're found.
Tag each one with the heuristic it breaks and a severity rating, and don't discuss it with the other evaluators yet.
Combine the lists.
Once everyone has finished independently, merge the lists and talk through any disagreement about whether something is really a problem.
Prioritize by severity
, not by how many evaluators happened to notice the same thing. A problem only one evaluator caught can still be the most serious one on the list.
082 min
Where Heuristic Evaluation falls short
Heuristic evaluation and usability testing do not reach the same problems by two different routes. They find different problems. An evaluator inspects a design against a checklist, using their own mental models of how an interface should behave. A user in a usability test is not inspecting anything. They are trying to finish a task, and a problem only shows up if it actually stops them partway through. An expert looking at a screen and a person trying to get something done are doing two different things, so heuristic evaluation and usability testing end up surfacing two different sets of failures.
One direction of that gap is false positives. Sauro's 2012 review of five published comparisons found that on average, 34% of the problems evaluators flagged in a heuristic evaluation were never seen when real users were watched doing the same tasks. Evaluators called these problems. The people actually using the product never ran into them.
The other direction is misses. The same review found heuristic evaluation turns up only around 36% of the problems a usability test finds. It misses roughly 49% of what watching real users uncovers. None of the evaluators were the person actually trying to complete the task, so those problems never came up during the evaluation.
Neither number makes one method better than the other on its own. It means a design reviewed only by experts still carries every failure that only shows up once someone tries to use it for real.
092 min
Heuristic Evaluation vs. usability testing
Both methods exist to catch problems before a design ships. What separates them is who does the judging and how.
| Heuristic Evaluation | Usability testing | |
|---|---|---|
| Who judges the interface | A small group of UX experts | People who match the product's real users |
| What they do | Check the interface against a fixed list of principles | Try to complete real tasks with the interface |
| What it costs | Low — no recruiting, no sessions to schedule | Higher — needs participants, a script, and time to run |
| What it finds well | Problems visible to someone who already knows what to look for | Problems that only appear once someone is actually trying to get something done |
| What it tends to miss | Problems that only show up mid-task, under real conditions | Rare problems the test's small number of participants happen not to hit |
The two methods do different jobs, not the same job twice. A heuristic evaluation is cheap enough to run early and often, catching the kind of problem an expert can spot on sight: a missing status indicator, an inconsistent label, a confirmation step with no way to undo it. A usability test is the only way to see whether people can actually use the thing, because it is the only one of the two where an actual user takes part.
102 min
Frequently asked questions about Heuristic Evaluation
What is heuristic evaluation in UX?
Heuristic evaluation is a usability inspection method where a small group of evaluators independently checks an interface against ten established usability principles, then combines what they each found. Jakob Nielsen and Rolf Molich introduced it in 1990.
What are the 10 usability heuristics?
Nielsen's ten heuristics cover visibility of system status, matching the real world, user control and freedom, consistency and standards, error prevention, recognition rather than recall, flexibility and efficiency, minimalist design, error recovery, and help and documentation. Nine came from a 1994 factor analysis; the tenth was added back afterward.
How many evaluators does a heuristic evaluation need?
Three to five. Nielsen calculated this from his own cost-benefit modelling of past projects — a single evaluator finds only about 35% of the problems present, and the cost of adding evaluators outpaces the benefit past about five.
How is heuristic evaluation different from usability testing?
An evaluator checks a design against a fixed list of principles. A usability test watches a real person try to complete a task. The two methods surface different problems because they are different activities, not two routes to the same result.
Can heuristic evaluation replace usability testing?
No. On average, 34% of the problems a heuristic evaluation flags turn out to be false positives users never hit, and it misses roughly half of what a usability test finds. The two methods are meant to run together, not instead of each other.
Who created heuristic evaluation?
Jakob Nielsen and Rolf Molich introduced the method in 1990, based on usability guidelines they had already been using as consultants. Nielsen refined the ten heuristics still used today in 1994.
What is an example of heuristic evaluation finding a problem?
In Nielsen's own case studies of six interfaces, the method found 59 major usability problems and 152 minor ones, most of them small issues like unclear labels or inconsistent icons, not failures that stopped a task outright.
?7 questions
Questions people ask
What is heuristic evaluation in UX?
What are the 10 usability heuristics?
How many evaluators does a heuristic evaluation need?
How is heuristic evaluation different from usability testing?
Can heuristic evaluation replace usability testing?
Who created heuristic evaluation?
What is an example of heuristic evaluation finding a problem?
§9 sources
Sources
Nielsen, J., & Molich, R. (1990). Heuristic Evaluation of User Interfaces. Proc. ACM CHI'90 (Seattle, WA), 249–256.
Nielsen, J. (1994). Enhancing the Explanatory Power of Usability Heuristics. Proc. ACM CHI'94 (Boston, MA), 152–158.
Nielsen, J., & Landauer, T. K. (1993). A Mathematical Model of the Finding of Usability Problems. Proc. ACM INTERACT'93 and CHI'93 (Amsterdam), 206–213.
Nielsen, J. (1994). 10 Usability Heuristics for User Interface Design. Nielsen Norman Group.
Show all 9 sourcesShow fewer sources
Moran, K., & Gordon, K. (2023). How to Conduct a Heuristic Evaluation. Nielsen Norman Group.
Nielsen, J. (1994). The Theory Behind Heuristic Evaluations. Nielsen Norman Group.
Nielsen, J. (1995). Usability Problems Found by Heuristic Evaluation. Nielsen Norman Group.
Sauro, J. (2012). How Effective are Heuristic Evaluations? MeasuringU.
Heuristic evaluation. Wikipedia.



