012 min
A false missile alert in Hawaii
On 13 January 2018, a warning officer at the Hawaii Emergency Management Agency sent a message to phones across the state. It said a ballistic missile was inbound and that people should seek shelter. No missile had been launched. The false alert went out at 8:07 a.m. The agency issued a correction message at 8:45 a.m., 38 minutes after the false alert.
The United States Federal Communications Commission (FCC) investigated and published a report in April 2018. The false alert began with a drill. At a shift change, a supervisor phoned the officers on duty and played a recording that was meant to start a no-notice exercise. The recording had the drill wording, but it also held the text of a real alert, including a line saying this was not a drill. The officer at the alert terminal believed the threat was real.
That belief explains why the officer sent an alert. It does not explain why sending one took so little effort. The officer chose a message from a drop-down list. The list held a test version of the missile alert beside the live one. The two options sat together on one screen. Live alerts and internal test alerts were sent from the same interface, with the same log-in. After the officer chose the live message, the software asked one question: was the officer sure they wanted to send it? The question read the same for a test as for a real alert, and it did not let the officer review the message that would be sent.
Nobody at the agency set out to build an alert system that is easy to trigger by mistake. Studying how screens like this one fail, and how to build screens that do not, is the work of human-computer interaction.
022 min
What human-computer interaction covers
The Hawaii case involves a list, a button and a question on a screen, yet the field covers much more than screens. Its scope was set in a 1992 curriculum report from ACM SIGCHI, the professional group for the subject. The 1992 report gives the field three jobs: to design interactive systems for people, to evaluate them, and to implement them. It adds a fourth job, which is to study the main things that happen around those systems.
Four parts are always in play, and the Hawaii case shows all of them:
- The person. A warning officer who had just heard what sounded like an attack order.
- The task. Send an alert to the public only when the threat is real.
- The tool. The alert software, with its list, its prompt and its log-in.
- The setting. A shift change, in the middle of an unannounced drill.
A failure can start in any of the four, and a fix that changes only the tool may miss the others. That is why the field draws on several disciplines at once. Computer science builds the systems. Psychology explains attention, memory and belief. Design shapes how a system looks and responds. Human factors and ergonomics, which began as the study of how to fit work to people, contribute methods for measuring performance and errors. Social science explains what happens when groups use a system together.
The field also goes by other names, including human-machine interaction and computer-human interaction. A 1983 book by Stuart Card, Thomas Moran and Allen Newell made the term common, and its title was The Psychology of Human–Computer Interaction.
032 min
Where human-computer interaction came from
The field exists because computers stopped being used only by the people who built them. Before Sketchpad, most communication with a computer was typed text, and the first change came from a doctoral student.
Sketchpad, 1963
In 1963 Ivan Sutherland finished a doctoral thesis at MIT on a program called Sketchpad. Sketchpad used a light pen, an early forerunner of the mouse, so a person could point at objects on the screen and work with them. The user did not describe a drawing in commands. The user acted on the drawing itself, and this is now called direct manipulation. Designers of the Xerox Star workstation, which reached the market in 1981, later wrote that Sketchpad influenced the Star's interface as a whole.
Personal computers and a new kind of user
Psychologists had been studying human information processing and performance since the 1950s, and engineers had long studied how to fit work to people. These two lines of work met computing when personal computers reached offices and homes in the early 1980s. The ACM SIGCHI report notes that computer sales had become more directly tied to the quality of interfaces than before. A buyer who was not a programmer could not use a machine that only a programmer could operate, so interface quality now affected what a company could sell.
The 1983 book by Card, Moran and Newell applied cognitive psychology to this problem. It treated the person and the computer as one system whose speed and errors could be measured and predicted. The CHI conference and SIGCHI itself grew from the same period.
042 min
The two gulfs between a person and a system
Sketchpad worked because a person could act on the thing they cared about. Many systems do not work that way, and a 1985 paper gave the gap two names. Edwin Hutchins, James Hollan and Donald Norman described two distances in a 1985 paper on direct manipulation. The gulf of execution is the distance between what a person wants to do and the commands the system offers. The gulf of evaluation is the distance between what the system shows and what the person needs to know about its state.
A system narrows the first gulf when its commands and mechanisms match the person's goals, and the second when its display gives a clear picture of what the system is doing. Both gulfs are measured in effort. The wider the gap, the more the person must think about the system instead of the task.
Hawaii had both. The first gulf was short: one choice and one confirmation sent an alert to the whole state. The second was long. The confirmation step did not clearly say whether the alert was a test or a live alert. The screen gave the officer little to check their belief against. Reading a system through these two gulfs, as in gulf of execution and gulf of evaluation, turns a vague complaint such as "the screen was confusing" into two questions a team can answer: could the person do it, and could the person tell what happened?
052 min
How teams practise human-computer interaction
Knowing where the gulfs are is half the work. The other half is finding them before real users do. Teams in the field usually run a loop: study the people and their tasks, build a version, watch real people use it, fix what the watching shows, and do it again.
Watching people use it
The central method is usability testing. A team gives a few people a realistic task and watches where they hesitate, choose wrongly or give up. The point is to see behaviour, because people often cannot describe their own difficulty in advance.
Ten rules of thumb
Reviewers can also check a design against a list without any users. Rolf Molich and Jakob Nielsen first developed such a list in 1990. Nielsen refined it in 1994, using a factor analysis of 249 usability problems, into ten broad rules of thumb called heuristics. Checking a screen against them is called heuristic evaluation. The fifth rule is error prevention. It says to remove error-prone conditions, or to check for them and offer a confirmation option before the person commits to the action.
Slips and mistakes
Nielsen's error prevention heuristic separates slips, which come from inattention, from mistakes, which come from a mismatch between the person's mental model and the design. The distinction tells a team which fix to try. A slip, such as tapping the wrong item in a crowded list, needs constraints and good defaults. A mistake needs a design that matches how people think the system works. The Hawaii officer's choice was deliberate, so a warning about inattention would not have helped. What helps with a deliberate wrong choice is a check that shows what will happen.
Predicting time before building anything
Some methods give numbers without any users. In 1980 Card, Moran and Newell published the keystroke-level model. It estimates how long a task takes by listing each small operation, such as pressing a key or pointing at a target, and adding up the time for each. A team can compare two designs on paper this way.
The model has clear limits. It describes an expert who makes no errors, and the wider family of models it belongs to treats all users as the same. It can say how long a routine task takes. It cannot say anything about a wrong choice like the one in Hawaii.
062 min
What changed in Hawaii, and what did not cause the alert
After the report, Hawaii changed more than its screen. The agency began to require two people for both tests and real alerts, where one credentialed officer had been enough. It wrote templates for correction messages, since it had no procedure or training for correcting a false alert. It also asked its software vendor to make test and live modes look different and to say on the confirmation step whether the alert about to be sent is live.
Why a confirmation step was not enough
The prompt in Hawaii asked a question whose answer was the same for a test and for a real alert. A confirmation step helps only when it says what will happen. The agency's own report suggested wording that names the action, such as asking whether the officer is sure they want to send a real-world ballistic missile alert. The FCC also noted that other alert software uses watermarks, colour coding and separate applications or log-in choices to keep test and live environments apart.
The interface was not the only cause
It would be wrong to say a menu caused the alert. The recording was wrong, the drill came without notice and at a shift change, and the staff had received only basic training on the software. The agency also had no process for correcting a false alert, which added to the delay before the correction. The FCC listed these as separate findings. A person trained in this field looks at the whole situation, which includes the person, the task, the tool and the setting, and not only at the screen.
071 min
How many people to watch
Every method above depends on watching people, which raises the question of how many to watch. Many teams follow a rule of thumb to test with about five people at a time, because early studies found that four or five people uncover most problems. The rule is cheap and works well for early, repeated tests.
It has a known weak point. In 2003 Laura Faulkner tested a time sheet application with 60 people and then drew many random groups from them. In the worst case, a group of five found 55 percent of the problems, and a group of ten found 82 percent. A team that happened to draw a poor group of five would have missed almost half of the problems and not known it.
What should a team do while this is unsettled? Test early with a few people to fix the obvious problems, then test again. Use more people when a missed problem is costly, as it is for an alert system, and when users differ in skill. A reasonable minimum is at least five people from each distinct group of users.
081 min
Human-computer interaction beyond the desktop
The tests and gulfs above were first built around one person at a screen with a keyboard and a mouse. The field no longer stops there. People now speak to phones, touch tablets, wear sensors and put on headsets. Each way of interacting changes what a person can do easily and what they can get wrong, so each one needs its own testing.
Accessibility is part of this work. The same alert can reach a person who can see and hear it and a person who cannot, and the FCC report records that the 38 minutes were particularly stressful for people with disabilities. Designing so that people with different abilities can use a system is covered in accessibility.
Designing for AI
Systems that learn and make predictions raise a new problem: they are sometimes wrong, and their users cannot see why. In 2019 Saleema Amershi and colleagues proposed 18 guidelines for human-AI interaction and tested them with 49 design practitioners against 20 AI products. The first two guidelines ask a system to make clear what it can do and how well it can do it. In the language of the gulfs, they narrow the gulf of evaluation for a system whose behaviour changes over time.
091 min
How the focus of human-computer interaction has widened
As devices spread, the questions changed. In 2007 Steve Harrison, Deborah Tatar and Phoebe Sengers argued that the field works from three paradigms: human factors, classical cognitivism, and one based on how meaning forms in context. Human factors asks how to fit a person and a machine together. Classical cognitivism builds models of the mind and the computer and tries to predict behaviour. The third paradigm asks how people make meaning from a system in their own situation. A study of the Hawaii alert could use each one: timing and error counts, a model of what the officer saw, and an account of the shift change and the drill.
In the same period the focus moved from usability, meaning whether a task can be completed, to experience, meaning everything a person feels and does around a product. Don Norman and Jakob Nielsen wrote in 1998 that user experience covers all aspects of a person's interaction with a company, its services and its products.
101 min
Human-computer interaction vs. user experience and nearby fields
People use the names of neighbouring fields as if they meant the same thing. They differ in the question each one asks and in who usually asks it.
| Field | Main question | Usual setting |
|---|---|---|
| Human-computer interaction | How do people and computer systems affect each other, and why? | Research groups and university courses |
| User experience design | How should this product feel and work across the whole journey? | Product teams in companies |
| User interface design | How should the screens and controls look and respond? | Design teams, often inside a UX team |
| Human factors | How do we fit work and equipment to people's limits? | Engineering, aviation, medicine |
Human-computer interaction is often called the forerunner of user experience design. It mostly concerns research and evidence, while user experience design mostly concerns shipping a product on a deadline. In practice the methods overlap: usability tests, heuristic reviews and task models appear in both.
111 min
Frequently asked questions about human-computer interaction
What do HCI researchers actually do? They run studies with people, build prototypes, create models that predict behaviour, and report results at venues such as the CHI conference. Many also study new devices or the social effects of systems.
Do I need a computer science degree to work in HCI? No. The field draws on computer science, psychology, design and social science, so people arrive from several backgrounds. What matters is skill with research methods and a habit of testing ideas with real users.
Does HCI cover voice assistants and virtual reality? Yes. Any system that a person operates is in scope, including voice, touch, gesture, headsets and wearable sensors. Each input and output method adds its own questions about what people find easy and what they get wrong.
How can a small team start applying HCI? Ask five people to finish one real task with your product while you watch silently. Note where they hesitate or choose wrongly, fix those points, and repeat the test with new people.
?4 questions
Questions people ask
What do HCI researchers actually do?
Do I need a computer science degree to work in HCI?
Does HCI cover voice assistants and virtual reality?
How can a small team start applying HCI?
§13 sources
Sources
ACM SIGCHI, Curricula for Human-Computer Interaction (Hewett et al., 1992), archived copy:
Federal Communications Commission, Report and Recommendations, Hawaii Emergency Management Agency January 13, 2018 False Alert (2018):
Ivan Sutherland, Sketchpad: A Man-Machine Graphical Communication System, University of Cambridge technical report 574:
Hutchins, Hollan and Norman, Direct Manipulation Interfaces, Human-Computer Interaction, 1985:
Show all 13 sourcesShow fewer sources
Card, Moran and Newell, The Keystroke-Level Model for User Performance Time with Interactive Systems, 1980:
Wikipedia, GOMS (a summary of the Card, Moran and Newell models):
Wikipedia, Human-computer interaction (a summary of the field and its history):
Jakob Nielsen, 10 Usability Heuristics for User Interface Design, Nielsen Norman Group:
Jean Fox, The Science of Usability Testing (2015), a review of sample-size studies including Faulkner (2003):
Amershi et al., Guidelines for Human-AI Interaction, CHI 2019:
Harrison, Tatar and Sengers, The Three Paradigms of HCI (2007):
Don Norman and Jakob Nielsen, The Definition of User Experience, Nielsen Norman Group (1998):
Brad Myers, A Brief History of Human-Computer Interaction Technology, interactions, 1998:





