011 min
Why the observer-expectancy effect matters
A product team often tests its own design. The person who built the prototype runs the sessions, writes down what happened, and writes the report. That person knows which result the team hopes for. Every condition for the observer-expectancy effect is already present.
The cost is a wrong decision backed by good-looking evidence. A report says most participants finished the task. The team ships the design. Later, the real completion rate turns out lower, because the sessions were run in a way that helped participants succeed. Nobody lied, and the report cannot show the problem, because the problem is in how the sessions were run and not in the numbers that were written down.
The same risk applies to a founder who interviews ten customers about their own idea, an engineer who labels test data for a model they built, and a researcher who reads open-ended answers after reading the hypothesis. In each case a person who wants one answer is the one collecting the answers.
022 min
How the observer-expectancy effect works
The effect works through two routes. One expectation drives both.
Route 1: the researcher changes what the participant does
People and animals respond to small cues. Examples are the tone of a question, a nod, a smile, how long the researcher waits before speaking, and how the researcher sits. A researcher who expects one outcome gives slightly different cues to participants in different conditions, usually without noticing. The participants respond to those cues. The data then shows a difference that the cues produced and the thing being tested did not.
Route 2: the researcher changes how the data is read
Many measurements need a judgment. Did the participant hesitate? Is this answer a failure or a near success? Does this comment count as a complaint? When the researcher knows which condition the data came from, they tend to make these calls in the direction they expected. This is confirmation bias, the tendency to look for information that supports a hypothesis and to overlook information that conflicts with it.
Both routes are unintentional. A researcher who sincerely wants an accurate answer can still produce a biased one. For that reason the standard control does not depend on good intentions. It removes the researcher's knowledge of who is in which condition.
The effect is a form of reactivity, which means that the act of studying something changes it. It is a significant threat to internal validity, which is how far a study's result was caused by the thing the study set out to test. The usual control is a double-blind design. In a double-blind design, neither the participants nor the people collecting the data know who is in which condition.
032 min
Clever Hans: the first documented case
Clever Hans was a horse in Germany in the early twentieth century. His owner, Mr. von Osten, said the horse could do arithmetic and answered by tapping a hoof. Pfungst's report lists this as one of the questions Hans answered: "If the eighth day of a month comes on Tuesday, what is the date for the following Friday?"
In September 1904 a commission examined the horse and found no use of tricks. The psychologist Oskar Pfungst, who worked with the philosopher and psychologist Carl Stumpf, then tested what the commission had not: whether the questioner was giving signals without knowing it. Pfungst published his results in 1907. The English translation by Carl L. Rahn appeared in 1911.
Pfungst ran four kinds of test.
- Hiding the questioner. Pfungst fitted blinders so that Hans could not see the person asking. In Pfungst's tally, 6% of Hans's answers were correct when the questioner was certainly not seen, and 89% were correct when the questioner was in sight.
- Questioners who did not know the answer. Hans failed the tests in which nobody present knew the answer.
- Pfungst as the horse. When Pfungst played the part of the horse, he learned to read small head movements that other people made without knowing it.
- Finding the cue. When the questioner gave the problem, they bent their head and trunk slightly forward, and the horse began to tap. When the horse had tapped the number the questioner wanted, the questioner made a slight upward jerk of the head, and the horse stopped.
Pfungst also tested himself as a questioner. When Pfungst counted the taps without knowing when the desired number was reached, the responses were always incorrect.
The case is useful because nobody needed to be dishonest. The questioner gave the cue involuntarily, and the horse was attentive. The error was in a person who knew the answer, standing in front of a subject that could respond to small movements. This is the pattern that Rosenthal later tested on purpose.
042 min
The Rosenthal and Fode rat experiment
In 1963 the psychologists Robert Rosenthal and Kermit Fode tested the idea on purpose. They wrote that the problem had been generally recognized and much discussed. They also wrote that there had been no systematic test of whether an experimenter can obtain from subjects the data the experimenter expects or wants to obtain.
The design was simple. Twelve of the thirteen students enrolled in a senior course in experimental psychology ran the rats. Each student received a group of rats to train in a maze. The students were told that some groups were "maze-bright", bred to learn mazes quickly, and that others were "maze-dull". Both labels were false. The rats in the two groups were the same kind of rat, and the labels had been assigned at random.
The students who had been told their rats were maze-bright reported better results than the students who had been told their rats were maze-dull. The rats did not differ. What differed was what the handlers expected. The paper's authors concluded that the handlers' expectations changed what the handlers did, and that this changed what the rats did.
Two limits apply to this result. The study used twelve student handlers with no previous experience of animal experiments. It is not a source of an effect size to quote for other settings. And the finding is narrow. It does not show that expectations always change outcomes. It shows that when the person collecting data knows the hypothesis, the data can follow the hypothesis.
The authors drew a practical rule from their work. Whenever possible, the people who run the subjects should not know what outcome is wanted. Almost every method in the next sections is a version of that rule.
052 min
How the observer-expectancy effect shows up in product, design and engineering work
An illustrative case. A UX researcher is running usability tests on a redesigned checkout. They believe the new design is better. In each session with the new design they say "great, go ahead" in a warm tone when the participant finds the next step, and they wait a little longer in silence when the participant hesitates. Participants finish the tasks at a higher rate than the team expected, and the report says the redesign works.
A second researcher then runs the same sessions without being told which design is new. This researcher says the same neutral phrase after every task. The completion rate is lower, and the gap between the designs is smaller. The participants in both rounds saw the same screens. The difference is what the first researcher's expectations did to the sessions: the warmth, the pauses, and which hesitations were counted as failures.
The same pattern appears wherever the person who knows the hypothesis also collects the data:
- Interviews. The interviewer asks follow-up questions after the answers they hoped for and moves on after the answers they did not.
- Usability tests. The facilitator helps, hints or waits, as in the case above.
- Data labelling. The person who labels examples for a machine learning model knows which group each example belongs to and settles unclear cases in the direction of their expectation.
- Experiments with a human scorer. Any score that depends on a person's judgment, such as "was this answer helpful", can move toward the scorer's expectation.
A result that software measures without a person's judgment is protected against both routes during collection. A click count in an A/B test is an example. The people who design the test, choose the metric and decide which results to report can still be affected, so the protection covers collection and nothing else.
061 min
How to reduce the observer-expectancy effect
Each step removes a chance for the researcher to act on an expectation. They follow from the mechanism above.
Blind the person who collects the data.
The most direct fix is a double-blind design. If that is not possible, use a person who does not know the hypothesis to run the sessions.
Script the session.
Use the same words, the same pauses and the same reaction to every participant, whatever they do.
Decide how to score before looking.
Write down what counts as a success or a failure before the first session, so that the rule is not chosen after seeing who is in which condition.
Separate the person who runs the sessions from the person who analyses them.
The second person reads the recordings or the raw data without knowing the group labels.
Prefer measures that do not need judgment.
Timed task completion from logs is less open to this bias than a rating given by the researcher.
Modern research practice adds more safeguards. They include preregistration of study hypotheses and analysis plans, registered reports in which a journal reviews the study protocol before data is collected, automated data collection, and blinded data analysis.
A researcher's good intentions are not a control. A control is anything that makes the bias impossible to act on.
071 min
When these controls fail or are not possible
Blinding is not always possible. In a design test, the researcher often has to know which prototype is on the screen, because they are the one who set it up. In that case, say so in the report as a limit, and use the other controls: a script, a scoring rule written in advance, and a second person who analyses the data without the labels.
A script does not remove every cue. A researcher can still change how they look at the screen, how quickly they move on, or how they breathe. A script reduces the cues. Blinding removes the knowledge that produces them.
Blinding the researcher does not help if the participants guess the hypothesis and act on it. That is a different problem, and the last section covers it.
The controls also cost time. A second person to run or score sessions is a real expense for a small team. The honest rule is to spend the most care where the decision is largest: when a result will decide a launch, a budget or a roadmap, run it blind. When the result is only a prompt for the next question, a script and a note about the limit are usually enough.
081 min
The observer-expectancy effect compared with similar effects
The key question is whose expectation causes the change. The observer-expectancy effect is distinct from related phenomena such as the subject-expectancy effect and demand characteristics.
| Term | Whose expectation | What it changes |
|---|---|---|
| Observer-expectancy effect | The researcher's | Participant behaviour or the reading of data, through small cues |
| Subject-expectancy effect | The participant's own beliefs about the study | The participant's own responses |
| Demand characteristics | Cues in the situation about what is wanted | The participant's responses, more broadly |
Two entries in this library are close but different. The Hawthorne effect is about participants who change their behaviour because they know they are being observed. No researcher expectation is needed for it. The Pygmalion effect is about a teacher's or manager's expectations changing another person's real performance in an ongoing relationship. In the observer-expectancy effect the concern is different: the change contaminates the measurement of a study, so that the result no longer reflects the thing under test.
?6 questions
Questions people ask
What is the observer-expectancy effect?
How do you prevent the observer-expectancy effect?
What is an example of the observer-expectancy effect?
What is the difference between the observer-expectancy effect and the Hawthorne effect?
Does the observer-expectancy effect mean the researcher is dishonest?
Can a UX researcher avoid the observer-expectancy effect when they know which design is on the screen?
§5 sources
Sources
Rosenthal, R. and Fode, K. L. (1963). "The effect of experimenter bias on the performance of the albino rat." Behavioral Science, 8(3), 183-189.
Rosenthal, R. and Fode, K. L. (1963). "The effect of experimenter bias on the performance of the albino rat." Open PDF copy of the same paper.
Pfungst, O. (1907). Clever Hans (The Horse of Mr. von Osten). English translation by C. L. Rahn, Henry Holt, 1911. Project Gutenberg edition.
Observer-expectancy effect, Wikipedia.
Show all 5 sourcesShow fewer sources
Observer bias, Wikipedia.


