Exploration / Exploitation

The trade-off between using an option already known to work and trying uncertain options that might work better.

15 min read

Β· Also in

  • Strategy

By Ravi SuranaUpdated 7 sources

Quick answer

~20 sec

Exploration / exploitation is the trade-off between using the option already known to work (exploitation) and testing uncertain options that might work better (exploration). Product managers and founders face it when they split roadmap time, traffic, or budget between a proven feature or channel and untested ideas. The right split depends on how long the team can still benefit.

011 min

Exploration / exploitation at a glance

  • What it is: every unit of effort either earns from what you know or buys information about what you do not know.
  • Origin: the multi-armed bandit problem in statistics, and James March's 1991 paper on organizational learning.
  • Produces: an explicit exploration budget, such as a share of traffic, sprint capacity, or spend.
  • The deciding variable: the time horizon. The more future decisions can use what you learn, the more exploration is worth.
  • Fails when: a team measures exploration with exploitation metrics and cuts it every quarter.

021 min

What breaks without an exploration-exploitation balance

The common failure is quiet. A team has a feature, a channel, or an onboarding flow that works. Each quarter, the roadmap goes to making that thing slightly better, because the result of a small improvement is fast and easy to measure. Each round usually gains less than the last, a pattern called diminishing returns. New ideas get the leftover time, which is usually none. Two years later the team is very good at a product the market has moved away from.

The opposite failure is louder. A team ships a new bet every sprint, drops each one before it has enough users to show a result, and never builds depth in anything. It pays the cost of each experiment without collecting what the experiment could teach.

James March described both failures in 1991. Systems that only explore "suffer the costs of experimentation without gaining many of its benefits." Systems that only exploit get stuck on a solution that is good enough today but not the best available.

The cost is concrete. A team stuck on the first failure will make every prioritisation meeting a choice between two versions of the same product. A team stuck on the second will reach its next funding or planning review with ten half-built features and no evidence about any of them.

033 min

How the exploration-exploitation trade-off works

The two options and what they return

Formal work models the problem as a multi-armed bandit. Each option is an "arm", named after the lever on a slot machine. Each arm pays out at an unknown average rate. Every time you choose, you either pull the arm with the best average so far (exploit) or pull a less-tested arm to learn its rate (explore).

March described the returns of exploration as uncertain, distant, and often negative, and the returns of exploitation as positive, proximate, and predictable. That asymmetry is the whole problem. The cost of exploring is paid now and is easy to see. The benefit arrives later, if at all.

ExploitationExploration
Product exampleOptimise the checkout that already convertsBuild a new pricing tier nobody has asked for
ReturnLikely, small, soonUncertain, possibly large, later
What you getRevenue or usage nowInformation that improves future choices
Feedback speedDaysWeeks to quarters

Why the time horizon decides the split

Information from exploration is only valuable if there are future decisions that can use it. The value of a test is roughly the improvement it might reveal, multiplied by how many future decisions will benefit. When few decisions remain, exploration cannot pay back its cost.

Illustrative worked example. A subscription app gets 10,000 sign-ups a week, and its current onboarding flow converts 10%. The team has a new flow that could convert anywhere from 7% to 13%, with even odds of being better or worse. Testing it means sending 2,000 users a week to the new flow for 4 weeks. If the new flow is worse by 3 points, the test loses about 240 conversions. If it is better by 3 points, the app gains about 300 conversions for every week after the test.

Weeks left before the app is retired or rebuiltExpected cost of the testExpected payoff after the test
3about 120 conversionsnothing, because the test ends after the last week
8about 120 conversionsabout 600 conversions (4 weeks Γ— 150)
100about 120 conversionsabout 14,400 conversions (96 weeks Γ— 150)

The same test is a mistake with 3 weeks left and an easy choice with 100 weeks left. Nothing about the new flow changed. Only the horizon changed.

Common exploration rules

RuleHow it exploresProduct translation
Epsilon-greedyA fixed small share of choices goes to random optionsReserve a fixed share of traffic or sprint capacity for new bets
Upper confidence boundPrefers options whose value is still uncertainFund the idea you know least about, if its best case is high
Thompson samplingChooses each option in proportion to its chance of being bestSplit effort in proportion to how likely each bet is to win

These rules do not produce a good decision when the options change faster than you can learn about them. A bandit that learns which article is popular is useless if the article pool is replaced before the learning finishes.

041 min

Where exploration / exploitation came from

The formal problem predates product management by decades. Allied scientists studied it during World War II. According to the statistician Peter Whittle, the problem was so hard that scientists proposed dropping it over Germany to waste German scientists' time. Herbert Robbins gave the version most people now study in a 1952 paper on the sequential design of experiments.

The best-known solution came later. In 1979 John Gittins published an index that gives an optimal policy for maximizing the expected discounted reward. The Gittins index gives each option a single score. That score counts both what the option is expected to pay and what you would learn by trying it. You then choose the option with the highest score. The word "discounted" matters: rewards far in the future count for less, so the index already contains a time horizon.

The business version comes from James March, a Stanford organization theorist. His 1991 paper in Organization Science, "Exploration and Exploitation in Organizational Learning", used simulations of organizations learning from their members. It argued that learning processes improve exploitation faster than exploration. March wrote the paper about firms, not algorithms, and it is the reason the terms are now used in strategy and product work.

052 min

Exploration / exploitation at MSN

Microsoft Research described a production system called the Decision Service in a 2016 paper. One of its first users was MSN, which used it to order news stories on the MSN homepage.

The decision MSN faced was a clean exploration-exploitation problem. Editors already chose and ranked tens of articles, and their ordering was the known-good option. Earlier machine-learning attempts had lost to it. The Decision Service instead ran a contextual bandit, a form of online learning that updates its model as new clicks arrive. It showed a share of visitors a randomised ordering to learn which stories each type of reader clicked, and it used what it had learned for everyone else. The paper reports that exploration used an epsilon-greedy rule with epsilon set to 33%. That is a large share, chosen because news goes stale within hours.

On the Slate segment of the MSN homepage, the bandit system beat the editors' ordering by more than 25% in click-through rate over two weeks. The larger Panel segment improved by 5.3%. The authors add that long-term engagement measures, such as sessions per user, held steady or improved. MSN then made the system the default for logged-in users in early 2016.

The product decision in this case was not "bandit or editors". It was how much traffic to spend on learning. A 33% exploration share would waste traffic on a product whose options change slowly. For a news page that changes every hour, a small share would never learn fast enough.

062 min

Exploration / exploitation at Yahoo's Today Module

Yahoo faced the same kind of problem earlier, and its team drew a different lesson from it. The Today Module was the main news panel on the Yahoo front page. Human editors refreshed its pool of articles every hour, and the product had to decide which article to feature for each visitor.

In a 2010 paper, Lihong Li and colleagues described how the team had already set aside a small random bucket of traffic. Articles in that bucket were chosen at random. That bucket was the team's exploration budget. It also turned out to be its most valuable data asset. Because the choices in it were random, the team could replay the logged events to test new algorithms offline without risking live traffic.

On a Yahoo! Front Page Today Module dataset of over 33 million events, the contextual bandit produced a 12.5% click lift over a standard context-free bandit. The paper adds that the advantage grew when data was more scarce.

The variable that differs from MSN is what the exploration traffic was for. At MSN, exploration mainly fed the live model. At Yahoo, the random bucket also served as an unbiased test set for comparing strategies. Together the two cases show that exploration buys two things. It buys information about the options, and it buys clean data to judge your own decision rules. A team that removes its exploration budget loses both.

072 min

Why the time horizon changes exploration / exploitation

The claim that exploration is worth more when more time remains is not only a mathematical result. People follow it too. Robert Wilson and colleagues at Princeton tested it in 2014 with a "Horizon task". Participants chose between two slot machines, with either one choice or six choices left in each game. Participants given six choices per game sought more information and made noisier choices than participants given one choice. With a longer future, people tried the less-known option more often.

For a product team, the horizon is not a fixed number. It is set by several things:

  • Product life stage. A product in its first year has years of decisions ahead. A feature scheduled for retirement next quarter has almost none.
  • How long a win lasts. A pricing insight can pay for years. A trick that competitors copy within a month has a short horizon even in a young product.
  • Runway. A startup with four months of cash has a short horizon, whatever its long-term plans. Exploration that pays back in a year is not affordable to it.
  • Rate of change. If users or competitors change quickly, what you learned expires quickly. That shortens the horizon of each lesson, even when the company plans to exist for a long time.

This is why the same exploration budget can be right for one team and wrong for the team next to it. A growth team working on a mature product that makes money can exploit most of its capacity. A new product line inside the same company should explore most of its capacity, because nearly all of its value lies in decisions it has not made yet.

082 min

How to set an exploration / exploitation budget this week

The artifact is a one-page exploration budget. It is a spreadsheet or document that states what share of capacity goes to exploration, what counts as exploration, and when the split will be reviewed.

  1. List where your capacity goes.

    Take the last two sprints or the current quarter's roadmap. Mark each item as exploit (improves something already proven) or explore (tests something whose value you do not know). Use a spreadsheet with one row per item and a column for estimated effort.

  2. Compute your actual split.

    Add up the effort in each column. Most teams find their exploration share is far lower than they believed.

  3. Estimate your horizon.

    Write down how many months of decisions the team will still make on this product, and how quickly user behaviour or the market changes. Use the four factors in the horizon section.

  4. Set a target share.

    Short horizon or tight runway: keep exploration small and aimed at ideas that pay back within that horizon. Long horizon and stable income: raise it. State the number, for example "20% of sprint capacity".

  5. Define exploration metrics separately.

    Judge an exploration item by what it taught you and how fast. Do not judge it by revenue in the quarter it ran. Write down, before each bet starts, the question it answers and the result that would end it.

  6. Put a review date on the split.

    Revisit the share every quarter, or when runway, product stage, or market speed changes.

If you run experiments on live traffic, the same steps apply to traffic instead of people. Decide what share of users sees untested variants, and keep a random slice for later evaluation, as Yahoo did.

092 min

Exploration / exploitation done badly

The competence trap

March's own warning is the most common failure. March warned that competence in an inferior activity can grow large enough to shut out a better activity the organization has little experience with. The tell: every new idea loses the prioritisation debate because "our current product does this better". That comparison is always made against a version that has had years of polish. This failure is built into how learning works, not a misuse of the idea.

Exploration judged by exploitation metrics

A team sets aside 20% for new bets, then reviews those bets on quarterly revenue. Each bet looks worse than the core product, so the budget is cut the next quarter. Loss aversion adds to this, because a failed bet feels worse than an equal win feels good. The tell is an exploration budget that exists on paper but has shrunk every planning cycle. This one is a misuse, and step 5 above exists to prevent it.

Exploration that never ends

The opposite misuse: bets are never stopped, because each one "just needs another quarter". Exploration has value only if you act on what it shows. The tell is a list of experiments with no written stop criteria and no dates.

Wrong model of the options

Bandit rules assume the options stay the same while you learn. When the options change quickly, the learning expires. The tell is an experiment that reaches a clear result for a product state that no longer exists. This one is a limit of the formal model.

101 min

Where the exploration / exploitation evidence is contested

March treated exploration and exploitation as competing for the same scarce resources, so more of one means less of the other. Later researchers disagreed about whether that is always true.

In a 2006 review for the Academy of Management Journal, Anil Gupta, Ken Smith and Christina Shalley set out the question directly. Are the two ends of one scale, where more exploration means less exploitation? Or are they two separate dimensions, where a company can be high on both? Their answer depends on the unit being studied. Within one person or one team, the two usually compete. Across loosely connected units, one unit can explore while another exploits. Gupta, Smith and Shalley argued in 2006 that no universal argument can be made for either view.

This matters for how you act. If exploration and exploitation compete inside your team, you need an explicit split, as in the budget above. If your company has several loosely connected teams, you can instead give whole teams different jobs. One team explores a new product line while another improves the core. Both approaches are defended in the research. While it remains unsettled, choose by structure. A single small team should split its own time. A larger organization can split by team, as long as what the exploring team learns reaches the teams that exploit.

112 min

Exploration / exploitation vs. A/B testing and risk-taking

Exploration / exploitation is a trade-off about how to divide effort over time. It is often confused with two narrower ideas.

Exploration / exploitationA/B testingRisk-taking
What it isA trade-off between learning and earning over many decisionsA method that compares variants with a fixed traffic splitAccepting a wider range of outcomes
Question it answersHow much effort should go to learning right now?Which of these variants is better?Is the possible upside worth the possible loss?
Uses the time horizonYes, it is centralNo, the test ends and then you chooseNot by itself

The deciding fact: an A/B test is one way to explore, and it keeps the split fixed until the test ends. A bandit changes the split as evidence arrives, so it exploits during the test. MSN's own team used a standard A/B test to confirm that its bandit beat the editors. The two work together.

Exploration is also not the same as taking risk. Exploiting can be risky, such as betting everything on one proven channel. Exploring can be low-risk, such as a small test on 1% of traffic. The difference is whether the choice is made to gain information.

People also confuse it with the sunk cost problem. Sunk cost is about letting past spending shape a decision. Exploration / exploitation is about the value of future information. A team can keep exploiting a weak product for sunk-cost reasons, but the fix is different.

?7 questions

Questions people ask

Who should decide the exploration / exploitation split?

The person who owns the roadmap or budget should decide it, usually the product lead or founder. They should state it as a number and review it each quarter. Leaving it implicit lets exploitation win by default, because its results arrive sooner.

What is an example of exploration / exploitation in product management?

A team that spends 80% of a sprint improving the checkout flow that already converts, and 20% testing a new pricing tier, is splitting between exploitation and exploration. MSN's homepage used a bandit to make the same split for news stories.

Is exploration / exploitation the same as A/B testing?

No. An A/B test is one way to explore, with a fixed traffic split until the test ends. Exploration / exploitation is the wider question of how much effort to spend on learning versus earning, and it changes with the time horizon.

How much should a team spend on exploration?

There is no universal share. Spend more when many future decisions will use what you learn: a young product, a long runway, a stable market. Spend less near a product's end of life or when cash is short.

Why does the time horizon matter so much?

Information has value only when a later decision uses it. With many decisions ahead, a lesson pays back many times. With few left, the cost of learning is paid in full and the benefit has no time to arrive.

What is the Gittins index?

The Gittins index is a score, published by John Gittins in 1979, for each option in a bandit problem. It combines expected payoff with the value of learning more. Choosing the highest-scoring option is optimal when future rewards are discounted.

What do you do if exploration never produces a winner?

Check whether bets had written stop criteria and were judged on what they taught. Also check that they ran long enough to show a result. If all three hold, shift exploration toward options nearer your current users rather than cutting it.

Β§7 sources

Sources for exploration / exploitation

  1. March, J. G. (1991). "Exploration and Exploitation in Organizational Learning." Organization Science 2(1), 71–87. (1991).pdf Β· JSTOR:

  2. Gittins, J. C. (1979). "Bandit Processes and Dynamic Allocation Indices." Journal of the Royal Statistical Society, Series B 41(2), 148–177.

  3. "Multi-armed bandit." Wikipedia, for the history of the problem (Whittle, Robbins 1952, Gittins 1979).

  4. Wilson, R. C., Geana, A., White, J. M., Ludvig, E. A., & Cohen, J. D. (2014). "Humans use directed and random exploration to solve the explore-exploit dilemma." Journal of Experimental Psychology: General 143(6), 2074–2081.

Show all 7 sources
  1. Agarwal, A., et al. (2016). "Making Contextual Decisions with Low Technical Debt" (Microsoft Decision Service, MSN deployment). arXiv:1606.03966.

  2. Li, L., Chu, W., Langford, J., & Schapire, R. E. (2010). "A Contextual-Bandit Approach to Personalized News Article Recommendation" (Yahoo! Today Module). arXiv:1003.0146.

  3. Gupta, A. K., Smith, K. G., & Shalley, C. E. (2006). "The Interplay between Exploration and Exploitation." Academy of Management Journal 49(4), 693–706. (2006).pdf

Keep reading

More from Product

All of Product
All of Product