Refactoring

Changing the internal structure of existing code, in small tested steps, without changing what the program does.

14 min read

· Also in

  • Code Quality
  • Maintenance

By Ravi SuranaUpdated 11 sources

Quick answer

~20 sec

Refactoring is changing the internal structure of existing code without changing what the program does. Martin Fowler's 1999 book defined it as a series of small steps that each keep behavior the same, checked by automated tests. The payoff is code that is cheaper to change later. The risk is new bugs when tests are weak.

011 min

Refactoring at a glance

  • What it is: Rearranging code in small, tested steps while every output stays identical.
  • Origin: William Opdyke's 1992 PhD thesis. Martin Fowler's 1999 book made it mainstream.
  • What it costs: Review time, harder merges, and the chance of a subtle regression.
  • When it doesn't apply: Code without tests, or any change meant to alter behavior.

021 min

The problem refactoring responds to

Software that succeeds keeps being changed. Each new feature goes wherever it is quickest to add. After a few years, a one-day change takes a week, and nobody can predict what it will break.

Here is how that looks as an event (illustrative). A product manager asks for a small change to how discounts are calculated. The engineer estimates two days. Then the engineer finds the discount logic copied into checkout, invoicing and the mobile app, each copy slightly different. The change ships two weeks late. One copy is missed, so invoices now disagree with the checkout total, and support tickets follow.

The team now sees two bad options. Leave the structure alone and pay the same delay on every future change. Or stop feature work and rewrite the module from scratch. Refactoring is a third option. The team improves the structure a little at a time and keeps the software working after every step. The work happens inside normal feature work, not as a separate project.

032 min

How refactoring actually works

Refactoring works by chaining many small changes, each of which does not change behavior. Fowler calls each change "a refactoring": a named transformation such as Extract Function (move a block of code into its own named function) or Rename Variable. One refactoring does very little. A sequence of fifty can reorganise a whole module.

"Does not change behavior" has a precise meaning. Given the same inputs, the program gives the same outputs before and after. What changes is how the code is arranged: the names, where one function ends and the next begins, where data is stored, and which module depends on which.

What keeps each step safe

  • Small steps. If something breaks, the cause is in the last few lines touched, so it is found in minutes.
  • Tests after every step. An automated test suite runs after each change. The tests turn "same behavior" into a checked fact instead of a hope. Change, run the tests, keep or undo: this short feedback loop is the core of the technique.
  • One activity at a time. Kent Beck's rule, which Fowler calls "two hats", is that a programmer is either adding function or refactoring, never both at once. A change that mixes the two cannot be checked, because some outputs are meant to change.

Most code editors used by professional programmers can now perform common refactorings, such as a rename across a whole codebase, as one automated action. Fowler treats these tools as useful but optional. Without them, he relies on small steps and frequent test runs.

The usual trigger is a code smell. Fowler and Beck use this term for a visible sign that the structure may be wrong. Examples are a very long function, the same logic in several places, or one file that changes for many unrelated reasons. A smell does not prove a problem. It tells the programmer where to look.

Much refactoring happens just before a feature is added. Beck summarised this order in one sentence, which Fowler quotes in his article on preparatory refactoring:

"for each desired change, make the change easy (warning: this may be hard), then make the easy change"

Kent Beck

041 min

Refactoring on a small scale: Extract Function

Fowler's online catalog shows each refactoring with a short before-and-after. His entry for Extract Function uses a function called printOwing. It prints a banner, calculates how much a customer owes, then prints the customer's name and the amount. A comment, //print details, marks where the printing starts.

The refactoring moves the two printing lines into a new function called printDetails. The original function now calls it by name, and the comment is deleted because the new name says the same thing.

BeforeAfter
Steps inside printOwingBanner, calculation, two print linesBanner, calculation, one call to printDetails
How a reader learns what the print lines doA commentThe function's name
What the program printsCustomer name and amountExactly the same

Nothing a user sees has changed. The benefit comes later. If the output format needs to change, or a second report needs the same details, there is now one named place to change or reuse.

The catalog lists Extract Function as the inverse of Inline Function, which folds a function back into its caller. That pairing matters. Refactoring has no fixed direction. The same technique that splits code apart can merge it back when a split turns out to be wrong.

051 min

Where the term refactoring comes from

The first detailed study was William Opdyke's PhD thesis, Refactoring Object-Oriented Frameworks, completed at the University of Illinois at Urbana-Champaign in 1992 under Ralph Johnson. Opdyke defined a set of program restructuring operations and gave a precise test for them. Refactorings do not change the behavior of a program: if the program is run with the same inputs before and after, it produces the same outputs. His focus was on automating these operations, so a tool could check that each one was safe before applying it.

Kent Beck took the practice into everyday programming. Fowler worked with Beck on the Chrysler C3 project, where Extreme Programming began, and saw Beck continually rework the code to keep it healthy. Fowler could not find a book to recommend on the technique, so he wrote one. Refactoring: Improving the Design of Existing Code was published in 1999, with examples in Java.

Fowler is clear that he did not invent the idea. In a 2004 post he calls himself "not the father or the inventor of refactoring - just a documenter."

061 min

How refactoring changed since 1999

Fowler published a second edition in 2018. The structure stayed the same: an opening example, principles, a list of code smells, a chapter on testing, then the catalog. The contents changed a lot. Of the first edition's 68 refactorings, 58 were kept and 17 new ones were added. Almost every page was rewritten.

Two changes show how programming itself moved on. The examples switched from Java to JavaScript, because Fowler wanted the language the most readers could follow. And the book moved away from classes as the main unit of code. The rename of Extract Method to Extract Function is the visible sign of that shift.

The meaning of the word moved too, in the other direction. By 2004 Fowler was already complaining that people used "refactoring" for any large restructuring, including work that left a system broken for days. He called this the "refactoring malapropism". His test is short:

"If somebody talks about a system being broken for a couple of days while they are refactoring, you can be pretty sure they are not refactoring."

In practice, most engineers today use the word more loosely than Fowler or Opdyke did. Survey data from Microsoft, covered in the contested-definition section, puts a number on that gap.

072 min

Refactoring GitHub's merge button

In December 2015, GitHub engineer Vicent Martí described how the team replaced the code that runs when a user clicks the Merge button on a pull request. The old code was a shell script built on an outdated Git merge strategy. It wrote temporary files to disk, it was slow, and it did not behave exactly like git merge on a developer's own computer. The new code used libgit2, a library that can merge in memory.

The team did not switch in one step. First, they refactored the existing merge code so the Git-specific part sat in its own method, and deployed only that change. Then they added a second method with the same inputs and outputs, built on libgit2.

To check that the two paths behaved the same, they used Scientist, a Ruby library GitHub wrote for testing refactorings and rewrites in production. On each sampled request, Scientist ran both paths, returned the old result to the user, and logged any difference. The team started at 1% of requests and fixed each mismatch it found.

Some mismatches were bugs in the new code. Some were bugs in the old code. In one case, the old script reported a file with exactly 768 conflicts as a clean merge. A shell keeps only the lowest 8 bits of an exit code, and 768 is a multiple of 256.

After about four days of fixes, the experiment ran on every request for 24 hours with no mismatches. Martí reports that Scientist verified tens of millions of merge commits to be virtually identical in the old and new paths. Only then was the old code deleted.

The lesson for other teams is where the proof came from. GitHub had a large test suite, but it treated production traffic as the final test of "same behavior".

081 min

What refactoring costs, even done right

Refactoring has costs even when it is done correctly. The clearest evidence comes from a 2012 field study at Microsoft by Miryung Kim, Thomas Zimmermann and Nachiappan Nagappan. They surveyed 328 engineers whose commit messages mentioned refactoring. 77% of the participants consider that refactoring comes with a risk of introducing subtle bugs and functionality regression. 12% said merging code became harder after a refactoring, and 24% named increased testing cost.

These costs fall on different people:

  • Reviewers. A refactoring touches many files at once. One engineer in the study said it "burdens code reviewers and increases the odds that your change will collide with someone else's change."
  • Parallel work. Anyone working on a separate branch of the code may find their change no longer applies cleanly after files are moved or renamed. Fowler says this is his main concern with feature branches that live longer than a few days: they discourage small refactorings.
  • The schedule. The payoff arrives later, on a future change, while the time is spent now. That makes refactoring easy to cut from a plan and hard to justify with a number.

There is also a quieter cost. A refactoring aimed at a change that never happens adds structure nobody needs. Engineers in the Microsoft study named "over-engineering" as a risk in its own right.

091 min

What the Windows refactoring data showed

The same Microsoft study looked at one large, planned refactoring effort. A dedicated team spent several years restructuring Windows to reduce the dependencies between its modules. The researchers compared that team's work with ordinary changes in the Windows 7 version history.

The binary modules the team refactored most often had a larger reduction in dependencies on other modules than the modules that were only changed most often. They also had fewer defects after release. The top 25% of frequently refactored binaries reduced the number of post-release defects 12.2% more than other modified binaries on average.

Two limits matter when reading that figure. First, it comes from one company and one codebase, measured by the people who also ran the survey. Second, this was a planned, multi-year programme with custom tools, not the small everyday refactoring Fowler describes. The study supports the claim that refactoring can reduce defects at scale. It does not show that every small cleanup pays for itself.

101 min

Refactoring vs. rewriting and restructuring

Three words get used as if they meant the same thing. The deciding axis is whether the software keeps working the whole time.

RefactoringRestructuringRewriting
What changesInternal structure onlyStructure, by any methodEverything, starting from scratch
Is the software working throughout?Yes, after every small stepNot necessarilyNo, until the new version is finished
How "same behavior" is checkedTests after each stepVariesAt the end, if at all

Fowler treats restructuring as the general term for rearranging parts of a whole. Refactoring is one specific way of doing it, with small steps that each keep behavior the same.

A full rewrite is the opposite end. In 2000, Joel Spolsky used Netscape as the warning case. The company decided to rewrite its browser from scratch, and there was never a version 5.0. The last major release, version 4.0, was released almost three years ago, he wrote, as Netscape 6.0 finally reached its first public beta. Spolsky's alternative was careful restructuring of the existing code.

111 min

When refactoring is the wrong tool

Refactoring is the wrong tool when there is no reliable way to check that behavior stayed the same. Without tests, each step is a guess. Engineers in the Microsoft study said that an inadequate regression test suite often stops them from starting at all. The right first move in that situation is to write tests that record the current behavior, then refactor.

It is also the wrong tool when the goal is to change behavior. Moving from one database to another, or replacing a system that must work differently, needs a migration plan. Fowler's Strangler Fig pattern fits this case better. A new system is built next to the old one. It takes over one function at a time until the old system can be switched off.

And it is wasted effort on code that is about to be deleted. Restructuring a module that the product plan retires next quarter improves nothing anyone will use.

A common failure has its own tell: a "refactoring" ticket that has run for weeks with the system broken. By Fowler's definition, that work has become a rewrite. It should be planned, estimated and reviewed as one, with the risk of an overrun stated openly.

122 min

How refactoring shows up in product and engineering work

For a product manager

A product manager planning a quarter sees two estimates for the same feature (illustrative). One engineer says eight days. Another says three days of refactoring, then two days for the feature. The second plan is shorter, and it leaves the next feature in the same area cheaper too. The decision is not "features or refactoring". It is whether refactoring goes into the feature estimate, where it belongs, or into a separate "cleanup" ticket that gets cut first when the schedule slips. A useful question to ask engineers is "what would make this change easy?"

For an engineer

A staff engineer reviewing a pull request sees a change that renames a core class, moves three files and adds a new payment option (illustrative). The right call is to ask for it to be split. One pull request holds the pure refactoring, where every test should pass unchanged. The other holds the feature, where test changes are expected. Each can then be reviewed separately.

A second decision is when to stop. Fowler warns that opportunistic refactoring can turn into fixing one thing after another. His advice is to leave the code better than it was, not perfect.

For a founder or engineering lead

A founder deciding whether to rewrite a slow, messy product should ask whether the code can be improved while it keeps shipping. Gall's Law is the general form of the argument: complex systems that work usually grew out of simpler systems that worked. A quick patch on every bug is the opposite habit, and it is the one that makes a rewrite look necessary.

132 min

Where the definition of refactoring is contested

Researchers and practitioners disagree about how strict the definition should be.

The strict view is Opdyke's and Fowler's. A refactoring preserves behavior exactly. Anything else is a different activity, and calling it refactoring hides the risk.

The Microsoft field study found that practice does not match. When asked how they define refactoring, 46% of developers did not mention preserving behavior at all. 78% described it as a change that improves some aspect of the program, such as readability, maintainability or performance. Many also said the named refactorings in Fowler's catalog are only the smallest unit of a larger restructuring effort. The paper links this to an argument by Ralph Johnson, Opdyke's own advisor, that refactoring preserves some behavior but not behavior in every respect. Timing, memory use and log output can all change while the results stay the same.

The strongest form of the loose view is this: real systems rarely have a complete definition of behavior. So "no change in behavior" always means "no change in the behavior we check".

Fowler's reply is practical. The word is useful only if it tells a reviewer how risky a change is. If "refactoring" can mean a two-week rewrite, it tells them nothing.

For a working team, the practical answer is to state which behavior is being preserved and how it is checked. "All tests pass unchanged" is one answer. GitHub's production comparison is a stronger one.

?8 questions

Questions people ask

What is refactoring in simple terms?

Refactoring is reorganising code so it is easier to understand and change, without changing what it does for users. It is done in small steps, and automated tests run after each step to confirm the program still gives the same results.

Who invented refactoring?

No single person invented it. William Opdyke wrote the first detailed study in his 1992 PhD thesis under Ralph Johnson, and Kent Beck made it part of daily practice. Martin Fowler's 1999 book documented the technique and made the term widely known.

What is the difference between refactoring and rewriting?

Refactoring keeps the software working after every small step. Rewriting replaces the code from scratch, and the new version works only when it is finished. Refactoring spreads the risk over many small checked changes. A rewrite concentrates it at the switch-over.

Does refactoring change functionality?

No, not by definition. If a change alters what users see or what the program outputs, it is a feature change or a bug fix, not a refactoring. In practice, many engineers use the word loosely, so it is worth asking which behavior was checked.

Can you refactor without tests?

You can, but each step then relies on reading the code carefully instead of checking it. The usual advice is to first write tests that capture how the code behaves today. Automated editor refactorings such as rename are the safest steps when tests are missing.

When should you not refactor?

Do not refactor code that is about to be deleted, code with no way to check its behavior, or a change whose purpose is to alter behavior. Refactoring also competes with urgent fixes. A production outage needs the smallest safe fix first, with restructuring afterwards.

How do you justify refactoring to a product manager?

Tie it to a specific upcoming change. "This feature takes eight days as the code is, or five with a three-day refactoring first" is a decision a product manager can make. Asking for general cleanup time with no linked feature is much harder to approve.

Is refactoring the same as paying down technical debt?

Refactoring is the main tool for it, but they are not the same thing. Technical debt is the extra cost that poorly structured code adds to future changes. Refactoring is one way to reduce that cost. Rewriting or deleting code are others.

§11 sources

Sources on refactoring

  1. Fowler, M. (2018). Refactoring: Improving the Design of Existing Code, 2nd ed. Addison-Wesley. Overview and definition:

  2. Fowler, M. (2018). The Second Edition of "Refactoring".

  3. Fowler, M. Refactoring catalog, Extract Function.

  4. Opdyke, W. F. (1992). Refactoring Object-Oriented Frameworks. PhD thesis, University of Illinois at Urbana-Champaign.

Show all 11 sources
  1. Fowler, M. (2004). Refactoring Malapropism.

  2. Fowler, M. (2015). An example of preparatory refactoring.

  3. Fowler, M. (2011). Opportunistic Refactoring.

  4. Kim, M., Zimmermann, T. & Nagappan, N. (2012). A Field Study of Refactoring Challenges and Benefits. FSE 2012.

  5. Martí, V. (2015). Move Fast and Fix Things. GitHub Blog.

  6. Spolsky, J. (2000). Things You Should Never Do, Part I.

  7. Fowler, M. (2024). Strangler Fig.

Keep reading

More from Engineering

All of Engineering
All of Engineering