011 min
Refactoring at a glance
- What it is: Rearranging code in small, tested steps while every output stays identical.
- Origin: William Opdyke's 1992 PhD thesis. Martin Fowler's 1999 book made it mainstream.
- What it costs: Review time, harder merges, and the chance of a subtle regression.
- When it doesn't apply: Code without tests, or any change meant to alter behavior.
021 min
The problem refactoring responds to
Software that succeeds keeps being changed. Each new feature goes wherever it is quickest to add. After a few years, a one-day change takes a week, and nobody can predict what it will break.
Here is how that looks as an event (illustrative). A product manager asks for a small change to how discounts are calculated. The engineer estimates two days. Then the engineer finds the discount logic copied into checkout, invoicing and the mobile app, each copy slightly different. The change ships two weeks late. One copy is missed, so invoices now disagree with the checkout total, and support tickets follow.
The team now sees two bad options. Leave the structure alone and pay the same delay on every future change. Or stop feature work and rewrite the module from scratch. Refactoring is a third option. The team improves the structure a little at a time and keeps the software working after every step. The work happens inside normal feature work, not as a separate project.
032 min
How refactoring actually works
Refactoring works by chaining many small changes, each of which does not change behavior. Fowler calls each change "a refactoring": a named transformation such as Extract Function (move a block of code into its own named function) or Rename Variable. One refactoring does very little. A sequence of fifty can reorganise a whole module.
"Does not change behavior" has a precise meaning. Given the same inputs, the program gives the same outputs before and after. What changes is how the code is arranged: the names, where one function ends and the next begins, where data is stored, and which module depends on which.
What keeps each step safe
- Small steps. If something breaks, the cause is in the last few lines touched, so it is found in minutes.
- Tests after every step. An automated test suite runs after each change. The tests turn "same behavior" into a checked fact instead of a hope. Change, run the tests, keep or undo: this short feedback loop is the core of the technique.
- One activity at a time. Kent Beck's rule, which Fowler calls "two hats", is that a programmer is either adding function or refactoring, never both at once. A change that mixes the two cannot be checked, because some outputs are meant to change.
Most code editors used by professional programmers can now perform common refactorings, such as a rename across a whole codebase, as one automated action. Fowler treats these tools as useful but optional. Without them, he relies on small steps and frequent test runs.
The usual trigger is a code smell. Fowler and Beck use this term for a visible sign that the structure may be wrong. Examples are a very long function, the same logic in several places, or one file that changes for many unrelated reasons. A smell does not prove a problem. It tells the programmer where to look.
Much refactoring happens just before a feature is added. Beck summarised this order in one sentence, which Fowler quotes in his article on preparatory refactoring:
"for each desired change, make the change easy (warning: this may be hard), then make the easy change"
041 min
Refactoring on a small scale: Extract Function
Fowler's online catalog shows each refactoring with a short before-and-after. His entry for Extract Function uses a function called printOwing. It prints a banner, calculates how much a customer owes, then prints the customer's name and the amount. A comment, //print details, marks where the printing starts.
The refactoring moves the two printing lines into a new function called printDetails. The original function now calls it by name, and the comment is deleted because the new name says the same thing.
| Before | After | |
|---|---|---|
Steps inside printOwing | Banner, calculation, two print lines | Banner, calculation, one call to printDetails |
| How a reader learns what the print lines do | A comment | The function's name |
| What the program prints | Customer name and amount | Exactly the same |
Nothing a user sees has changed. The benefit comes later. If the output format needs to change, or a second report needs the same details, there is now one named place to change or reuse.
The catalog lists Extract Function as the inverse of Inline Function, which folds a function back into its caller. That pairing matters. Refactoring has no fixed direction. The same technique that splits code apart can merge it back when a split turns out to be wrong.
051 min
Where the term refactoring comes from
The first detailed study was William Opdyke's PhD thesis, Refactoring Object-Oriented Frameworks, completed at the University of Illinois at Urbana-Champaign in 1992 under Ralph Johnson. Opdyke defined a set of program restructuring operations and gave a precise test for them. Refactorings do not change the behavior of a program: if the program is run with the same inputs before and after, it produces the same outputs. His focus was on automating these operations, so a tool could check that each one was safe before applying it.
Kent Beck took the practice into everyday programming. Fowler worked with Beck on the Chrysler C3 project, where Extreme Programming began, and saw Beck continually rework the code to keep it healthy. Fowler could not find a book to recommend on the technique, so he wrote one. Refactoring: Improving the Design of Existing Code was published in 1999, with examples in Java.
Fowler is clear that he did not invent the idea. In a 2004 post he calls himself "not the father or the inventor of refactoring - just a documenter."
061 min
How refactoring changed since 1999
Fowler published a second edition in 2018. The structure stayed the same: an opening example, principles, a list of code smells, a chapter on testing, then the catalog. The contents changed a lot. Of the first edition's 68 refactorings, 58 were kept and 17 new ones were added. Almost every page was rewritten.
Two changes show how programming itself moved on. The examples switched from Java to JavaScript, because Fowler wanted the language the most readers could follow. And the book moved away from classes as the main unit of code. The rename of Extract Method to Extract Function is the visible sign of that shift.
The meaning of the word moved too, in the other direction. By 2004 Fowler was already complaining that people used "refactoring" for any large restructuring, including work that left a system broken for days. He called this the "refactoring malapropism". His test is short:
"If somebody talks about a system being broken for a couple of days while they are refactoring, you can be pretty sure they are not refactoring."
In practice, most engineers today use the word more loosely than Fowler or Opdyke did. Survey data from Microsoft, covered in the contested-definition section, puts a number on that gap.
072 min
Refactoring GitHub's merge button
In December 2015, GitHub engineer Vicent Martí described how the team replaced the code that runs when a user clicks the Merge button on a pull request. The old code was a shell script built on an outdated Git merge strategy. It wrote temporary files to disk, it was slow, and it did not behave exactly like git merge on a developer's own computer. The new code used libgit2, a library that can merge in memory.
The team did not switch in one step. First, they refactored the existing merge code so the Git-specific part sat in its own method, and deployed only that change. Then they added a second method with the same inputs and outputs, built on libgit2.
To check that the two paths behaved the same, they used Scientist, a Ruby library GitHub wrote for testing refactorings and rewrites in production. On each sampled request, Scientist ran both paths, returned the old result to the user, and logged any difference. The team started at 1% of requests and fixed each mismatch it found.
Some mismatches were bugs in the new code. Some were bugs in the old code. In one case, the old script reported a file with exactly 768 conflicts as a clean merge. A shell keeps only the lowest 8 bits of an exit code, and 768 is a multiple of 256.
After about four days of fixes, the experiment ran on every request for 24 hours with no mismatches. Martí reports that Scientist verified tens of millions of merge commits to be virtually identical in the old and new paths. Only then was the old code deleted.
The lesson for other teams is where the proof came from. GitHub had a large test suite, but it treated production traffic as the final test of "same behavior".
081 min
What refactoring costs, even done right
Refactoring has costs even when it is done correctly. The clearest evidence comes from a 2012 field study at Microsoft by Miryung Kim, Thomas Zimmermann and Nachiappan Nagappan. They surveyed 328 engineers whose commit messages mentioned refactoring. 77% of the participants consider that refactoring comes with a risk of introducing subtle bugs and functionality regression. 12% said merging code became harder after a refactoring, and 24% named increased testing cost.
These costs fall on different people:
- Reviewers. A refactoring touches many files at once. One engineer in the study said it "burdens code reviewers and increases the odds that your change will collide with someone else's change."
- Parallel work. Anyone working on a separate branch of the code may find their change no longer applies cleanly after files are moved or renamed. Fowler says this is his main concern with feature branches that live longer than a few days: they discourage small refactorings.
- The schedule. The payoff arrives later, on a future change, while the time is spent now. That makes refactoring easy to cut from a plan and hard to justify with a number.
There is also a quieter cost. A refactoring aimed at a change that never happens adds structure nobody needs. Engineers in the Microsoft study named "over-engineering" as a risk in its own right.
091 min
What the Windows refactoring data showed
The same Microsoft study looked at one large, planned refactoring effort. A dedicated team spent several years restructuring Windows to reduce the dependencies between its modules. The researchers compared that team's work with ordinary changes in the Windows 7 version history.
The binary modules the team refactored most often had a larger reduction in dependencies on other modules than the modules that were only changed most often. They also had fewer defects after release. The top 25% of frequently refactored binaries reduced the number of post-release defects 12.2% more than other modified binaries on average.
Two limits matter when reading that figure. First, it comes from one company and one codebase, measured by the people who also ran the survey. Second, this was a planned, multi-year programme with custom tools, not the small everyday refactoring Fowler describes. The study supports the claim that refactoring can reduce defects at scale. It does not show that every small cleanup pays for itself.
101 min
Refactoring vs. rewriting and restructuring
Three words get used as if they meant the same thing. The deciding axis is whether the software keeps working the whole time.
| Refactoring | Restructuring | Rewriting | |
|---|---|---|---|
| What changes | Internal structure only | Structure, by any method | Everything, starting from scratch |
| Is the software working throughout? | Yes, after every small step | Not necessarily | No, until the new version is finished |
| How "same behavior" is checked | Tests after each step | Varies | At the end, if at all |
Fowler treats restructuring as the general term for rearranging parts of a whole. Refactoring is one specific way of doing it, with small steps that each keep behavior the same.
A full rewrite is the opposite end. In 2000, Joel Spolsky used Netscape as the warning case. The company decided to rewrite its browser from scratch, and there was never a version 5.0. The last major release, version 4.0, was released almost three years ago, he wrote, as Netscape 6.0 finally reached its first public beta. Spolsky's alternative was careful restructuring of the existing code.
111 min
When refactoring is the wrong tool
Refactoring is the wrong tool when there is no reliable way to check that behavior stayed the same. Without tests, each step is a guess. Engineers in the Microsoft study said that an inadequate regression test suite often stops them from starting at all. The right first move in that situation is to write tests that record the current behavior, then refactor.
It is also the wrong tool when the goal is to change behavior. Moving from one database to another, or replacing a system that must work differently, needs a migration plan. Fowler's Strangler Fig pattern fits this case better. A new system is built next to the old one. It takes over one function at a time until the old system can be switched off.
And it is wasted effort on code that is about to be deleted. Restructuring a module that the product plan retires next quarter improves nothing anyone will use.
A common failure has its own tell: a "refactoring" ticket that has run for weeks with the system broken. By Fowler's definition, that work has become a rewrite. It should be planned, estimated and reviewed as one, with the risk of an overrun stated openly.
122 min
How refactoring shows up in product and engineering work
For a product manager
A product manager planning a quarter sees two estimates for the same feature (illustrative). One engineer says eight days. Another says three days of refactoring, then two days for the feature. The second plan is shorter, and it leaves the next feature in the same area cheaper too. The decision is not "features or refactoring". It is whether refactoring goes into the feature estimate, where it belongs, or into a separate "cleanup" ticket that gets cut first when the schedule slips. A useful question to ask engineers is "what would make this change easy?"
For an engineer
A staff engineer reviewing a pull request sees a change that renames a core class, moves three files and adds a new payment option (illustrative). The right call is to ask for it to be split. One pull request holds the pure refactoring, where every test should pass unchanged. The other holds the feature, where test changes are expected. Each can then be reviewed separately.
A second decision is when to stop. Fowler warns that opportunistic refactoring can turn into fixing one thing after another. His advice is to leave the code better than it was, not perfect.
For a founder or engineering lead
A founder deciding whether to rewrite a slow, messy product should ask whether the code can be improved while it keeps shipping. Gall's Law is the general form of the argument: complex systems that work usually grew out of simpler systems that worked. A quick patch on every bug is the opposite habit, and it is the one that makes a rewrite look necessary.
132 min
Where the definition of refactoring is contested
Researchers and practitioners disagree about how strict the definition should be.
The strict view is Opdyke's and Fowler's. A refactoring preserves behavior exactly. Anything else is a different activity, and calling it refactoring hides the risk.
The Microsoft field study found that practice does not match. When asked how they define refactoring, 46% of developers did not mention preserving behavior at all. 78% described it as a change that improves some aspect of the program, such as readability, maintainability or performance. Many also said the named refactorings in Fowler's catalog are only the smallest unit of a larger restructuring effort. The paper links this to an argument by Ralph Johnson, Opdyke's own advisor, that refactoring preserves some behavior but not behavior in every respect. Timing, memory use and log output can all change while the results stay the same.
The strongest form of the loose view is this: real systems rarely have a complete definition of behavior. So "no change in behavior" always means "no change in the behavior we check".
Fowler's reply is practical. The word is useful only if it tells a reviewer how risky a change is. If "refactoring" can mean a two-week rewrite, it tells them nothing.
For a working team, the practical answer is to state which behavior is being preserved and how it is checked. "All tests pass unchanged" is one answer. GitHub's production comparison is a stronger one.
?8 questions
Questions people ask
What is refactoring in simple terms?
Who invented refactoring?
What is the difference between refactoring and rewriting?
Does refactoring change functionality?
Can you refactor without tests?
When should you not refactor?
How do you justify refactoring to a product manager?
Is refactoring the same as paying down technical debt?
§11 sources
Sources on refactoring
Fowler, M. (2018). Refactoring: Improving the Design of Existing Code, 2nd ed. Addison-Wesley. Overview and definition:
Fowler, M. (2018). The Second Edition of "Refactoring".
Fowler, M. Refactoring catalog, Extract Function.
Opdyke, W. F. (1992). Refactoring Object-Oriented Frameworks. PhD thesis, University of Illinois at Urbana-Champaign.
Show all 11 sourcesShow fewer sources
Fowler, M. (2004). Refactoring Malapropism.
Fowler, M. (2015). An example of preparatory refactoring.
Fowler, M. (2011). Opportunistic Refactoring.
Kim, M., Zimmermann, T. & Nagappan, N. (2012). A Field Study of Refactoring Challenges and Benefits. FSE 2012.
Martí, V. (2015). Move Fast and Fix Things. GitHub Blog.
Spolsky, J. (2000). Things You Should Never Do, Part I.
Fowler, M. (2024). Strangler Fig.


