011 min
Why the confusion matrix matters
A single score such as "98.8% correct" hides which mistakes a classifier makes. A classifier can make two different kinds of mistake. It can flag something that is fine, or it can miss something that is not. These mistakes usually cost different amounts, and one number cannot show both.
The confusion matrix keeps the two mistakes apart. A PM deciding whether to ship a fraud detector, a spam filter or a content moderation model needs to know how many real cases it misses and how many good cases it blocks. The matrix answers both questions with plain counts.
The decision it supports is concrete. If a team reads only accuracy, it can ship a model that finds none of the cases it was built to find. Section "A numeric example" below shows how this happens with a model that looks 98.8% correct.
022 min
The four cells of a confusion matrix
A binary classifier answers yes or no for each item. The positive class is the answer the team cares about finding, such as "fraud" or "spam". The negative class is everything else. For each item, the matrix compares the classifier's answer with the true label. That gives four outcomes:
- True positive (TP): the classifier said positive and the item is positive.
- False positive (FP): the classifier said positive but the item is negative. This is also called a false alarm or a Type I error.
- False negative (FN): the classifier said negative but the item is positive. This is also called a miss or a Type II error.
- True negative (TN): the classifier said negative and the item is negative.
The four counts add up to the total number of items tested. Powers (2011) defines true and false positives as the numbers of predicted positives that were correct and incorrect, and the four cells sum to the total number of cases.
The two cells on the diagonal, TP and TN, are the correct decisions. The other two cells are the errors. Fawcett (2006) puts it the same way: the major diagonal holds the correct decisions, and the other cells hold the errors, which he calls the confusion between classes. That is where the name comes from.
| Actual positive | Actual negative | |
|---|---|---|
| Predicted positive | TP | FP |
| Predicted negative | FN | TN |
Use this layout as the default, because it is the one Fawcett and Powers both draw, with predictions as rows. The next section explains why you must check it before you read any matrix.
032 min
How to read the matrix: check the axes first
There is no single standard for which axis holds the true labels. Libraries and papers disagree, and a matrix read with the axes swapped gives swapped conclusions: false positives look like false negatives.
scikit-learn, a widely used Python machine-learning library, puts the true labels on the rows. Its documentation says the function computes the confusion matrix with each row corresponding to the true class, and warns that Wikipedia and other references may use a different convention for axes. In scikit-learn, entry (i, j) is the number of items that are actually in class i but were predicted to be class j. For two classes, the library returns the cells in the order TN, FP, FN, TP when the matrix is flattened.
Powers (2011) and Fawcett (2006) put the predictions on the rows instead. Both layouts are correct. They are transposes of each other.
Two habits prevent mistakes:
- Look for axis titles before you read any number. If they are missing, ask whoever made the table.
- State which class is positive. For a fraud model, positive means fraud. If another team flips the labels so that positive means legitimate, every cell moves.
scikit-learn can also report ratios instead of counts. The documentation says the confusion matrix can be normalized three ways, by the sum of each column, each row, or the entire matrix. Dividing each row by its total gives, for the positive row, the share of real positives that were found. Dividing each column by its total gives the share of the classifier's flags that were correct. Those are two different questions, so be clear about which one a normalized table answers.
041 min
Where the confusion matrix comes from
The idea is older than machine learning. It is a contingency table, the standard way statisticians cross-tabulate two categorical variables. Here one variable is the true class and the other is the predicted class. Fawcett (2006) calls the two-by-two form a confusion matrix and notes it is also called a contingency table.
Two papers are the practical references for how the matrix is used to judge classifiers. Fawcett's "An introduction to ROC analysis" appeared in Pattern Recognition Letters in 2006. It shows the matrix with the common metrics calculated from it, and then builds the ROC graph from two of those metrics. Powers' "Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation" was published in the Journal of Machine Learning Technologies in 2011. It starts from the same four counts and argues about which measures to trust. The copy cited here is Powers' own upload to arXiv in 2020.
This article does not attribute the name "confusion matrix" to one first author. The sources checked here use the term without saying who coined it, and a claim about the first use would be unverified.
052 min
A numeric example: a fraud detector
All numbers here are illustrative. They are not from a real system. They are chosen so the arithmetic is easy to follow.
A subscription app tests a fraud detector on 10,000 transactions. 100 are fraudulent and 9,900 are legitimate. The detector flags 160 transactions. Of those, 70 are really fraud and 90 are not.
That fixes all four cells:
- TP = 70, the fraud cases it flagged.
- FP = 90, the legitimate transactions it flagged.
- FN = 30, the fraud cases it missed (100 fraud minus 70 found).
- TN = 9,810, the legitimate transactions it left alone (9,900 minus 90).
| Actual fraud | Actual legitimate | Row total | |
|---|---|---|---|
| Flagged as fraud | 70 | 90 | 160 |
| Not flagged | 30 | 9,810 | 9,840 |
| Column total | 100 | 9,900 | 10,000 |
Now calculate the common metrics from the four counts:
- Accuracy = (TP + TN) / total = (70 + 9,810) / 10,000 = 98.8%.
- Precision = TP / (TP + FP) = 70 / 160 = 43.8%. Fewer than half of the flags are real fraud.
- Recall = TP / (TP + FN) = 70 / 100 = 70%. The detector finds seven in ten fraud cases. Recall is also called sensitivity and the true positive rate.
- Specificity = TN / (TN + FP) = 9,810 / 9,900 = 99.1%. It finds 99.1% of the legitimate transactions as legitimate. The false positive rate is the remainder, 90 / 9,900 = 0.9%.
- F1 score = 2 x precision x recall / (precision + recall) = 2 x 0.4375 x 0.70 / 1.1375 = 0.54.
The accuracy of 98.8% sounds excellent. Now compare a detector that never flags anything. It has TP = 0 and FP = 0, so its accuracy is 9,900 / 10,000 = 99.0%. It is more "accurate" than the real detector and finds no fraud at all. Accuracy cannot tell these two apart. The confusion matrix can, because its FN cell shows 100 missed cases.
The matrix also shows the trade-off the team has to price. Each of the 90 false positives is a customer whose payment is blocked or reviewed. Each of the 30 false negatives is a fraud loss. The right balance depends on those two costs, and neither is a statistical question.
062 min
How the threshold changes the four counts
Most classifiers do not output yes or no directly. They output a score, and the team picks a threshold: items scoring above it are flagged. A confusion matrix describes the classifier at one threshold only. Change the threshold and the four counts change.
Take the same detector and the same 10,000 transactions. Illustratively, lower the threshold so more transactions are flagged:
| Threshold | TP | FP | FN | TN | Precision | Recall | Accuracy |
|---|---|---|---|---|---|---|---|
| Strict (before) | 70 | 90 | 30 | 9,810 | 43.8% | 70% | 98.8% |
| Loose (after) | 85 | 400 | 15 | 9,500 | 17.5% | 85% | 95.85% |
The loose setting finds 15 more fraud cases and flags 310 more legitimate ones. Recall rises from 70% to 85%. Precision falls from 43.8% to 17.5%. Accuracy falls from 98.8% to 95.85%, even though the detector now catches more fraud.
This is why one matrix is a snapshot, not a verdict on the model. If someone reports a single matrix, ask what threshold produced it.
Plotting recall against the false positive rate for every threshold gives the ROC curve. Each point on the curve is the true positive rate and false positive rate from one confusion matrix. Fawcett (2006) writes that ROC curves are insensitive to changes in class distribution. Each axis of the curve divides within one column of the matrix, so the class mix cancels out. That property matters in the next section.
072 min
Same detector, different mix of cases
The next example changes one variable. Keep the detector's behaviour fixed: it finds 70% of fraud and wrongly flags 0.9% of legitimate transactions. Now suppose fraud is 10% of the transactions, not 1%. That is 1,000 fraud cases and 9,000 legitimate ones.
| Fraud share | TP | FP | FN | TN | Precision | Recall | Accuracy |
|---|---|---|---|---|---|---|---|
| 1% (100 fraud) | 70 | 90 | 30 | 9,810 | 43.8% | 70% | 98.8% |
| 10% (1,000 fraud) | 700 | 82 | 300 | 8,918 | 89.5% | 70% | 96.2% |
The detector did not change. Recall stayed at 70% and the false positive rate stayed at 0.9%. Precision doubled, from 43.8% to 89.5%, and accuracy dropped by more than two points. Only the mix of cases changed.
The reason is where each metric draws its numbers. Recall uses only the actual-positive column. The false positive rate uses only the actual-negative column. Precision uses a row, which mixes both columns. Fawcett (2006) notes that metrics such as accuracy, precision, lift and F score use values from both columns of the confusion matrix. Those metrics change when the class mix changes, even when the classifier does not.
The practical rule: before comparing a metric from a test set to a metric from production, check that both have the same share of positives. A precision of 89.5% measured on a balanced test set does not predict precision on live traffic with 1% fraud. The related problem of very uneven classes is covered under class imbalance.
082 min
More than two classes
A confusion matrix is not limited to yes and no. For n classes it has n rows and n columns. Fawcett (2006) describes the n by n matrix as holding n correct classifications on the major diagonal and the remaining entries as errors, so the number of error cells grows quickly as classes are added.
The scikit-learn documentation gives a small checkable example with three classes. The true labels are 2, 0, 2, 2, 0, 1 and the predictions are 0, 0, 2, 2, 0, 2. The function returns this matrix, with true classes as rows:
| True \ Predicted | 0 | 1 | 2 |
|---|---|---|---|
| 0 | 2 | 0 | 0 |
| 1 | 0 | 0 | 1 |
| 2 | 1 | 0 | 2 |
Reading it takes three steps:
- The diagonal is 2, 0 and 2. Four of six items were classified correctly.
- Row 1 is the one item whose true class is 1. It was predicted as class 2. The classifier never predicted class 1 at all, so column 1 is empty.
- The one error in row 2 is an item of true class 2 that was predicted as class 0.
For a multi-class matrix, precision and recall are calculated per class by treating that class as positive and all others as negative. For class 2, the true positives are the 2 on the diagonal. The false negatives are the other cells in its row (1), and the false positives are the other cells in its column (1). So recall for class 2 is 2 / 3 and precision is 2 / 3. Class 1 has recall 0, because its only item was missed.
A matrix like this shows which classes get confused with which. A support team that routes tickets into categories can see that billing tickets are sent to the account category. That tells them where to add training examples. A single accuracy figure would only say that some tickets are misrouted.
092 min
How teams use the confusion matrix
Different roles read the same four cells for different decisions.
- A PM uses it to set the operating point. If a false positive costs a support call and a false negative costs a lost customer, the PM picks the threshold where the total cost is lowest. The matrix at each candidate threshold turns that into a count of calls and losses.
- An engineer uses it as a regression test. After retraining, compare the new matrix to the old one on the same test set. If accuracy is unchanged but false negatives went from 30 to 45, the model got worse on the cases that matter.
- A designer uses it to decide what the interface says. A high false positive rate means a flag should be worded as a suggestion with an easy way to dismiss it. A high false negative rate means users cannot rely on the absence of a flag.
- A founder uses it to check a vendor's claim. Ask for the matrix on a test set with the same share of positives as your real data. A claim of "95% accurate" without the matrix means little.
Libraries produce the matrix in one call. In scikit-learn the function is confusion_matrix(y_true, y_pred). The documentation shows the binary case: for the labels [0, 0, 0, 1, 1, 1, 1, 1] and the predictions [0, 1, 0, 1, 0, 1, 0, 1], flattening the result gives TN = 2, FP = 1, FN = 2 and TP = 3. That small case is easy to verify by hand and is a good first test of whether you have the axes right.
101 min
When the confusion matrix misleads
The matrix is exact about what it counts, but it has limits.
- It describes one threshold. A different threshold gives a different matrix, as the threshold section showed. Comparing two models at their default thresholds can favour the one whose default happens to be set better.
- It depends on the class mix. The test set must resemble production. A matrix from a set with 50% positives does not describe traffic with 1% positives.
- It treats all errors in one cell as equal. Ten false positives on low-value cases and ten on high-value cases look the same. If error costs vary a lot, weight the cells by cost.
- It does not say why the errors happen. The matrix shows where the model is wrong. Finding the cause needs a look at the individual items in the FP and FN cells.
- The derived metrics have known biases. Powers (2011) argues that recall, precision, F-measure and Rand accuracy are biased and should not be used without understanding the biases. Powers adds that a system that does worse on his chance-corrected measure, informedness, can appear better under those common measures. Report the four counts, not only one ratio from them.
- A small test set gives unstable cells. If there are only 20 real positives, one extra miss moves recall by 5 points. Report the counts so readers can see how small they are.
111 min
Confusion matrix versus ROC curve and accuracy
A confusion matrix is often mixed up with the metrics and graphs built from it.
| Confusion matrix | Accuracy | ROC curve | |
|---|---|---|---|
| What it is | Four counts | One ratio from the counts | A curve over many thresholds |
| Thresholds covered | One | One | All |
| Shows the type of error | Yes | No | Partly |
The matrix is the raw table. Accuracy, precision, recall and the F1 score are single numbers calculated from it. The ROC curve is a set of points, each one taken from the matrix at a different threshold. Use the matrix when you need to see the counts at the threshold you will actually ship. Use the curve when you need to compare models across thresholds.
?8 questions
Questions people ask
What is a confusion matrix?
How do you read a confusion matrix?
What is the difference between a false positive and a false negative?
How do you calculate accuracy, precision and recall from a confusion matrix?
Why can a model with high accuracy be useless?
Does the confusion matrix work with more than two classes?
Does the threshold change the confusion matrix?
What is the difference between a confusion matrix and an ROC curve?
§4 sources
Sources
Powers, D. M. W. (2011). Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation. Journal of Machine Learning Technologies 2(1), 37-63.
Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters 27(8), 861-874. (open copy used for quotations: )
scikit-learn developers. sklearn.metrics.confusion_matrix.
scikit-learn developers. Metrics and scoring: quantifying the quality of predictions, confusion matrix section.


