
DS-GA 3001 - Fall 2026
NYU Center for Data Science
2026-09-22
Confusion Matrices and ROC Curves
Visual Analytics Systems for Model Performance
Calibration
Consider a disease prediction model. Suppose the hypothetical disease has a 5% prevalence in the population
The given model converges on the solution of predicting that nobody has the disease (i.e., the model predicts “0” for every observation)
Our model is 95% accurate
Yet, public health officials are stumped
Consider a model to identify handwritten digits. All digits are equally probable and equally represented in the training and test datasets.
The model correctly identifies all of the digits, except for digit \(5\), classifying half of the \(5\)s samples as \(6\) and the other half is correctly identified
The accuracy of this model is \(95\%\). Is this information enough to determine whether the model is good or not?

Table from Wikipedia: Confusion matrix

scikit-learn digits: label spreading trained on 40 labelled images, tested on 300 (accuracy 90%)
Reading it: rows are the true label, columns the predicted label. The diagonal is correct.
Pros
Cons
Squares, Neo and the confusion wheel, later today, each attack one of these.

from sklearn import datasets
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
X,y = datasets.make_classification(5000, 10, random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
clf = LogisticRegression(random_state=0)Görtler, J., Hohman, F., Moritz, D., et al. (2022). Neo: Generalizing Confusion Matrix Visualization to Hierarchical and Multi-Output Labels. CHI.
ROC analysis is another way to assess a classifier’s output
ROC analysis developed out of radar operation in the second World War, where operators were interested in detecting signal (enemy aircraft) versus noise
We create an ROC curve by plotting the true positive rate (TPR) against the false positive rate (FPR) at various thresholds
True Positive Rate (TPR), also known as Recall or Sensitivity, is the proportion of actual positives that are correctly identified as such (TP / (TP + FN)).
False Positive Rate (FPR) is the proportion of actual negatives that are incorrectly identified as positive (FP / (FP + TN)).




Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861-874.


Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861-874.
Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861-874.
Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861-874.

Micro-average: Pool every one-vs-rest decision into a single curve. Each sample counts equally, so the large classes dominate.
Macro-average: Compute a curve per class, then average them. Each class counts equally, so a small, badly served class still shows up.
Davis, J., & Goadrich, M. (2006). The relationship between Precision-Recall and ROC curves. ICML. Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot on imbalanced datasets. PLOS ONE.
Ren, D., Amershi, S., Lee, B., Suh, J., & Williams, J. D. (2016). Squares: Supporting interactive performance analysis for multiclass classifiers. IEEE TVCG.
Alsallakh, B., Hanbury, A., Hauser, H., Miksch, S., & Rauber, A. (2014). Visual methods for analyzing probabilistic classification data. IEEE TVCG.
Alsallakh, B., Hanbury, A., Hauser, H., Miksch, S., & Rauber, A. (2014). Visual methods for analyzing probabilistic classification data. IEEE TVCG.
Beauxis-Aussalet, E., & Hardman, L. (2014). Visualization of confusion matrix for non-expert users. IEEE VAST.

Beauxis-Aussalet, E., & Hardman, L. (2014). Visualization of confusion matrix for non-expert users. IEEE VAST.
Talbot, J., Lee, B., Kapoor, A., & Tan, D. S. (2009). EnsembleMatrix: interactive visualization to support machine learning with multiple classifiers. CHI.
When performing classification, we often are interested not only in predicting the class label, but also in the probability of the output
This probability gives us a kind of confidence score on the prediction
However, a model can separate the classes well (having a good accuracy/AUC), but be poorly calibrated. In this case, the estimated class probabilities are far from the true class probabilities
We can calibrate the model, changing the scale of the predicted probabilities
Weather forecasters started thinking about calibration a long time ago (Brier, 1950): a forecast of “70% chance of rain” should be followed by rain 70% of the time. Let’s consider a small toy example:

This forecast is doing well at predicting the rain:
Slides based on classifier-calibration.github.io


Slides based on classifier-calibration.github.io


Slides based on classifier-calibration.github.io


Slides based on classifier-calibration.github.io
import matplotlib.pyplot as plt
from sklearn.calibration import CalibrationDisplay
fig = plt.figure()
ax = fig.add_subplot(111)
CalibrationDisplay.from_estimator(lg, X_test, y_test, n_bins=10, ax=ax,
label='Logistic Regression')
CalibrationDisplay.from_estimator(nb, X_test, y_test, n_bins=10, ax=ax,
label='Naive Bayes')
Underconfidence: a classifier thinks it’s worse at separating classes than it actually is.
Overconfidence: a classifier thinks it’s better at separating classes than it actually is
A classifier can be overconfident for one class and underconfident for the other
Slides based on classifier-calibration.github.io

Let \(N\) be the total number of samples, \(B\) the number of bins, \(n^b\) the samples in bin \(b\), \(acc(b)\) the fraction of positives in bin \(b\), and \(conf(b)\) the average predicted probability in bin \(b\).
\(ECE = \sum_{b=1}^B \frac{n^b}{N}|acc(b) - conf(b)|\)
\(MCE = \underset{b \in \{1,2,\dots,B\}}{\text{max}} |acc(b) - conf(b)|\)
Predict the base rate for everyone: “there is a 5% chance you have the disease”, for every patient alike
Every group’s observed frequency matches its prediction, so ECE = 0 and the reliability diagram sits on the diagonal
Yet it ranks nobody above anybody: AUC = 0.5, and it never identifies a single patient
Calibration is being right on average; sharpness is committing to a prediction. A useful model needs both
Recalibration can repair a sharp but miscalibrated model. Nothing repairs an unsharp one

Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. ICML.


Architecture and pre-training, not size, drive calibration: ViT and MLP-Mixer are both the most accurate and among the best calibrated. In-distribution, calibration degrades slightly with size; under distribution shift the effect reverses
Image taken from Minderer, Matthias, et al. (2021). Revisiting the calibration of modern neural networks. Advances in Neural Information Processing Systems.
The observation held: those CNNs really were overconfident, and temperature scaling really does fix them cheaply
The extrapolation did not. A benchmark over 117,000 architectures finds no universal accuracy-calibration trade-off, and attributes “calibration decreases with model size” to survivorship bias in the 2017 model grid
The direction has flipped for foundation models: ConvNeXt, EVA and BEiT come out underconfident in-distribution, though better calibrated under distribution shift
For LLMs: pre-trained next-token probabilities are close to calibrated; SFT and RLHF degrade that. For RLHF’d models, asking for a stated confidence often beats the token probabilities
So: measure calibration for your model on your distribution. Do not inherit a trend from a paper
Tao, L. et al. (2024). A Benchmark Study on Calibration. ICLR. Hekler, A. et al. (2025). Beyond Overconfidence. arXiv preprint. Tian, K. et al. (2023). Just Ask for Calibration. EMNLP.
Proper scoring rules are calculated at the observation level, whereas ECE is binned
Think of them as putting each item in its separate bin, then computing the average of some loss for each predicted probability and its corresponding observed label

Slides based on classifier-calibration.github.io
You believe \(q\), you report \(p\). Your expected score is \(\ q\,S(p,1) + (1-q)\,S(p,0)\)
Proper: \(p = q\) is optimal. Strictly proper: \(p = q\) is the only optimum, so honesty is uniquely best
Accuracy is not strictly proper. Believing 0.6, claiming \(1.0\) scores exactly the same: every report above \(0.5\) ties. It cannot tell \(0.51\) from \(0.99\)
Absolute error \(|p - y|\) is worse: its expectation \(q + p(1 - 2q)\) is linear in \(p\), so the unique optimum is \(p = 0\) or \(1\). It pays you to overclaim
A metric that is indifferent to your uncertainty, or rewards overstating it, cannot judge probabilities. Next: two that do neither
\[BS = \frac{1}{N} \sum_{i=1}^N (\hat{y}_i - y_i)^2\]
Log-loss/Cross entropy:
Frequently used as the training loss of machine learning methods, such as neural networks
Only penalises the probability given to the true class
\[LL = -\frac{1}{N} \sum_{i=1}^N [y_i \text{log}(\hat{y}_i) + (1-y_i)\text{log}( 1 - \hat{y}_i)]\]
Slides based on classifier-calibration.github.io
An intuitive way to decompose proper scoring rules is into refinement and calibration losses
Refinement loss: is the loss due to producing the same probability for instances from different classes
Calibration loss: is the loss due to the difference between the probabilities predicted by the model and the proportion of positives among instances with the same output
Slides based on classifier-calibration.github.io
Parametric calibration involves modelling the score distributions within each class
Platt scaling: Logistic calibration can be derived by assuming that the scores within both classes are normally distributed with the same variance (Platt, 2000)
Beta calibration: employs Beta distributions instead, to deal with scores already on a [0, 1] scale (Kull et al., 2017)
Dirichlet calibration for more than two classes (Kull et al., 2019)
Non-parametric calibration often ignores scores and employs ranks
Slides based on classifier-calibration.github.io
\(p(y_i = 1 | \hat{y}_i) = \frac{1}{1 + exp(\mathbf{A}\hat{y}_i + \mathbf{b})}\)
Works best if the calibration error is symmetrical (classifier output for each binary class is normally distributed with the same variance)
This can be a problem for highly imbalanced classification problems, where outputs do not have equal variance
In general it is most effective when the un-calibrated model is under-confident and has similar calibration errors for both high and low outputs
Fits a non-parametric isotonic regressor, which outputs a step-wise non-decreasing function
Isotonic regression is more general when compared to Platt scaling, as the only restriction is that the mapping function is monotonically increasing
Is more powerful as it can correct any monotonic distortion of the un-calibrated model
However, it is more prone to overfitting, especially on small datasets
\[\hat{q}_i = \max_k \ \sigma_{SM}(z_i / T)^{(k)}\]
\(T > 1\) softens the distribution (raises its entropy); \(T = 1\) leaves the network unchanged; \(T \to \infty\) gives uniform
\(T\) is fitted by minimising NLL on a held-out validation set. The network’s weights are never touched
Dividing every logit by the same number cannot change which one is largest, so accuracy is unchanged
One parameter, so it rarely overfits: the special case of Platt scaling used for deep nets today
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. ICML.


Reliability diagrams are a standard way to visualize calibration
ECE is a summary of what reliability diagrams show
Proper scoring rules (Log loss, Brier score) measure different aspects of probability correctness
However, proper scoring rules cannot tell us where a model is miscalibrated
Calibration is not quality: the base-rate predictor scores ECE 0. Always report it beside AUC or a proper scoring rule
Every number here depends on the binning, which papers rarely report. The rest of this section is about that
Image taken from Xenopoulos, P., Rulff, J., Nonato, L. G., Barr, B., & Silva, C. (2022). Calibrate: Interactive analysis of probabilistic model output. IEEE Transactions on Visualization and Computer Graphics.

Xenopoulos, P., Rulff, J., Nonato, L. G., Barr, B., & Silva, C. (2022). Calibrate: Interactive analysis of probabilistic model output. IEEE TVCG.
Xenopoulos, P., Rulff, J., Nonato, L. G., Barr, B., & Silva, C. (2022). Calibrate: Interactive analysis of probabilistic model output. IEEE TVCG.
Xenopoulos, P., Rulff, J., Nonato, L. G., Barr, B., & Silva, C. (2022). Calibrate: Interactive analysis of probabilistic model output. IEEE TVCG.
Błasiok, J., & Nakkiran, P. (2023). Smooth ECE: Principled Reliability Diagrams via Kernel Smoothing. arXiv preprint arXiv:2309.12236.
Błasiok, J., & Nakkiran, P. (2023). Smooth ECE: Principled Reliability Diagrams via Kernel Smoothing. arXiv preprint arXiv:2309.12236.
Vaicenavicius, J., Widmann, D., Andersson, C., et al. (2019). Evaluating model calibration in classification. AISTATS.
Niculescu-Mizil, A., & Caruana, R. (2005). Predicting good probabilities with supervised learning. ICML.
Nixon, J., Dusenberry, M. W., Zhang, L., Jerfel, G., & Tran, D. (2019, June). Measuring Calibration in Deep Learning. In CVPR Workshops (Vol. 2, No. 7).
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. ICML.
Vaicenavicius, J., Widmann, D., Andersson, C., Lindsten, F., Roll, J., & Schön, T. (2019, April). Evaluating model calibration in classification. In The 22nd International Conference on Artificial Intelligence and Statistics (pp. 3459-3467). PMLR.
Kull, M., & Flach, P. (2015). Novel decompositions of proper scoring rules for classification: Score adjustment as precursor to calibration. ECML-PKDD, pp. 68-85.
Silva Filho, T. et al. (2020). ECML/PKDD 2020 Tutorial: Evaluation metrics and proper scoring rules. Full tutorial at classifier-calibration.github.io.
Minderer, M. et al. (2021). Revisiting the calibration of modern neural networks. NeurIPS.
Xenopoulos, P., Rulff, J., Nonato, L. G., Barr, B., & Silva, C. (2022). Calibrate: Interactive analysis of probabilistic model output. IEEE TVCG.
Błasiok, J., & Nakkiran, P. (2023). Smooth ECE: Principled reliability diagrams via kernel smoothing. arXiv. Python package: relplot