DS-GA 3001 - Fall 2026
NYU Center for Data Science
2026-09-29
“Interpretability is the degree to which a human can understand the cause of a decision”
Molnar, C. (2022). Interpretable Machine Learning. 2nd Edition.
🔧 Debugging & Validation
🔬 Knowledge Discovery
🤝 Building Trust
⚖️ Compliance & Ethics


Raccuglia, P., et al. (2016). Machine-learning-assisted materials discovery using failed experiments. Nature, 533, 73-76.
Human explanations are naturally:
Molnar, C. (2022). Interpretable Machine Learning. 2nd Edition.

Intrinsic (White-box)
Post-hoc (Black-box)
Additional dimensions: Model-specific vs Model-agnostic | Local vs Global | Feature importance vs Feature effects
We discuss the following models that are intrinsically interpretable:
Linear models can be used to model the dependence of a regression target y on some features x in a format as below: \[\begin{equation} y = \beta_0 + \beta_1 x_1 + \ldots + \beta_n x_n + \varepsilon\end{equation}\]
The predicted target \(y\) is a linear combination of the weighted features \(\beta_i x_i\). The estimated linear equation is a hyperplane in the feature/target space (a simple line in the case of a single feature).
The weights specify the slope (gradient) of the hyperplane in each direction.
How do you interpret the influence of each property on the prediction of housing price?
Basic interpretation: An increase in feature \(x_j\) by one unit changes the prediction by \(\beta_j\) units
Molnar, C. (2022). Interpretable Machine Learning. Chapter 4.1.
Linear models make strong assumptions:
Linearity: Effects are additive (no interactions unless explicitly added)
Independence: Features are not strongly correlated
Homoscedasticity: Constant error variance
No multicollinearity: Correlated features can flip coefficient signs!
Example: Housing model with both “square footage” AND “number of rooms”
Molnar, C. (2022). Interpretable Machine Learning. Chapter 4.1.
Notation:
\(R^2\) (R-squared): Proportion of variance explained \[\begin{equation} R^2 = 1 - \frac{\sum (y_i - \hat{y}_i)^2}{\sum (y_i - \bar{y})^2} \end{equation}\]
Mean Square Error (MSE)/Root Mean Square Error (RMSE) \[\begin{equation} MSE = \frac{1}{N} \sum_{i=1}^{N} (y_i - \hat{y}_i)^2, \quad RMSE = \sqrt{MSE} \end{equation}\]
Mean Absolute Error (MAE) \[\begin{equation} MAE = \frac{1}{N} \sum_{i=1}^{N} |y_i - \hat{y}_i| \end{equation}\]

Core Visualizations:
Key Insights:
Mühlbacher, T., & Piringer, H. (2013). A Partition-Based Framework for Building and Validating Regression Models. IEEE TVCG. Best Paper Award, IEEE VAST 2013.
What if your dataset does not follow these assumptions?
GAMs extend linear models by replacing linear terms with flexible shape functions:
\[\begin{equation} g(\mathbb{E}[y|X]) = \beta_0 + \sum_{j=1}^{p} f_j(x_{j}) \end{equation}\]
Key idea: Replace \(\beta_j x_j\) (linear) with \(f_j(x_j)\) (flexible smooth function)
Molnar, C. (2022). Interpretable Machine Learning. Chapter 4.2.
Linear Model: \(y = \beta_j x_j\)
GAM: \(y = f_j(x_j)\)
Visualization: PDPs show \(f_j(x_j)\) - the contribution of feature \(x_j\) to the prediction across its range
GAMs use splines (piecewise polynomial functions) to approximate smooth curves:
Technical approach:
Interpretation:
Molnar, C. (2022). Interpretable Machine Learning. Chapter 4.2.
\[\begin{equation} Wage = f(year, age, education) = b_0 + f_1(year) + f_2(age) + f_3(education) \end{equation}\]

\[\begin{equation} g(\mathbb{E}[y]) = \beta_0 + \sum f_j(x_j) \end{equation}\]
\[\begin{equation} g(\mathbb{E}[y]) = \beta_0 + \sum f_j(x_j) + \sum f_{ij}(x_i, x_j) \end{equation}\]
What if we have a lot of interactions? How do we choose our interactions?
What PDPs show: The marginal effect of a feature on the predicted outcome
Mathematical idea: Average the model’s predictions across all data points while varying one feature

Molnar, C. (2022). Interpretable Machine Learning. Chapter 8.1.
Molnar, C. (2022). Interpretable Machine Learning. Chapter 8.1.
Partial dependency plot
Partial dependency plot
Partial dependency plot
Partial dependency plot
Hohman, F., Head, A., Caruana, R., DeLine, R., & Drucker, S. M. (2019). Gamut: A Design Probe to Understand How Data Scientists Understand Machine Learning Models. CHI 2019.
Drucker, S. (2020). Data Visualization: Bridging the Gap Between Users and Information. Microsoft Research Webinar.

Human-in-the-Loop Features:
Key Innovation: Bridges data-driven learning with expert knowledge through interactive visualization
Wang, Z. J., Kale, A., Nori, H., Stella, P., Nunnally, M., Chau, D. H., … & Caruana, R. (2021). GAM Changer: Editing Generalized Additive Models with Interactive Visualization. arXiv:2112.03245. Demo | Code
Decision trees recursively split data based on feature thresholds:
Prediction: Follow path from root to leaf

Example: 3 splits, 4 leaves, depth 2
Molnar, C. (2022). Interpretable Machine Learning. Chapter 4.4.
Reading a tree: “If feature \(x_j\) is [smaller/larger] than threshold \(c\) AND … then predict \(\hat{y}\)”
→ This instability motivates ensemble methods (Random Forests, Gradient Boosting) which average many trees
Trade-off: Ensembles gain accuracy but lose white-box interpretability → Need for global surrogates (discussed later)
Molnar, C. (2022). Interpretable Machine Learning. Chapter 4.4.
A decision tree of diabetes diagnosis

It shows the flow of different class, and the class distribution along the feature values.
van den Elzen, S., & van Wijk, J. J. (2011). BaobabView: Interactive Construction and Analysis of Decision Trees. IEEE VAST 2011.
iForest
Zhao, X., Wu, Y., Lee, D. L., & Cui, W. (2019). iForest: Interpreting Random Forests via Visual Analytics. IEEE TVCG, 25(1), 407-416.

Decision paths colored by class and features
Decision rules with feature splits
Definition: A decision rule is a simple IF-THEN statement consisting of a condition (antecedent) and a prediction (consequent).
Decision Trees:
Rule Systems:
Implication: Rule systems need strategies to handle overlapping rules (majority vote, highest confidence, first match)
Evaluation Metrics:
Learning Approaches:
Molnar, C. (2022). Interpretable Machine Learning. Chapter 4.7.
Clearly see how the decision is made and which rule is more important.

The final decision is made based on a voting mechanism.
A recent user study shows that “if-then structure without any connecting else statements enables users to easily reason about the decision boundaries of classes.”
Disjunctive normal form (DNF, OR-of-ANDs) Conjunctive normal form (CNF, AND-of-ORs)
What form does this rule set follow?

Research Questions:
Can different visualizations of rules lead to different levels of understanding?
What visual factors influence understanding and how do they affect rule comprehension?
Key findings: Visual encoding choices significantly impact interpretability
Yuan, J., Nov, O., & Bertini, E. (2021). An Exploration and Validation of Visual Factors in Understanding Classification Rule Sets. arXiv:2109.09160.
Given a rule below:
If \(X\), then class \(Y\).
Support / Coverage of a rule:
\[\begin{equation} \text{Support} = \frac{\text{number of instances that match the conditions in } X}{\text{total number of instances}} \end{equation}\]
Confidence / Accuracy of a rule:
\[\begin{equation} \text{Confidence} = \frac{\text{number of instances that match conditions in } X \text{ and belong to class } Y}{\text{number of instances that match conditions in } X} \end{equation}\]
Imagine that we have a black-box model (too complex to understand the internal structure), can we use white-box models to help us understand the model behavior of the black-box model?

Open the black box by understanding a “surrogate model” that approximate the behavior of the original black-box model.

The Fidelity-Interpretability Trade-off: A fundamental challenge in XAI where increasing model interpretability often decreases fidelity to the original model’s behavior
What you want:

Simple, interpretable surrogate with high fidelity to the black-box model
What you get:

Either low fidelity (simple but inaccurate) or low interpretability (accurate but complex)
RuleMatrix
Ming, Y., Qu, H., & Bertini, E. (2019). RuleMatrix: Visualizing and Understanding Classifiers with Rules. IEEE TVCG, 25(1), 342-352.
Explainable Matrix
Popolin Neto, M., & Paulovich, F. V. (2021). Explainable Matrix – Visualization for Global and Local Interpretability of Random Forest Classification Ensembles. IEEE TVCG, 27(2), 1427-1437.
Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Vaughan, J. W., & Wallach, H. (2021). Manipulating and Measuring Model Interpretability. CHI 2021.
Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5), 206-215.
Chung, Y., Kraska, T., Polyzotis, N., Tae, K. H., & Whang, S. E. (2019). Slice Finder: Automated Data Slicing for Model Validation. IEEE ICDE 2019, 1550-1553. Extended version: arXiv:1807.06068.

How about we use whether the model prediction is wrong or not to train a “surrogate tree”?
InterpretML: https://github.com/interpretml/interpret
Notebook: https://colab.research.google.com/drive/1nKE6WIApebHi67yfhH6k5mZN86evLZOM?usp=sharing
Some other libraries for PDP visualization: https://scikit-learn.org/stable/modules/partial_dependence.html https://interpret.ml/docs/pdp.html
Notebook: https://colab.research.google.com/drive/12LV2Z_1BbP3efACYp2QxzsPaOrIn8a8l?usp=sharing