van der Schaar Lab

Interpretable machine learning

What do we mean with interpretable?

There are several reasons to make a “black box” machine learning model interpretable.

First, an interpretable output can be more readily understood and trusted by its users (for example, clinicians deciding whether to prescribe a treatment), making its outputs more actionable.

Second, a model’s outputs often need to be explained by its users to the subjects of its outputs (for example, patients deciding whether to accept a proposed treatment course).

Third, by uncovering valuable information that otherwise would have remained hidden within the model’s opaque inner workings, an interpretable output can empower users such as researchers with powerful new insights.

The value of interpretability as a broad concept is, therefore, clear. Yet despite this, the meaning of the term itself is too seldom discussed and too often oversimplified. There is no single “type” of interpretability, after all, since there are many potential ways to extract and present information from the output of a model, and many types of information to choose to extract.

Current Research Highlights

Inspiration Exchange Session 15

Inspiration Exchange Session 21

Revolutionizing Healthcare: making ML output useful and actionable for clinicians and researchers

Our powerful software repository: Interpretability Suite

Multiple stakeholders drive diverse interpretability requirements for machine learning in healthcare (2023)

Our Research

Our lab has been researching interpretability methods and approaches (for application in healthcare and beyond) for many years. Our work so far has led us to a unique but powerful framework for considering the multiple types of interpretability.

Each of these types of interpretability represents a distinct set of challenges from a model development perspective and can benefit different users in a variety of applications.

Type 1 interpretability: feature importance

Type 2 interpretability: similarity classification

Type 3 interpretability: transparent mathematical equations

Type 4 interpretability: concept-based explainability

Robust and trustworthy interpretations