Data-Centric AI Research Area
Our approach is to build a reference guide for both clinical practitioners and researchers, highlighting specific data-centric AI challenges and research opportunities.
Here, we have created a comprehensive overview graphic in order to communicate the design of data-centric ML pipelines.
Click on the specific areas in the graphic to learn more about our current work.

Data quality - Data characterisation
Unlock the power of data characterisation: Learn how techniques like Data-IQ, TRIAGE, and Datagnosis help audit datasets, boost model efficiency, and improve training by identifying easy, ambiguous, and unlearnable samples.
Explore these tools to enhance your machine learning projects.
Synthetic data for data improvement
Discover how synthetic data can enhance real datasets by augmenting small samples and ensuring fairness. By applying data characterisation techniques, we can guide synthetic data generation to reflect real-world complexities while using novel metrics to ensure quality.
Uncertainty Estimation
Uncertainty estimation is vital to ensure trustworthy predictions from ML models. In particular, we desire that models ‘know what they don’t know’.
Read more: https://www.vanderschaar-lab.com/trustworthy-predictions/
Data monitoring
Identification and appropriate handling of inconsistencies in data at deployment time is crucial to using machine learning models reliably.
Read more: https://www.vanderschaar-lab.com/data-monitoring/
Model testing — stress test scenarios and evaluations beyond subgroups
Machine learning models fail in real-world scenarios! This is especially worrying in high-stakes settings like healthcare where we might have unreliable model performance for different subgroups or across different distributions.



