This MICCAI Tutorial will be presented by Mihaela van der Schaar, Nabeel Seedat, and Camila González (see presenter bios below) during the 27th INTERNATIONAL CONFERENCE ON MEDICAL IMAGE COMPUTING
AND COMPUTER ASSISTED INTERVENTION, running from 6 – 10 October, 2024.
Title
Clinical AI in the Real-World: From Data-Centric AI to Dynamic Learning
Logistics
The live tutorial will take place on 6 October (13:30 – 18:00 PM), in Marrakesh. It will be located in the Topaze room of Palmeraie Palace.
About
Recent advances in AI have brought tremendous progress, yet also revealed new challenges around data and model utilization over time. This tutorial examines two emerging and important areas – data-centric AI and continual learning – unpacking issues and state-of-the-art solutions tailored for medical imaging contexts.
In part one, we will cover Data-centric AI which seeks to place the previously undervalued nuances of data at the center of AI development and articulate its transformative potential. We will explore the motivation behind the data-centric approach, highlighting the power to improve model performance, as well as engender more trustworthy, fair, and unbiased AI systems. Our examination extends to standardized documentation frameworks, exposing how they form the backbone of this new paradigm. We will cover state-of-the-art methodologies in (1) data characterization to audit datasets, (2) synthetic data and (3) data-centric AI in the era of Foundation models.
In part two of the tutorial, we move from a static perspective to the dynamic nature of medical AI systems, extending our view to systems that adapt over their lifetime in situations where we have spatial and temporal data availability constraints. We will outline the process of building and deploying federated and continual learning medical AI systems, with a focus on leveraging trained Foundation models. These techniques could extend the lifespan of medical software solutions but also signify technical and regulatory challenges.
Our integrated tutorial will equip participants with a comprehensive understanding and practical skills around two important real-world medical AI challenges. The tutorial aims to provide an interactive and hands-on experience via software tools and interactive coding sessions, thereby enabling practical engagement for participants.
Goals of the lab
Machine Learning (ML) models, despite their prowess, often encounter challenges when deployed in the real world. Model failure cases are especially consequential in high-stakes domains like healthcare. While it is well-known that the quality of data used to train Machine Learning models is crucial to their success or failure, it is often undervalued. The richness and diversity of the training set become particularly important in dynamic scenarios, where the data distribution shifts as time goes on. When this happens, models must incorporate new developments without forgetting previous knowledge. The emergence of Data-Centric AI gives the data used in AI/ML and its quality center stage and seeks to develop tools for systematic characterization, evaluation, and monitoring of the data used to train and evaluate ML models.
We will address the following topics:
Motivate and introduce Data-centric AI, a topic of emerging importance for AI.
Familiarize attendees with the challenges that we encounter when deploying ML models in dynamic, changing, scenarios; and present solutions from the field of federated and continual learning.
Increase hands-on practical coding experience via demonstrations with recent Data-centric and Continual AI tools/methods.
By the end of this tutorial, participants will have a comprehensive understanding of the key concepts in Data-Centric and Dynamic AI, and be equipped with knowledge of state-of-the-art tools and what lies ahead. Examples will focus on the healthcare context given the presenters’ expertise and MICCAI’s focus. Importantly, the coding demos will provide participants with hands-on experience.
Presenter bios
The tutorial will be presented by Mihaela van der Schaar, Nabeel Seedat, and Camila González.

Mihaela van der Schaar
Mihaela van der Schaar is the John Humphrey Plummer Professor of Machine Learning, Artificial Intelligence and Medicine at the University of Cambridge. In addition to leading the van der Schaar Lab, Mihaela is founder and director of the Cambridge Centre for AI in Medicine (CCAIM).
Mihaela was elected IEEE Fellow in 2009 and Fellow of the Royal Society in 2024. She has received numerous awards, including the Johann Anton Merck Award (2024), the Oon Prize on Preventative Medicine from the University of Cambridge (2018), a National Science Foundation CAREER Award (2004), 3 IBM Faculty Awards, the IBM Exploratory Stream Analytics Innovation Award, the Philips Make a Difference Award and several best paper awards, including the IEEE Darlington Award. She was a Turing Fellow at The Alan Turing Institute in London between 2016 and 2024.
Mihaela is personally credited as inventor on 35 USA patents (the majority of which are listed here), many of which are still frequently cited and adopted in standards. She has made over 45 contributions to international standards for which she received 3 ISO Awards. In 2019, a Nesta report determined that Mihaela was the most-cited female AI researcher in the UK.

Nabeel Seedat
Nabeel Seedat is a PhD candidate at the University of Cambridge. Nabeel’s research is focused on Data-Centric AI, uncertainty quantification and synthetic data. He has published papers on Data-Centric AI in leading ML conferences including, NeurIPS, ICML and AISTATS. Nabeel has recently given talks and presentations on Data-Centric AI to both industry: AstraZeneca, Discovery Limited and academic research groups: Queen Mary University of London, University of Cape Town). He also has experience giving talks to diverse audiences at conferences including IEEE conferences, KDD and PyData.
Beyond Nabeel’s academic background in data-centric AI, he also has extensive industry experience working on data-centric problems. He has worked as a Machine Learning engineer across two multinational corporations (in the USA and South Africa), building real-world computer vision and NLP systems that currently serve millions of customers daily.

Camila González
Camila is a postdoctoral scholar at the Computational Neuroscience Laboratory at Stanford University, where she develops continual learning methods suitable for dynamic settings with ongoing data collection. Last year, she co-organized the first MICCAI tutorial on Dynamic AI in the Clinical Open World. She also presented her translational research at the ContinualAI Society seminar series and helped organize the first ContinualAI Unconference from her roles as Diversity, Equity, and Inclusion (DEI) chair and session chair. Additionally, she participated in an expert panel on Applications of Continual Learning. Her work has received multiple distinctions, including the MICCAI Young Scientist Award, the Francois Erbsmann Award, and the Best Presentation Award at the EuSoMII annual meeting. She has been featured in outlets such as the Computer Vision News magazine and the AI-Ready Healthcare podcast. Outside her research, she presided over the MICCAI student board for two years.









