A collection of our research, publications and resources exploring how AI is transforming clinical trials.
Overview
Clinical trials are essential for bringing new treatments to patients, but they remain one of the most challenging stages of medical innovation. They are often slow, expensive and operationally complex, making it difficult to identify early which patients are most likely to benefit, which may be harmed, and which trial designs are most likely to succeed. The problem is not simply that trials need to be faster. The deeper challenge is that clinical development remains limited in its ability to continuously learn and adapt from accumulating evidence.
For more than a decade, the van der Schaar Lab has been developing the machine learning foundations needed to change this. Our work has spanned causal AI, digital twins, synthetic data, adaptive clinical trials, AI for pharmacology, individualized treatment-effect estimation, and, more recently, agentic AI systems for clinical development. Together, these advances contribute to a new generation of AI-enabled clinical trials that are more adaptive, personalised and efficient.
This page brings together a selection of our publications, articles and resources exploring how advances in AI and machine learning can transform clinical development, accelerate the delivery of better treatments to patients, and help shape the future of medicine.

Causal AI
At the centre of this vision is causal AI. Clinical trials are not only about predicting outcomes, but they are also about making better decisions. What would happen if we changed the dose, the eligibility criteria, the comparator, the endpoint, the site mix, or the treatment strategy? Which patients are likely to benefit, and under what assumptions? What evidence is needed before a decision can be trusted?
The van der Schaar Lab has pioneered methods for causal effect inference and individualized treatment-effect estimation that help answer these questions. Our work enables researchers to better understand how treatments affect different patients, predict the consequences of interventions and generate evidence that can support more informed clinical and regulatory decisions.
Resources:
Causal Effect Inference (Research Pillar)
Key Publications:
- Nonparametric estimation of heterogeneous treatment effects: From theory to learning algorithms (AISTATS 2021)
- Using Machine Learning to Individualize Treatment Effect Estimation: Challenges and Opportunities (Clinical Pharmacology & Therapeutics)
- Causal machine learning for predicting treatment outcomes (Nature Medicine)
- Bayesian Inference of Individualized Treatment Effects using Multi-task Gaussian Processes (NeurIPS 2017)
- From Real-World Patient Data to Individualized Treatment Effects Using Machine Learning: Current and Future Methods to Address Underlying Challenges (Clinical Pharmacology & Therapeutics, 2020)
Recent Publications:
- Gradient-Based Causal Tree Ensembles: A Backbone Architecture for Heterogeneous Treatment Effects (ICML 2026)
- Identifiable Nonlinear Differentiable Causal Discovery via Independence and Adaptive Group Sparsity (ICML 2026)
- Overlap-weighted orthogonal meta-learner for treatment effect estimation over time (ICLR 2026)
- Treatment Effect Estimation for Optimal Decision-Making (NeurIPS 2025)
- AutoCATE: End-to-End, Automated Treatment Effect Estimation (ICML 2025)
- Active Feature Acquisition for Personalised Treatment Assignment (AISTATS 2025)
- Quantifying Aleatoric Uncertainty of the Treatment Effect: A Novel Orthogonal Learner (NeurIPS 2024)
- Causal Deep Learning
- ODE Discovery for Longitudinal Heterogeneous Treatment Effects Inference (ICLR 2024)
Digital Twins
This causal foundation naturally leads to digital twins. In our vision, digital twins are not generic simulators or decorative AI tools. They are dynamic, adaptive, causally grounded models of patients, diseases, populations, trials and healthcare systems. They allow us to ask: what could happen before we run the trial? Which assumptions are fragile? Which patients should be enrolled? How might disease trajectories evolve? How should we adapt if new evidence emerges?
The lab’s work on digital twins provides a route to stress-test trials before launch, support in-silico evidence generation, and support more informed clinical development decisions. Together, these technologies offer a path towards clinical trials that can continuously learn, adapt and improve over time.
Resources:
- Understanding Digital Twins (Sep 2025)
- Nature Biotech Q&A (Oct 2025) “Applications of digital twins in medicine”
- Inspiration Exchange 40
- Revolutionising Healthcare 38
Key Publications:
- Decision-Targeted Digital Twins (DT²) (ICML 2026)
- Continuously Updating Digital Twins using Large Language Models (ICML 2025)
- HDTwin – Automatically Learning Hybrid Digital Twins of Dynamical Systems (NeurIPS 2024)
- SyncTwin: Treatment Effect Estimation with Longitudinal Outcomes (NeurIPS 2021)
- HDTwinGen – Automatically Learning Hybrid Digital Twins of Dynamical Systems (NeurIPS 2024)
Synthetic Data
A third major pillar is synthetic data. Clinical development is constrained by limited access to high-quality, representative, privacy-preserving datasets. Over many years, the van der Schaar Lab has developed methods and tools for generating, evaluating and applying synthetic healthcare data to help address these challenges.
Synthetic data can support privacy-preserving data access, data augmentation, fairness, benchmarking and machine learning model development, while enabling researchers to work with realistic healthcare data in a safe and scalable way. This work includes open-source software such as SynthCity and more recent work on making synthetic data generation accessible to clinicians and healthcare researchers.
Resources:
Key Publications:
- Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes (ICML2024)
- Synthetic data in biomedicine via generative artificial intelligence (Nature Reviews Bioengineering, 2024)
- How Faithful is your Synthetic Data? Sample-level Metrics for Evaluating and Auditing Generative Models (ICML 2022)
- Time-series Generative Adversarial Networks (NeurIPS 2019)
- PATE-GAN: Generating Synthetic Data with Differential Privacy Guarantees (ICLR 2019)
- Synthcity: facilitating innovative use cases of synthetic data in different data modalities (NeurIPS 2023)
- Synthetic data for privacy-preserving clinical risk prediction (Scientific Reports, 2024)
- SynthCraft: An AI partner for synthetic data generation to support data access and augmentation in healthcare (PLOS Digital Health, 2026)
Adaptive Clinical Trials
The lab has also been at the forefront of adaptive clinical trials. Traditional trials often rely on decisions and assumptions made at the outset, with limited ability to respond to new information as it emerges. Adaptive trials offer a different paradigm: learning during the trial while preserving scientific validity, transparency and trust.
Over the years, our research has developed machine learning approaches that can support better patient allocation, cohort enrichment, treatment adaptation, and decision-making under uncertainty. The goal is not uncontrolled flexibility, but the ability to adapt when the evidence supports it, helping clinical trials become more efficient, informative and responsive.
Resources:
Key Publications:
- Revolutionizing Clinical Trials: A Manifesto for AI-Driven Transformation (2025)
- Adaptive Identification of Populations with Treatment Benefit in Clinical Trials: Machine Learning Challenges and Solutions (ICML 2023)
- Adaptive Experiment Design with Synthetic Controls (AISTATS 2024)
- Towards Regulatory-Confirmed Adaptive Clinical Trials: Machine Learning Opportunities and Solutions (AISTATS 2025)
- Machine learning for clinical trials in the era of COVID-19 (Stat Biopharm Res, 2020)
AI for Pharmacology
Another key area of research is AI for pharmacology and pharmacometrics. Clinical development depends on understanding dose, response, toxicity, pharmacokinetics, pharmacodynamics, disease progression and patient heterogeneity. The van der Schaar Lab’s work combines mechanistic modelling with modern machine learning to improve prediction of drug response, enable precision dosing, learn dynamical systems, and connect pharmacological theory with clinical data. By connecting pharmacological theory with clinical data, this work aims to support not only more efficient clinical trials, but also the scientific and biological questions at the heart of drug development.
Resources:
Key Publications:
- Data-Driven Discovery of Dynamical Systems in Pharmacology using Large Language Models (NeurIPS 2024)
- From Real-World Patient Data to Individualized Treatment Effects Using Machine Learning: Current and Future Methods to Address Underlying Challenges (Clinical Pharmacology & Therapeutics, 2020)
- Synthetic Model Combination: A new machine-learning method for pharmacometric model ensembling (CPT Pharmacometrics Syst Pharmacol, 2023)
- The potential and pitfalls of artificial intelligence in clinical pharmacology (CPT Pharmacometrics Syst Pharmacol, 2023)
- Bridging the Worlds of Pharmacometrics and Machine Learning (Clin Pharmacokinet, 2023)
- Integrating Expert ODEs into Neural ODEs: Pharmacology and Disease Progression (NeurIPS 2021)
Recent Publications:
Agentic AI for clinical trials
Most recently, the lab has been developing agentic AI for clinical trials. Bringing a new treatment to patients requires decisions across many areas, including biology, statistics, operations, regulation, safety, recruitment, and real-world evidence. As the volume and complexity of information continue to grow, there is an increasing interest in how AI systems can help researchers make sense of that evidence and support better decision-making.
In our vision, AI agents can become active reasoning partners: monitoring evidence, identifying fragile assumptions, stress-testing designs, integrating data streams, detecting drift, and helping teams ask better questions before failures occur. In our vision, these agents do not replace clinical, statistical or regulatory expertise. Rather they augment it, helping to create a new layer of clinical-development intelligence.
Resources:
White Paper: Clinical Trials as Continuously Learning Systems.
The opportunity now is to bring these advances together. Clinical trials are not transformed by any single technology, but by combining complementary approaches that address different parts of the development process. Causal AI helps answer the right questions, digital twins enable simulation and scenario testing, synthetic data expands access to information while protecting privacy, adaptive methodologies support better decision-making during trials and pharmacological AI connects mechanisms and treatment response, and AI agents helps researchers navigate increasingly complex evidence, assumptions and decisions across the development lifecycle.
Together, these technologies offer a new foundation for transforming clinical trials from isolated, rigid studies into continuously learning evidence systems. Before a trial begins, we can stress-test its design. During the trial, we can monitor assumptions and adapt responsibly. After the trial, we can feed the evidence back into digital twins, causal models, synthetic-data engines and agentic systems that improve the next trial.
The van der Schaar Lab has spent the past decade developing the scientific foundations for this transformation. The challenge now is translation: building validated, auditable, regulator-aligned tools that can be used in real clinical-development settings. If successful, this could make trials faster, safer, more predictive, more efficient, and ultimately better able to bring the right treatments to the right patients sooner.
This page brings together a selection of the publications, articles, software tools and resources that contribute to this vision of AI-enabled clinical development.









