van der Schaar Lab

Understanding Digital Twins

Digital twins are emerging as one of the most exciting frontiers in medicine and machine learning. Unlike traditional models that make static predictions, digital twins are dynamic, evolving representations of individual patients or one can think of them as “living models” that can generate multiple plausible futures. They hold the promise of transforming healthcare from reactive treatment to proactive planning. This blog is an introduction to digital twins: what they are, where they’re used, and how the van der Schaar Lab is pushing the field forward. Explore more in the links and videos here.

To learn more, see our feature in Nature Biotechnology’s Q&A on digital twins in medicine, which explores how these models are being developed and applied across medicine.

(Updated October 2025 to include our Nature Biotechnology Q&A on digital twins in medicine.)

Digital Twin schematic

Do you want to learn more about what digital twins are, and why they matter?

We’ve explored this topic in our public engagement series from two complementary angles:

Inspiration Exchange #40 – aimed at the machine learning community. Prof van der Schaar outlines the foundations of digital twins, how they differ from prediction and synthetic data, why they are generative and the challenges ahead. You can watch the video below:


Revolutionising Healthcare #38 – aimed at clinicians and healthcare professionals. Here, Prof van der Schaar introduces digital twins as an empowering new technology for medicine, highlighting their potential for personalised care, early diagnosis, operational efficiency and innovation. Watch the talk below:

Together, these sessions show how digital twins are both a technical frontier in machine learning and a practical opportunity for transforming healthcare. You can join these ongoing engagement series to take part in the discussions and explore the future of AI in healthcare with us.


Do you want to know more about the technical details?

For those interested in the research foundations, our lab has developed several approaches that advance digital twin methodology:

HDTwin (Hybrid Digital Twin)

Traditional digital twins either rely too heavily on mechanistic models (which struggle with flexibility) or purely data-driven neural networks (which can fail in data-scarce settings). HDTwin combines the strengths of both through a hybrid architecture: mechanistic components embed domain knowledge, while neural components capture complex, data-driven patterns.

What makes HDTwin distinctive is its evolvability. The architecture is modular, so it can be extended and adapted as new information or requirements emerge. We also developed HDTwinGen, an evolutionary algorithm that uses large language models (LLMs) to automatically propose, evaluate, and optimise hybrid models. Instead of relying on experts to handcraft architectures, HDTwinGen enables the discovery of new, effective hybrid twins by generating and refining model specifications iteratively.

The result is a digital twin framework that is generalizable, sample-efficient, and adaptive, advancing the field well beyond static or rigid designs.

CALM-DT (Context-Adaptive Language Model-based Digital Twin)

A major limitation of most digital twins is that they require a fixed, well-defined modelling environment. Adding new variables (such as a novel biomarker or therapy) typically requires re-designs and retraining, making them brittle in practice. CALM-DT reframes digital twinning as an in-context learning problem.

By using LLMs as adaptive engines, CALM-DT can integrate new variables, state-action spaces, and sources of knowledge at inference time, without retraining or parameter updates. To support this, we designed fine-tuned encoders that retrieve relevant samples and contextualise them for the LLM, enabling the twin to update seamlessly in real-world settings.

In our empirical work, CALM-DT not only performs competitively with existing digital twin approaches but uniquely demonstrates the ability to adapt dynamically as its modelling environment changes. This is what allows CALM-DT to function as a truly living model that evolves alongside the system it represents.

SyncTwin Treatment Effect Estimation with
Longitudinal Outcomes

Most medical observational studies aim to estimate causal treatment effects using electronic health records (EHR), where both a patient’s covariates and outcomes are observed longitudinally. However, many existing approaches adjust only for covariates while neglecting the temporal structure of outcomes, which can be critical in real-world healthcare data.

SyncTwin (NeurIPS 2021) bridges this gap by learning a patient-specific, time-constant representation from pre-treatment observations. Using this representation, it constructs a synthetic twin that closely matches the target patient in its pre-treatment trajectory. This synthetic twin enables counterfactual prediction – estimating how the patient would have responded under alternative treatments – in a way that reflects the underlying temporal dynamics.

The reliability of each estimate can be verified by comparing observed and synthetic pre-treatment outcomes, while interpretability is achieved by identifying which individuals most influenced the synthetic twin’s construction.

In real-world experiments, SyncTwin successfully reproduced the findings of a randomized controlled trial using only observational data, showing that machine learning can generate accurate, interpretable digital twins that serve as synthetic control arms, a major step toward trustworthy, data-driven clinical research.

The next frontier. Digital Twin Agents

HDTwin and CALM-DT push digital twins toward adaptability and evolvability. The next step is Digital Twin Agents : biologically grounded, high-fidelity twins that bring together a patient’s molecular profile, clinical history, and therapeutic options in a unified model. Unlike current twins that primarily react to clinical data, agents will be able to anticipate, reason, and intervene, serving as active decision-support systems. Our aim is to build genomics-informed, high-actionability twins that integrate biology, clinical practice, and AI reasoning to transform personalised care. This vision connects closely with our broader genomics work, including a recent MLGenX 2025 paper co-authored with pharma collaborators on interpretable DNA sequence analysis. Advances like these provide the foundations for building biologically grounded, genomics-informed digital twins.


Do you want to see how digital twins could transform clinical trials?

Our latest research sets out a manifesto for reimagining clinical trials with AI. By combining casual inference and digital twins, trials can become faster, more personalised and more inclusive. This paper outlines a structured path forward, from robust validation and best practices for data sharing and standardisation, to fostering collaboration across academia, industry and regulators, all while ensuring success and regulatory compliance.

Mihaela van der Schaar

Mihaela van der Schaar is the John Humphrey Plummer Professor of Machine Learning, Artificial Intelligence and Medicine at the University of Cambridge and a Fellow at The Alan Turing Institute in London.

Mihaela has received numerous awards, including the Oon Prize on Preventative Medicine from the University of Cambridge (2018), a National Science Foundation CAREER Award (2004), 3 IBM Faculty Awards, the IBM Exploratory Stream Analytics Innovation Award, the Philips Make a Difference Award and several best paper awards, including the IEEE Darlington Award.

In 2019, she was identified by National Endowment for Science, Technology and the Arts as the most-cited female AI researcher in the UK. She was also elected as a 2019 “Star in Computer Networking and Communications” by N²Women. Her research expertise span signal and image processing, communication networks, network science, multimedia, game theory, distributed systems, machine learning and AI.

Mihaela’s research focus is on machine learning, AI and operations research for healthcare and medicine.

Marika Niihori

Marika is our communications manager since joining in 2025. Marika is a trained physicist with a PhD in NanoPhotonics from the University of Cambridge.

Alongside her scientific background, she has extensive experience in science communication through content creation, outreach, and public engagement. She has also gained industry experience in biotech, further broadening her perspective on how research translates into real-world applications.

Marika works to share the group’s cutting-edge AI and machine learning research with both scientific and wider audiences, making complex ideas clear, engaging, and impactful.