Women’s health has too often been obscured by evidence built around the average patient. Machine learning gives us an opportunity to reveal overlooked risks, personalise care, and make the clinical decisions shaping women’s health more open to scrutiny, and ultimately improve them.
Women may be present in a dataset and still be poorly served by the model learned from it. A single risk score, diagnostic threshold or treatment pathway can perform well across a population while overlooking that age changes risk, that different information matters for different patients, or that apparently similar patients may experience different routes through care. Closing the historical data gap remains essential, but more data alone can still produce a more confident average.
Research across the van der Schaar Lab shows how carefully designed machine learning (ML) can do something more ambitious. It can uncover patterns that conventional models were not built to find; personalise screening and treatment decisions while representing uncertainty; and make the human decisions behind clinical pathways open to inspection. From cardiovascular risk prediction to breast-cancer screening and treatment, the goal is not merely higher average accuracy. It is to discover who is being overlooked, why, and what should change.
The story begins where population averages break down. Cardiovascular studies show that age, treatment delay and care setting can change the apparent relationship between sex and outcome. The ML research then tackles a harder problem: how can we systematically uncover clinically important differences without specifying in advance every subgroup or interaction that might matter?
Revealing what population averages conceal
Women are often reported to have worse outcomes after a heart attack. But that headline hides a more revealing truth: the women at greatest risk may not be the ones an average comparison leads us to expect. In “Sex Differences in Outcomes After STEMI: Effect Modification by Treatment Strategy and Age”, a large study of patients with ST-segment elevation myocardial infarction (STEMI), a life-threatening form of heart attack, found higher unadjusted 30-day mortality in women than men. After age and other clinical differences were considered, the remaining excess risk was concentrated among women under 60; older age groups did not show a statistically significant difference. The finding was therefore not simply that women fared worse: age changed the observed relationship between sex and outcome.
A related study, “Sex-Specific Treatment Effects After Primary Percutaneous Intervention”, examined patients receiving primary percutaneous coronary intervention (PCI)—the emergency procedure used to reopen a blocked coronary artery. Women had higher 30-day mortality and were more likely to have blood flow that remained below optimal after treatment. Among patients who reached hospital within 120 minutes, the mortality difference was no longer statistically significant, but the post-procedure blood-flow difference persisted. This suggests that delayed arrival may contribute to part of the pattern, while additional factors remain important after treatment begins.
These were targeted clinical analyses: once age or treatment delay is suspected, conventional statistical methods can test its importance. The harder problem is discovering consequential relationships that researchers have not already thought to test. In “Cardiovascular Disease Risk Prediction Using Automated Machine Learning”, AutoPrognosis analysed 473 variables from 423,604 UK Biobank participants without cardiovascular disease at baseline. It predicted five-year cardiovascular risk more accurately than the Framingham score and conventional proportional-hazards models.
The methodological advance was not simply the use of a more complex predictor. Rather than assuming one modelling approach in advance, AutoPrognosis selected and tuned entire clinical ML pipelines: how missing information should be handled, how variables should be processed, which predictive models should be combined and how their risk estimates should be calibrated. This distinction matters because different modelling choices can reveal different patterns and some ML models offer little improvement. The system searched those choices systematically and evaluated them against the clinical task.
The study also surfaced information that standard cardiovascular scores do not normally consider. Among women, hormone-replacement therapy and measured “ankle-spacing width” appeared among the more informative predictors. Among participants with diabetes, the variables that mattered most differed substantially from those in the population overall, with microalbuminuria emerging as particularly informative. These are predictive associations, not evidence that the variables cause cardiovascular disease. Their importance is that they open questions a short, fixed list of conventional risk factors might not have prompted researchers to ask. ML can help turn what was previously overlooked into something measurable and testable.
Personalising care beyond a single equation
Discovering an overlooked pattern is only the first step. An equally important challenge is allowing for the possibility that different patterns govern risk for different patients. Breast cancer provides some of the clearest examples of how ML can support that form of personalisation across a clinical pathway.
Screening must balance missed cancers against false positives, unnecessary procedures, cost and anxiety. A single guideline makes that trade-off for a population, even though the balance of risks and benefits may differ substantially between women.
Our ConfidentCare system addressed this limitation by identifying groups of women with similar characteristics and learning a separate screening policy for each group. Importantly, good average performance was not enough: every group-specific policy had to meet a predefined accuracy requirement with a specified level of confidence. In its evaluation, ConfidentCare was more cost-efficient than the clinical guidelines used for comparison and achieved a lower false-positive rate than a single decision-tree model. It therefore asked a different question from a conventional risk score: not simply whether a woman was at risk, but which screening approach was appropriate for women with her characteristics.
The next decision is which examinations need further expert attention. In our mammography research, the system classified a mammogram as negative only when sufficiently confident and sent uncertain or abnormal cases for specialist review. The retrospective proof-of-concept used data from more than 7,000 women recalled for assessment at six NHS breast-screening centres. While maintaining a negative predictive value of 0.99, it identified 34% of negative mammograms in a test setting with 15% cancer prevalence and 91% in a screening-like setting with 1% prevalence. These figures came from constructed test settings rather than prospective deployment, but they demonstrate the potential to reduce unnecessary radiologist workload without lowering the required confidence in a negative result.
Personalisation can continue after cancer is detected. In “Machine Learning to Guide the Use of Adjuvant Therapies for Breast Cancer”, Adjutorium used automated and interpretable ML to predict breast-cancer-specific and all-cause mortality and to estimate the potential survival benefit of therapies given after initial treatment. It was developed and validated using registry data from nearly one million women in the UK and United States. In those evaluations, Adjutorium was more accurate than PREDICT v2.1 and improved accuracy in groups known to be underserved by existing models. Because these were observational registry data, the estimates should not be interpreted as establishing the causal effect of treatment. Their value is in providing more individualised prognostic and treatment-benefit estimates to inform a clinical decision.
These direct applications reflect a broader ML principle illustrated by the lab’s Trees of Predictors method. In cardiac transplantation, the method did not fit one survival equation to every patient. It identified groups of patients and learned a different predictive model for each group, allowing the variables that mattered to change across patients and time horizons. This was not a women’s-health study and did not establish differences between women and men. Its relevance is methodological: it shows how ML can discover that risk is structured differently for different patients without requiring researchers to define every important group and interaction in advance.
Together, these projects describe a connected vision of personalised care: choosing how to screen, determining which cases require expert attention and estimating which treatments may offer the greatest predicted benefit. What becomes possible is not merely a more accurate population score, but a sequence of decisions adapted to the individual woman while remaining open to clinical judgement and scrutiny.
Making the clinical pathway visible
Women’s outcomes are shaped not only by biology, but also by what happens as they move through healthcare: when they seek help, whether their symptoms are recognised, which tests are ordered and what level of care they receive.
In “Sex Differences and Disparities in Cardiovascular Outcomes of COVID-19”, an international study involving our lab examined unvaccinated patients hospitalised with COVID-19. After balancing measured clinical characteristics, overall in-hospital mortality was similar in women and men. Beneath that average, however, women were less likely to be admitted to intensive care, and outcomes differed by care setting: among patients treated on general wards, women had greater risks of acute heart failure and in-hospital mortality, while these sex differences were not seen in intensive care. The study was observational and cannot establish that unequal access caused these outcomes; its entirely White, unvaccinated cohort also limits generalisation. Nevertheless, it shows how the reassuring claim that women were generally protected from severe COVID-19 could conceal both a vulnerable group of women and a clinically specific cardiovascular risk.
The COVID-19 study showed that disparities may be hidden inside the clinical pathway. INTERPOLE offers a way to investigate such pathways by reconstructing how evidence is accumulated and translated into decisions. In its Alzheimer’s disease application, it examined sequences of MRI ordering and diagnosis and identified apparent “belated diagnoses”: cases in which ordering an MRI was the most likely action under the learned policy, but no scan was ordered until a later visit, when it led to a near-certain diagnosis. This pattern was observed in 7.19% of female patients, compared with 5.97% of male patients, suggesting that women in this dataset were diagnosed less promptly and, in this specific sense, less well, than men. Although the study did not establish that this difference represented a statistically significant sex disparity, it shows how interpretable ML can make a possible inequality in diagnostic practice visible, measurable and open to scrutiny.
An opportunity we should not miss
None of this happens automatically. Models trained on incomplete or biased records can reproduce existing disparities. Predictive patterns require external validation; associations must not be presented as causes; and performance must be examined through calibration, uncertainty and clinically harmful errors – not average accuracy alone.
Yet designing ML to discover who is being overlooked is more than a fairness objective. It can produce better science: revealing differences in risk, treatment response and clinical pathways; generating new hypotheses; and exposing gaps in existing knowledge. More and better data on women remain essential. But without methods capable of representing meaningful differences, more data can still produce a more confident average.
The goal is not to replace the traditional “average patient” with an “average woman.” It is to ensure that no woman’s risk, symptoms or potential benefit from treatment disappears into the average, and to build better medicine for everyone.
Note on terminology: In this article, “women” and “men” reflect the language used in the original studies. Where sex was analysed, the available data generally classified participants in binary female/male categories and did not separately record gender identity. Accordingly, references to women in the reported statistical findings refer to participants classified as female in those datasets. We recognise that biological sex characteristics and gender identity are distinct, may not align and both include variation not captured by these categories.
References
Sex Differences in Outcomes After STEMI: Effect Modification by Treatment Strategy and Age. JAMA Internal Medicine.
Sex-Specific Treatment Effects After Primary Percutaneous Intervention: A Study on Coronary Blood Flow and Delay to Hospital Presentation. Journal of the American Heart Association.
ConfidentCare: A Clinical Decision Support System for Personalized Breast Cancer Screening. IEEE Transactions on Multimedia.
Improving Workflow Efficiency for Mammography Using Machine Learning. Journal of the American College of Radiology.
Machine learning to guide the use of adjuvant therapies for breast cancer. Nature Machine Intelligence.
Personalized survival predictions via Trees of Predictors: An application to cardiac transplantation. PLOS ONE.
Sex differences and disparities in cardiovascular outcomes of COVID-19. Cardiovascular Research.
Explaining by Imitating: Understanding Decisions by Interpretable Policy Learning. International Conference on Learning Representations.









