Synthetic data is a powerful tool to improve real data – whether it’s augmenting small sample data (a major challenge in healthcare) or debiasing and ensuring fairness of data. To ensure high quality generation we both need novel metrics to quantify quality, but also if we want to ensure the synthetic data is truly reflective of real data (with all it’s issues like like noise etc), we can guide synthetic data generation with ideas from data characterisation.
For more of our work on Synthetic data and other uses see our Synthetic Data Research Pillar.
Read more:
Augmentation: CLLM: https://arxiv.org/abs/2312.12112
Fairness: DECAF: https://arxiv.org/abs/2110.12884
Metrics: https://arxiv.org/abs/2110.12884
Hardness-guided synthetic data generation: https://arxiv.org/abs/2310.16981









