van der Schaar Lab

Data quality – Data characterisation

Data characterisation methods aim to assess samples in a dataset in a principled manner — to understand which samples in a dataset are easy and learnable for a machine learning model, which are ambiguous and which samples are hard, mislabeled and unlearnable. This can be used both to audit a dataset from a trustworthy ML perspective by pruning problematic samples, but also to select samples to improve training efficiency. The below works highlight data characterisation methods applicable to any data modality and both for classification (Data-IQ) and regression (TRIAGE) datasets. H-CAT then benchmarks different methods for hardness characterisation and resulted in our open-source hardness characterisation package Datagnosis.

Read more:

Software:

Watch our Inspiration Exchange session on TRIAGE and Datagnosis: