Does synthetic data help classifiers on a tiny dataset?

CodeReport (PDF)

University of Mannheim, Data Mining course. A team of six.

Problem

Health datasets are often small and imbalanced, and adding synthetic records is a popular fix whose effect is debated. We asked whether augmenting a small obesity-risk dataset with synthetic data improves classification into obesity levels, or only adds noise, and which models benefit.

The data came in three layers: 477 real survey records, 1,634 records generated from them with SMOTE, and 20,759 records generated by a deep learning model.

Approach

Results

Training dataEffect on the models
477 real records onlyWorst. Too small and too imbalanced; decision trees mostly predicted the majority class
Real plus SMOTE recordsLarge improvement for most models
Plus deep-learning recordsDropped again for most models; the generated data mostly added noise

What I’d do differently

How to run it

The code and instructions are in the project repository. The full write-up is linked as the report above.