Does synthetic data help classifiers on a tiny dataset?
University of Mannheim, Data Mining course. A team of six.
Problem
Health datasets are often small and imbalanced, and adding synthetic records is a popular fix whose effect is debated. We asked whether augmenting a small obesity-risk dataset with synthetic data improves classification into obesity levels, or only adds noise, and which models benefit.
The data came in three layers: 477 real survey records, 1,634 records generated from them with SMOTE, and 20,759 records generated by a deep learning model.
Approach
- Two experiments with one fixed test set each. Holding the test set constant while changing only the training data is the only way to isolate what synthetic data does. The initial experiment trains on the real and SMOTE records with and without the deep-learning data; the extra experiment adds the real-only training set. The test set in the extra experiment comes from real records only.
- Leakage control. The deep-learning data was generated before any train/test split, so it contained records derived from test records. We removed near-duplicates of test records from it using Gower distance, with a threshold chosen so that every test record had at least one match without shrinking the small datasets further.
- Models. Nine classifiers (decision tree, random forest, gradient boosting, XGBoost, LightGBM, k-NN, naive Bayes, SVM and an MLP) plus a majority-class baseline, each tuned with 25 iterations of randomized search. Cross-validation was leave-one-out for the 477 real records and stratified 5-fold otherwise, scored with macro F1 so that rare classes count.
Results
| Training data | Effect on the models |
|---|---|
| 477 real records only | Worst. Too small and too imbalanced; decision trees mostly predicted the majority class |
| Real plus SMOTE records | Large improvement for most models |
| Plus deep-learning records | Dropped again for most models; the generated data mostly added noise |
- XGBoost and LightGBM gained 5 to 10% F1 from the synthetic data in the extra experiment, and the MLP also improved with more synthetic data. Decision trees and random forests got worse.
- Simple SMOTE data helped more than the generative model’s data for most models.
- Synthetic data can improve performance, but only with a model that tolerates noise and a generator that preserves the real patterns.
What I’d do differently
- The real dataset is tiny, and some leakage risk remains despite the Gower filtering. I would validate on real data only and collect more of it.
- Filtering the generated data for quality before training, and combining several generation methods, are the next experiments the report proposes.
- Feature importances differed across both models and datasets, and we could only cover a few of the combinations. A systematic comparison would explain the model-specific effects.
How to run it
The code and instructions are in the project repository. The full write-up is linked as the report above.