A systematic experimental study demonstrating how flexible ML models can produce misleadingly high training accuracy on data with zero signal. Across neural networks, logistic regression, and random forests, training accuracy climbs as high as 64–100% on pure noise — while held-out accuracy stays firmly at 50% (chance).
A model reporting 64% accuracy on a binary task may sound meaningful, but if that number comes only from training data, it can be entirely explained by overfitting to noise. Always evaluate on a held-out set.
| Experiment | Script | What it tests |
|---|---|---|
| Core overfitting | pytorch_experiment.py |
2-layer NN across 9 input sizes (2–50 features) on pure noise |
| Baselines | baseline_experiment.py |
Logistic regression & random forest on the same data |
| Model complexity | complexity_experiment.py |
Hidden units × depth sweep (h10d1 → h500d2) |
| Sample size | sample_size_experiment.py |
Training set sizes from 50 to 1,500 |
| Regularisation | regularisation_experiment.py |
Dropout, weight decay, and combinations |
| Signal injection | signal_experiment.py |
0–25 informative features to show what real learning looks like |
Each experiment runs 500 bootstrap seeds for statistical robustness.
Training accuracy rises with feature count while test accuracy stays at chance — the textbook signature of overfitting.
Random forest achieves 100% training accuracy yet ~50% test accuracy — the clearest demonstration that training accuracy can be meaningless.
More data does not help when signal is zero.
Larger models overfit more but still cannot generalise.
Regularisation shrinks the train–test gap but cannot create signal.
With even 1 real feature out of 25, test accuracy jumps to 83% and train/test metrics converge — this is what genuine learning looks like.
See interpretation.md for the full analysis and statistical tests.
- Python 3.10+ with PyTorch (CUDA recommended)
- R with
tidyverse - scikit-learn, numpy, pandas
bash run_sh.shThis generates the data, runs all 6 experiments (500 seeds each), and produces CSV result files.
Rscript r_visualization.r├── generate_data.py # Creates pure-noise dataset (100 features, coin-flip labels)
├── common.py # Shared model, training loop, and utilities
├── pytorch_experiment.py # Core NN overfitting experiment
├── baseline_experiment.py # Logistic regression & random forest baselines
├── complexity_experiment.py # Model capacity sweep
├── sample_size_experiment.py # Training set size sweep
├── regularisation_experiment.py # Dropout & weight decay sweep
├── signal_experiment.py # Informative features injection
├── r_visualization.r # R script producing all violin plots
├── run_sh.sh # Master script to run everything
├── interpretation.md # Full analysis and interpretation
└── data/ # Generated .npy data files
- 100 random features from a standard normal distribution
- Labels: coin-flip random (50/50 binary) — zero signal
- Splits: 1,800 training pool, 200 validation, 500 test
- Each seed draws a bootstrap sample of 1,000 train and 100 test rows
MIT






