Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

30 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

False Discovery in Machine Learning: Training Accuracy on Pure Noise

A systematic experimental study demonstrating how flexible ML models can produce misleadingly high training accuracy on data with zero signal. Across neural networks, logistic regression, and random forests, training accuracy climbs as high as 64–100% on pure noise — while held-out accuracy stays firmly at 50% (chance).

Key Finding

A model reporting 64% accuracy on a binary task may sound meaningful, but if that number comes only from training data, it can be entirely explained by overfitting to noise. Always evaluate on a held-out set.

Experiments

Experiment Script What it tests
Core overfitting pytorch_experiment.py 2-layer NN across 9 input sizes (2–50 features) on pure noise
Baselines baseline_experiment.py Logistic regression & random forest on the same data
Model complexity complexity_experiment.py Hidden units × depth sweep (h10d1 → h500d2)
Sample size sample_size_experiment.py Training set sizes from 50 to 1,500
Regularisation regularisation_experiment.py Dropout, weight decay, and combinations
Signal injection signal_experiment.py 0–25 informative features to show what real learning looks like

Each experiment runs 500 bootstrap seeds for statistical robustness.

Results

Neural network on pure noise

Training accuracy rises with feature count while test accuracy stays at chance — the textbook signature of overfitting.

Logistic regression on pure noise Random forest on pure noise

Random forest achieves 100% training accuracy yet ~50% test accuracy — the clearest demonstration that training accuracy can be meaningless.

Sample size sweep

More data does not help when signal is zero.

Complexity sweep

Larger models overfit more but still cannot generalise.

Regularisation comparison

Regularisation shrinks the train–test gap but cannot create signal.

Signal injection

With even 1 real feature out of 25, test accuracy jumps to 83% and train/test metrics converge — this is what genuine learning looks like.

See interpretation.md for the full analysis and statistical tests.

Quick Start

Prerequisites

  • Python 3.10+ with PyTorch (CUDA recommended)
  • R with tidyverse
  • scikit-learn, numpy, pandas

Run everything

bash run_sh.sh

This generates the data, runs all 6 experiments (500 seeds each), and produces CSV result files.

Generate plots

Rscript r_visualization.r

Project Structure

├── generate_data.py              # Creates pure-noise dataset (100 features, coin-flip labels)
├── common.py                     # Shared model, training loop, and utilities
├── pytorch_experiment.py         # Core NN overfitting experiment
├── baseline_experiment.py        # Logistic regression & random forest baselines
├── complexity_experiment.py      # Model capacity sweep
├── sample_size_experiment.py     # Training set size sweep
├── regularisation_experiment.py  # Dropout & weight decay sweep
├── signal_experiment.py          # Informative features injection
├── r_visualization.r             # R script producing all violin plots
├── run_sh.sh                     # Master script to run everything
├── interpretation.md             # Full analysis and interpretation
└── data/                         # Generated .npy data files

Data Setup

  • 100 random features from a standard normal distribution
  • Labels: coin-flip random (50/50 binary) — zero signal
  • Splits: 1,800 training pool, 200 validation, 500 test
  • Each seed draws a bootstrap sample of 1,000 train and 100 test rows

License

MIT

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages