Skip to content
sphsnyPublic

About

ML and EA in Jupyter Notebook

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

COMP2002 Coursework — Machine Learning & Optimisation

Grade: 85% · University of Plymouth
A two-part coursework covering supervised regression and metaheuristic optimisation, implemented end-to-end in Python.


Overview

This notebook contains two independent problems:

Part 1 — Machine Learning: Predict nitrite concentration in ocean water samples using three regression models, trained on a real-world nutrient chemistry dataset (e1_nutrients.csv).

Part 2 — Optimisation: Implement and compare two metaheuristic algorithms - a Hill Climber and an Evolutionary Algorithm - tasked with minimising the McCormick benchmark function.


Part 1: Nitrite Regression

The Problem

Nitrite (NO₂⁻) is a chemically unstable intermediate in the nitrogen cycle, appearing briefly as ammonia oxidises into nitrate. Because it's reactive and transient, predicting its concentration from co-occurring chemicals is non-trivial - and exactly the kind of task where ML shines over hand-crafted rules.

Target variable: NITRITE
Features: NITRATE+NITRITE, AMMONIA, SILICATE, PHOSPHATE, DEPTH

Data Pipeline

The dataset contains 2,496 samples across 6 numeric columns with no missing values. The following preprocessing steps were applied:

  • Log1p normalisation - applied before splitting to handle the heavily right-skewed distributions dominating each feature
  • 80/20 train/test split - with all subsequent transformations fit on training data only
  • IQR-based outlier capping - outliers clipped to lower/upper fence values rather than discarded, preserving rare high-nitrite samples
  • MinMax scaling - all features scaled to [0, 1] to prevent any single feature dominating by magnitude

Feature correlation analysis (Seaborn heatmap) confirmed NITRATE+NITRITE and PHOSPHATE share the strongest relationship (r = 0.80), while AMMONIA shows a negative correlation with NITRATE+NITRITE — chemically interpretable as ammonia not yet converted.

Models

Model Notes
Linear Regression Assumes linear input–output relationships; serves as interpretable baseline
Random Forest Ensemble of decision trees; naturally handles non-linear patterns and feature interactions
MLP (Neural Network) Multi-layer perceptron; learns non-linear mappings via backpropagation

A DummyRegressor (predicting the training mean) was used as a sanity-check baseline.

Hyperparameter tuning was performed via RandomizedSearchCV with 5-fold cross-validation (60–100 iterations) for the Random Forest and MLP. Linear Regression takes no hyperparameters.

Results

Model Untuned MSE Tuned MSE vs. Baseline
Dummy (baseline) 0.0370 - -
Linear Regression 0.0303 0.0303 −19%
MLP 0.0259 0.0228 −43%
Random Forest 0.0145 0.0134 −65%

All three models outperform the dummy baseline, confirming genuine learning. The Random Forest achieved the best accuracy and the most consistent cross-validation scores (smallest IQR). The MLP was a close second, though residual plots reveal it averages predictions and struggles with rare high-nitrite spikes. Linear Regression confirmed the dataset's non-linearity.

A shared weakness across all models: underprediction of high nitrite values - a direct consequence of the dataset's imbalanced distribution.


Part 2: McCormick Function Optimisation

The Problem

The McCormick function is a standard optimisation benchmark with a known analytic minimum:

$$f(x, y) = \sin(x + y) + (x - y)^2 - 1.5x + 2.5y + 1$$

The canonical minimum is f ≈ −1.913 at (−0.547, −1.547). Within the extended search range of [−5, 5] used here, a deeper global minimum exists at f ≈ −5.055, adding an extra layer of difficulty.

Algorithms

Both algorithms use additive Gaussian mutation as their exploration operator, with solutions clipped to [−5, 5] bounds.

Hill Climber

  • Maintains a single solution at a time
  • Applies mutation, then keeps whichever candidate (parent or child) has lower fitness - greedy selection
  • Fast per iteration, but highly susceptible to local minima

Evolutionary Algorithm

  • Maintains a population of 10 solutions
  • Each generation: 50 children are generated by mutating randomly-selected parents, then the best 10 from the combined pool are kept - elitism selection
  • Broader exploration through population diversity

Results (30 runs each)

Both algorithms were evaluated over 30 independent runs across sigma values of 0.1, 0.3, 0.5, 0.7, and 1.0.

Algorithm Best σ Global min found Notes
Hill Climber 0.5 6/30 runs Often trapped at −1.91 or local minima
Evolutionary Algorithm 0.7 30/30 runs Std ≈ 0, perfectly consistent

At σ = 0.7–1.0, the EA achieved a standard deviation of ~0 across 30 runs - finding the global minimum every time. The HC remained inconsistent regardless of sigma: too small and it gets trapped; too large and solutions become noisy.

The EA's advantage comes from simultaneous multi-point exploration: diversity in the population makes it resilient to local minima that reliably catch single-solution approaches.


Tech Stack

Python 3.x
├── pandas          — data loading, DataFrame handling
├── numpy           — numerical operations, array processing
├── matplotlib      — all plotting (histograms, boxplots, contour maps, residuals)
├── seaborn         — correlation heatmap
└── scikit-learn    — preprocessing, model training, hyperparameter search, evaluation

Repository Structure

.
├── comp2002_coursework.ipynb   # Full notebook with code, analysis and plots
├── e1_nutrients.csv            # Ocean nutrient dataset (Part 1, provided by University, not included in GitHub)
└── README.md

Key Takeaways

  • Real-world datasets require careful, stepwise preprocessing - choices like capping vs. removing outliers meaningfully affect downstream model behaviour.
  • Random Forest is a strong default for tabular regression with non-linear structure; it naturally captures feature interactions without requiring architectural decisions.
  • Population-based search (EA) fundamentally outperforms single-solution search (HC) for multi-modal landscapes - the McCormick results make this vivid and concrete.
  • Benchmark problems are genuinely useful: having a known ground truth makes algorithm comparison rigorous rather than speculative.

About

ML and EA in Jupyter Notebook

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages