Grade: 85% · University of Plymouth
A two-part coursework covering supervised regression and metaheuristic optimisation, implemented end-to-end in Python.
This notebook contains two independent problems:
Part 1 — Machine Learning: Predict nitrite concentration in ocean water samples using three regression models, trained on a real-world nutrient chemistry dataset (e1_nutrients.csv).
Part 2 — Optimisation: Implement and compare two metaheuristic algorithms - a Hill Climber and an Evolutionary Algorithm - tasked with minimising the McCormick benchmark function.
Nitrite (NO₂⁻) is a chemically unstable intermediate in the nitrogen cycle, appearing briefly as ammonia oxidises into nitrate. Because it's reactive and transient, predicting its concentration from co-occurring chemicals is non-trivial - and exactly the kind of task where ML shines over hand-crafted rules.
Target variable: NITRITE
Features: NITRATE+NITRITE, AMMONIA, SILICATE, PHOSPHATE, DEPTH
The dataset contains 2,496 samples across 6 numeric columns with no missing values. The following preprocessing steps were applied:
- Log1p normalisation - applied before splitting to handle the heavily right-skewed distributions dominating each feature
- 80/20 train/test split - with all subsequent transformations fit on training data only
- IQR-based outlier capping - outliers clipped to lower/upper fence values rather than discarded, preserving rare high-nitrite samples
- MinMax scaling - all features scaled to [0, 1] to prevent any single feature dominating by magnitude
Feature correlation analysis (Seaborn heatmap) confirmed NITRATE+NITRITE and PHOSPHATE share the strongest relationship (r = 0.80), while AMMONIA shows a negative correlation with NITRATE+NITRITE — chemically interpretable as ammonia not yet converted.
| Model | Notes |
|---|---|
| Linear Regression | Assumes linear input–output relationships; serves as interpretable baseline |
| Random Forest | Ensemble of decision trees; naturally handles non-linear patterns and feature interactions |
| MLP (Neural Network) | Multi-layer perceptron; learns non-linear mappings via backpropagation |
A DummyRegressor (predicting the training mean) was used as a sanity-check baseline.
Hyperparameter tuning was performed via RandomizedSearchCV with 5-fold cross-validation (60–100 iterations) for the Random Forest and MLP. Linear Regression takes no hyperparameters.
| Model | Untuned MSE | Tuned MSE | vs. Baseline |
|---|---|---|---|
| Dummy (baseline) | 0.0370 | - | - |
| Linear Regression | 0.0303 | 0.0303 | −19% |
| MLP | 0.0259 | 0.0228 | −43% |
| Random Forest | 0.0145 | 0.0134 | −65% |
All three models outperform the dummy baseline, confirming genuine learning. The Random Forest achieved the best accuracy and the most consistent cross-validation scores (smallest IQR). The MLP was a close second, though residual plots reveal it averages predictions and struggles with rare high-nitrite spikes. Linear Regression confirmed the dataset's non-linearity.
A shared weakness across all models: underprediction of high nitrite values - a direct consequence of the dataset's imbalanced distribution.
The McCormick function is a standard optimisation benchmark with a known analytic minimum:
The canonical minimum is f ≈ −1.913 at (−0.547, −1.547). Within the extended search range of [−5, 5] used here, a deeper global minimum exists at f ≈ −5.055, adding an extra layer of difficulty.
Both algorithms use additive Gaussian mutation as their exploration operator, with solutions clipped to [−5, 5] bounds.
Hill Climber
- Maintains a single solution at a time
- Applies mutation, then keeps whichever candidate (parent or child) has lower fitness - greedy selection
- Fast per iteration, but highly susceptible to local minima
Evolutionary Algorithm
- Maintains a population of 10 solutions
- Each generation: 50 children are generated by mutating randomly-selected parents, then the best 10 from the combined pool are kept - elitism selection
- Broader exploration through population diversity
Both algorithms were evaluated over 30 independent runs across sigma values of 0.1, 0.3, 0.5, 0.7, and 1.0.
| Algorithm | Best σ | Global min found | Notes |
|---|---|---|---|
| Hill Climber | 0.5 | 6/30 runs | Often trapped at −1.91 or local minima |
| Evolutionary Algorithm | 0.7 | 30/30 runs | Std ≈ 0, perfectly consistent |
At σ = 0.7–1.0, the EA achieved a standard deviation of ~0 across 30 runs - finding the global minimum every time. The HC remained inconsistent regardless of sigma: too small and it gets trapped; too large and solutions become noisy.
The EA's advantage comes from simultaneous multi-point exploration: diversity in the population makes it resilient to local minima that reliably catch single-solution approaches.
Python 3.x
├── pandas — data loading, DataFrame handling
├── numpy — numerical operations, array processing
├── matplotlib — all plotting (histograms, boxplots, contour maps, residuals)
├── seaborn — correlation heatmap
└── scikit-learn — preprocessing, model training, hyperparameter search, evaluation
.
├── comp2002_coursework.ipynb # Full notebook with code, analysis and plots
├── e1_nutrients.csv # Ocean nutrient dataset (Part 1, provided by University, not included in GitHub)
└── README.md
- Real-world datasets require careful, stepwise preprocessing - choices like capping vs. removing outliers meaningfully affect downstream model behaviour.
- Random Forest is a strong default for tabular regression with non-linear structure; it naturally captures feature interactions without requiring architectural decisions.
- Population-based search (EA) fundamentally outperforms single-solution search (HC) for multi-modal landscapes - the McCormick results make this vivid and concrete.
- Benchmark problems are genuinely useful: having a known ground truth makes algorithm comparison rigorous rather than speculative.