Skip to content

About

N2O Digital Twin for wastewater treatment — predicts N2O emissions up to 60 minutes ahead (6-step multi-horizon) using 10-minute interval time-series data. Built with XGBoost + LSTM ensemble model, FastAPI backend, and a static frontend dashboard for operational decision-making (e.g., DO control scenarios).

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

7 Commits

Folders and files

Repository files navigation

AI Festival: N2O Digital Twin

This project predicts N2O emissions (up to 60 minutes ahead) using time-series data sampled every 10 minutes from a wastewater treatment plant, and is intended to support operational decision-making (e.g., DO control scenarios).

At a glance

  • Goal: 6-step multi-horizon forecast of N2O-N [kg N/h] (10, 20, 30, 40, 50, 60 minutes ahead)
  • Input: past 36 steps (6 hours) sequence
  • Core model: XGBoost + LSTM ensemble (src/ensemble_xgb_v2_lstm_v6.py)
  • Service: FastAPI backend + static frontend dashboard (webapp/)

Directory structure

ai_festival/
├── data/
│   ├── raw/                # Raw data
│   ├── processed/          # Preprocessing outputs
│   └── reports/            # Reports/visualizations
├── src/                    # Production training/evaluation code
├── experiments/
│   ├── src/                # Experiment scripts
│   └── results/            # Experiment outputs
├── results/                # Operational/final artifacts
├── scripts/                # Batch pipelines
├── webapp/
│   ├── backend/            # FastAPI
│   └── frontend/           # HTML/JS frontend
└── docs/                   # Project documentation

Quick start

1) Prepare the environment

cd /Volumes/a3122a1/ai_festival
python3 -m venv .venv
source .venv/bin/activate
pip install -r webapp/backend/requirements.txt

2) Run the preprocessing pipeline

python scripts/n2o_control_pipeline.py

Main outputs:

  • data/processed/N2O_10min_phase_preserved.csv
  • data/reports/outlier_flag_report.csv

3) Train / evaluate models

python src/ensemble_xgb_v2_lstm_v6.py

Example artifacts:

  • results/ensemble_xgb_v2_lstm_v6/test_metrics.csv
  • results/ensemble_xgb_v2_lstm_v6/xgb_models/*.pkl
  • results/lstm_v6_phase/best_model.pth

Running the webapp

Backend

cd webapp/backend
source ../../.venv/bin/activate
OMP_NUM_THREADS=1 KMP_DUPLICATE_LIB_OK=TRUE MKL_NUM_THREADS=1 \
  uvicorn main:app --reload --port 8000

Frontend

cd webapp/frontend
python3 -m http.server 3000

Open in browser:

  • http://localhost:3000

Main API endpoints:

  • GET /health
  • POST /api/upload
  • GET /api/metrics
  • GET /api/backtest

Experiment code guide

experiments/src contains separate experiment runners for XGBoost/LightGBM/CatBoost/LSTM/GRU/TimesNet/Hybrid/Ensemble, etc.

  • XGBoost family: experiments/src/experiments/xgb/
  • LSTM family: experiments/src/lstm/
  • GRU family: experiments/src/gru/
  • Hybrid: experiments/src/experiments/hybrid/
  • Ensemble: experiments/src/experiments/ensemble/

For model comparisons and detailed parameters, see docs/PROJECT_OVERVIEW.md.

Git policy for CSV files

The current .gitignore policy is as follows:

  • By default, all CSV files are ignored.
  • However, to preserve results for reproducibility, CSV files in the following paths are tracked:
    • results/**/*.csv
    • experiments/results/**/*.csv

In other words, large raw CSV data are excluded from version control, while result CSVs needed for reproducibility are kept in the repository.

Reference documents

  • docs/PROJECT_OVERVIEW.md: Detailed descriptions and performance by model
  • docs/CLAUDE_CODE_PROMPT.md: Development prompts/rules

About

N2O Digital Twin for wastewater treatment — predicts N2O emissions up to 60 minutes ahead (6-step multi-horizon) using 10-minute interval time-series data. Built with XGBoost + LSTM ensemble model, FastAPI backend, and a static frontend dashboard for operational decision-making (e.g., DO control scenarios).

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages