This project predicts N2O emissions (up to 60 minutes ahead) using time-series data sampled every 10 minutes from a wastewater treatment plant, and is intended to support operational decision-making (e.g., DO control scenarios).
- Goal: 6-step multi-horizon forecast of N2O-N [kg N/h] (10, 20, 30, 40, 50, 60 minutes ahead)
- Input: past 36 steps (6 hours) sequence
- Core model: XGBoost + LSTM ensemble (
src/ensemble_xgb_v2_lstm_v6.py) - Service: FastAPI backend + static frontend dashboard (
webapp/)
ai_festival/
├── data/
│ ├── raw/ # Raw data
│ ├── processed/ # Preprocessing outputs
│ └── reports/ # Reports/visualizations
├── src/ # Production training/evaluation code
├── experiments/
│ ├── src/ # Experiment scripts
│ └── results/ # Experiment outputs
├── results/ # Operational/final artifacts
├── scripts/ # Batch pipelines
├── webapp/
│ ├── backend/ # FastAPI
│ └── frontend/ # HTML/JS frontend
└── docs/ # Project documentation
cd /Volumes/a3122a1/ai_festival
python3 -m venv .venv
source .venv/bin/activate
pip install -r webapp/backend/requirements.txtpython scripts/n2o_control_pipeline.pyMain outputs:
data/processed/N2O_10min_phase_preserved.csvdata/reports/outlier_flag_report.csv
python src/ensemble_xgb_v2_lstm_v6.pyExample artifacts:
results/ensemble_xgb_v2_lstm_v6/test_metrics.csvresults/ensemble_xgb_v2_lstm_v6/xgb_models/*.pklresults/lstm_v6_phase/best_model.pth
cd webapp/backend
source ../../.venv/bin/activate
OMP_NUM_THREADS=1 KMP_DUPLICATE_LIB_OK=TRUE MKL_NUM_THREADS=1 \
uvicorn main:app --reload --port 8000cd webapp/frontend
python3 -m http.server 3000Open in browser:
http://localhost:3000
Main API endpoints:
GET /healthPOST /api/uploadGET /api/metricsGET /api/backtest
experiments/src contains separate experiment runners for XGBoost/LightGBM/CatBoost/LSTM/GRU/TimesNet/Hybrid/Ensemble, etc.
- XGBoost family:
experiments/src/experiments/xgb/ - LSTM family:
experiments/src/lstm/ - GRU family:
experiments/src/gru/ - Hybrid:
experiments/src/experiments/hybrid/ - Ensemble:
experiments/src/experiments/ensemble/
For model comparisons and detailed parameters, see docs/PROJECT_OVERVIEW.md.
The current .gitignore policy is as follows:
- By default, all CSV files are ignored.
- However, to preserve results for reproducibility, CSV files in the following paths are tracked:
results/**/*.csvexperiments/results/**/*.csv
In other words, large raw CSV data are excluded from version control, while result CSVs needed for reproducibility are kept in the repository.
docs/PROJECT_OVERVIEW.md: Detailed descriptions and performance by modeldocs/CLAUDE_CODE_PROMPT.md: Development prompts/rules