Productizing 2D DEM earthquake-rupture analyses as interactive dashboards and a companion web site.
- Companion site: https://harvardrc.github.io/eps-ground-rupture/
- Dashboards: https://public.tableau.com/app/profile/michael.bouzinier (free to view, no login)
- The paper: Chiama et al. (2025), Earthquake Spectra 41(5), 3977–4014, DOI 10.1177/87552930251346434 — not open access. Six of its figures are shown on the site from the authors' accepted manuscript under the publisher's sharing terms (see Licensing below); nothing from the typeset version is reproduced.
- Related work (the 3D models; not a source for these dashboards): Chiama, Plesch & Shaw (2025), Seismological Research Letters 96(6), 3473–3489, DOI 10.1785/0220250173
The legacy material — two Jupyter notebooks and the accompanying paper —
lives under legacy/ as local reference artifacts (gitignored). This repo
turns that notebook-driven workflow into:
- A Python data pipeline (modules + CLI; no notebooks) that ingests the raw measurement sets, writes tidy Parquet tables, defines every derived product as a DuckDB SQL view, pins the analytical results with tests, and exports the CSVs the dashboards consume.
- Interactive Tableau Public dashboards — five published, covering all six chart families: the model-vs-reality scatter (+ coverage matrix), response curves, per-event boxplots, the slip regression with Kern County inference, and the faceted distributions with mean ± σ summary.
- A MkDocs companion site on GitHub Pages that embeds the dashboards and reads as an interactive version of the paper — glossary, data documentation, figure-by-figure crosswalk.
Architecture decisions live in docs/adr/: nine active ADRs, plus
the story of the dead ends — the earlier
SQL-engine/AWS/Superset architecture, retired or parked.
# Project-level (root owns these)
README.md, LICENSE, TODO.md, .gitignore, .gitattributes, .env.example
docs/
setup.md the developer manual: layout, pipeline, tools, setup
datasets.md the input datasets and the derived views
adr/ active ADRs + dead-ends.md (the retired architecture)
dashboards/ per-dashboard developer docs: data contracts,
calculated fields, editing traps
data/
raw/ raw CSVs (gitignored; drop inputs here — data/README.md)
interim/ intermediate cleaning artifacts (gitignored)
processed/ dir-per-table Parquet outputs (gitignored)
e.g. data/processed/dem/data.parquet
dist/ build outputs (gitignored)
csv/ egr-csv exports — what the published workbooks read
python/ the wheel
dashboards/
tableau/ seven .twb: five published `-public` workbooks, one
per dashboard, plus desktop copies of the two June
families (Athena, pre-pivot); README = workbook
index + publish procedure
duckdb/ generated views-only DuckDB file (gitignored)
sql/ generated DDL for the parked AWS lane (gitignored)
sheets/ dormant Google Sheets push — the central `dem` view
exceeds the Sheets cell cap (see dead-ends.md)
superset/ retired; README only (see dead-ends.md)
deploy/
terraform/ AWS data layer — parked; revival triggers in TODO.md
notes/ roadmap, chart inventory, dashboard build specs
(tracked — the site cites them as provenance);
design reviews, multi-machine notes and dated
working notes are local only (gitignored), so
references to those read as context, not links
resources/local/ local secrets, e.g. the Sheets service-account key
(gitignored; see .env.example)
ai/ initial scoping conversation (gitignored)
legacy/ original notebooks and 2025 paper PDF (gitignored)
# Gradle root — orchestrator only (ADR-0001)
settings.gradle.kts lists :subprojects:python, :subprojects:mkdocs,
:deploy:terraform
build.gradle.kts cross-cutting tasks (just the `base` plugin today)
.github/workflows/ mkdocs.yml — builds and deploys the site to Pages
# Code modules (Gradle subprojects)
subprojects/
python/ Poetry-managed pipeline package
(see subprojects/python/README.md)
mkdocs/ the companion site (MkDocs Material)
See ADR-0001 for the layout rationale.
Via Gradle (orchestrates Poetry behind the scenes):
./gradlew :subprojects:python:pytest # tests
./gradlew :subprojects:python:egrBuild # pipeline: Parquet + views
./gradlew :subprojects:python:egrBuildAndExport # …then every CSV into dist/csv/Or directly via Poetry (activate the project venv first — see
docs/setup.md):
cd subprojects/python
poetry install
poetry run pytest
poetry run egr-build # data/processed/<table>/ + dashboards/duckdb/eps.duckdb
poetry run egr-csv --view dem # one view -> dist/csv/dem.csv; the workbooks read
# several each — csvExportAll (Gradle) does them allThen:
- Dashboards: open a workbook from
dashboards/tableau/in the Tableau app and refresh extracts so they rebuild from yourdist/csv/. The published versions live on the Tableau Public profile linked above. - Companion site, locally: one-time
poetry install --only docs --no-root(fromsubprojects/python), thencd subprojects/mkdocs && mkdocs serve. Deployment to GitHub Pages is automatic via.github/workflows/mkdocs.yml(ADR-0009). - Optional desktop lanes: Tableau Desktop can connect straight to
dashboards/duckdb/eps.duckdb; the AWS/Athena lane is parked — status and revival triggers inTODO.md→ Deployment.
docs/setup.md— the developer manual: layout, pipeline, theegr-*and Gradle tool surface, setup, known gapsdocs/adr/— active decisions + dead-ends.mddocs/datasets.md— the input datasets (DEM, FDHI, SURE, Kern) and the thirteen derived viewsdocs/dashboards/— per-dashboard developer docs: data contracts, calculated fields, how to edit a workbook safely;tableau-editing-notes.mdcollects the.twbtraps that apply to every workbookdashboards/tableau/README.md— workbook index (files, dashboards, published slugs) and the publish/republish proceduredata/README.md— the raw input filesegr-buildexpectsnotes/Roadmap.md— build plan and statuses;notes/chart-families.md— chart inventory;notes/dashboard-*-build-spec.md— per-dashboard specs;notes/multi-machine.md— working across two machines (absolute paths in workbooks, therapture/rupturefolder-name story)subprojects/python/README.md— pipeline package usage and the IDEA setupsubprojects/mkdocs/DEPLOY.md— site deployment, the byline, and the terms figures are shown under;subprojects/mkdocs/EMBEDS.md— the Tableau embed pattern, the view ↔ page ↔ size map, and the "How to cite" markupdeploy/terraform/README.md— the parked AWS lane;dashboards/duckdb/README.md— connecting a desktop client to the DuckDB file
Three kinds of material live here under three different terms. Reuse the one that matches what you are taking.
| Material | Terms |
|---|---|
| Code — the Python pipeline, the Tableau workbooks, the site's build configuration, scripts and styling | Apache-2.0 — see LICENSE |
| Prose — the companion site's written content and this repository's documentation | CC BY 4.0 — see LICENSE-docs |
| Figures from the paper — the pre-typeset originals shown on the companion site | © The Author(s) 2025. Shown under the publisher's author-sharing terms: non-commercial use, no derivatives, with the full citation alongside each figure. |
Two things those terms do not cover:
- The published article. Earthquake Spectra published it under a "© The Author(s) 2025" line with no Creative Commons licence. Nothing from the typeset version is reproduced in this repository or on the site.
- The input datasets. The DEM experiments, the FDHI flatfile, the SURE
database and the Kern County compilation are published elsewhere under
their own terms and are not redistributed here.
data/README.mdnames each source, and the companion site's How to cite this data lists them with DOIs.
If you use anything you found through this project, cite the paper — the citation is at the top of this README, and behind the "How to cite" button on every page of the site. A DOI issued for this repository or the site would identify the software, not the research.