A local demo for investigating renewable energy assets with specialised agents, cited evidence and inspectable execution traces. All assets, telemetry, forecasts, alarms and documents are synthetic. This is a prototype on synthetic data, not a production deployment.
The demo answers an operational question: why did an asset underperform, which records support the explanation, and what still needs human verification?
Python 3.12 and uv are required. From this directory:
# Fast, deterministic baseline. No LLM and no model downloads.
powershell -ExecutionPolicy Bypass -File scripts/start-demo.ps1
# Actual local tool-calling model and vector retrieval.
# First run downloads a portable Ollama runtime, qwen3:8b and nomic-embed-text.
powershell -ExecutionPolicy Bypass -File scripts/start-demo.ps1 -Backend ollamaOpen http://127.0.0.1:8000. The page displays the active backend and retrieval mode. Ctrl+C stops the API. The portable Ollama service remains available for subsequent runs; its PID and logs are under runtime/. Model files stay inside this project. No paid API or credentials are required.
On the prepared Windows machine, double-click Start Demo.cmd for the actual local model or Start Replay.cmd for the baseline. Run one API instance at a time. The launchers print the local URL; open it in your browser.
After installation, already-present models are reused without a registry request; routine launches can run offline.
If Ollama is already listening on port 11434, setup uses that service; its existing model storage configuration applies. Budget roughly 10 GB disk space for the portable runtime, download archive, model files and Python environment. Runtime memory and latency depend on the machine; 8 GB VRAM is not an unconditional fit guarantee.
- Solar / Asset Manager / August 18–22: ask “Explain the energy underperformance and any conflicting records.” Follow the
asset → omhand-off, inspectINV-OT, and compare the closure note with the alarm window. - Wind / Analyst / August 25–26: ask about forecast error and missing telemetry. Only 18 of 24 hours match. MAE excludes the six missing observations.
- Battery / Asset Owner / August 18–22: inspect charge, discharge and net energy. Net energy is explicitly not round-trip efficiency.
Windows are UTC, start-inclusive and end-exclusive, aligned to hours. Dataset coverage is June 1 through August 31, 2026. The interface and agent prompts use English.
flowchart TD
U[Validated role, asset, UTC window, question] --> R[Deterministic policy router]
R --> M[Asset specialist]
R --> O[O&M specialist]
R --> A[Analyst specialist]
R --> X[Unsupported request refusal]
M -->|Overlapping alarm and shared evidence| O
M & O & A --> T[Validated read-only tools]
T --> D[(DuckDB: telemetry, alarms, forecasts)]
T --> C[(Chroma + local embeddings / lexical baseline)]
T --> E[Exact evidence report + optional model interpretation]
E --> V[Viewer, JSON trace and feedback]
- LangGraph implements routing, specialist nodes, shared evidence and a bounded hand-off. Routing policy is deliberately explicit rather than another model invocation. Asset Owner uses the asset specialist; Trader is unsupported because no market data exists. Roles are workflow hints, not authenticated identities.
- In Ollama mode, each specialist chooses tools using a bounded tool-calling loop. A separate structured synthesis call interprets evidence and lists source IDs. Unsupported citations are rejected. Source existence is not proof that the interpretation is true.
- A specialist that returns prose without its required evidence receives a missing-tool correction within the four-turn budget. Completion requires that specialist's evidence checklist, including shared evidence from a hand-off. This is an explicit workflow policy, not a hidden fallback to replay.
- In replay mode, specialists invoke a fixed tool sequence. It exercises data, graph, validation and trace contracts; it is not an AI performance benchmark and never silently replaces failed Ollama calls.
- The four tools accept typed arguments. SQL is parameterised; connections are read-only. Model-selected asset and dates must match the request. No command, trade, maintenance execution or work-order tool exists.
- Numeric evidence is rendered directly from tool results. Model interpretation is shown separately with uncertainty and a human-review label. There is no automated semantic verification of model prose.
- Battery net energy uses discharge minus charge. Python computes the net-flow label; the assessment's structured battery direction must agree. This targeted consistency check does not certify the rest of the model's prose.
- Each short document is one retrieval chunk with a stable ID. Chroma collections are versioned by corpus content and embedding-model name. Lexical retrieval is an explicit baseline, not embeddings. Queries are filtered to the selected asset plus shared methodology/policy.
uv sync --frozen --link-mode copy
.venv/Scripts/python -m horizon seed
.venv/Scripts/python -m horizon doctor
.venv/Scripts/python -m horizon ingest
.venv/Scripts/python -m horizon --backend ollama --retrieval chroma ask "Investigate inverter alarm and conflicting logbook records" --role "O&M Manager"
.venv/Scripts/python -m pytest -q
.venv/Scripts/python -m horizon eval --runs 3
.venv/Scripts/python -m horizon --backend ollama --retrieval chroma eval --cases eval/smoke.jsonHORIZON_BACKEND (replay or ollama), HORIZON_RETRIEVAL (lexical or chroma), OLLAMA_URL, OLLAMA_MODEL, OLLAMA_EMBED_MODEL, and HORIZON_DATA_DIR configure the service. CLI backend/retrieval flags override environment defaults. Regeneration replaces the synthetic tables; do it while the API is stopped. runtime/ is ignored by Git.
The HTTP API has /health, /assets, /ask, /traces, /traces/{id} and /traces/{id}/feedback; interactive API documentation is at /docs. Only one investigation runs at a time to bound GPU pressure. Traces include full user prompts and retrieved text, so use synthetic inputs only. The service does not supply authentication, per-user data isolation, retention enforcement or distributed rate limits.
The 22-case suite covers seeded anomalies, tool selection and argument scope, hand-offs, arithmetic bounds, no-data windows, missingness, alarm boundaries, unsupported operations and invalid input. Repeated runs report latency and pass-rate variation. JSON results are in eval/results.
The source bundle includes the referenced synthetic traces under eval/results/traces/, so a reviewer can inspect measured runs without installing the model. After new evaluations, run python scripts/export-eval-traces.py and python scripts/package-demo.py to refresh the shareable source ZIP in dist/.
Citation integrity and contract success are not semantic groundedness or a hallucination-rate measurement. Model quality needs human review against source records and a broader held-out set. No LLM judge is presented as objective truth. Test doubles verify protocol handling and rejection paths separately from actual local-model runs.
Measured on the prepared Windows / RTX 5060 Laptop machine:
| Check | Result |
|---|---|
| Python tests | 20 passed |
| Replay, 22 cases × 3 runs | 66/66 contract passes |
| Ollama + Chroma, 22 cases × 2 runs | 44/44 contract passes |
| Subset that actually invoked the model | 24/24 contract passes; mean 40.96 seconds |
The remaining 20 actual-backend checks exercised validation/refusal policy without inference. The report also records remaining semantic errors in model prose; the above scores do not conceal or measure them.
See report.md for measured results and production gaps, and docs/demo-script.md for a short presentation script.
docker build -t horizon-agent-demo .
docker run --rm -p 127.0.0.1:8000:8000 horizon-agent-demoThis starts the replay baseline as a non-root user. Mount a writable /app/runtime volume to preserve traces. Connecting a container to host Ollama requires explicit network configuration and a prebuilt Chroma index. The primary supported workflow is the Windows local script. CI runs Python checks and the replay suite; it does not download an LLM or validate a GPU deployment.
Implementation follows the official Ollama tool-calling protocol, embedding API, LangGraph Graph API, and Chroma explicit-vector query API.