Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
30 commits
Select commit Hold shift + click to select a range
481f4e8
refactor(core): extract internal/chain and internal/cosmovisor packag…
SegovChik May 16, 2026
9478025
feat(core): ModelRecipe registry + arch-aware recommendConfig + Kimi …
SegovChik May 16, 2026
4b8f44a
feat(cmd): setup --model flag + Blackwell-only menu filter + phase 02…
SegovChik May 16, 2026
9dad8cc
feat(chain): PoCSafetyChecker with 2s fail-closed timeout for hotfix …
SegovChik May 16, 2026
894ed70
feat(cosmovisor): atomic api-binary installer with sha256, rollback, …
SegovChik May 16, 2026
ed2939e
feat(api): wire cosmovisor installer into gonka-nop update --api-only
SegovChik May 16, 2026
f64200b
feat(cmd): gonka-nop govmodel {list|delegate|refuse|declare-intent}
SegovChik May 16, 2026
42671fa
feat(cmd): ml-node update — PoC-gated read-merge-write PUT
SegovChik May 16, 2026
6052a54
feat(status): add Governance Models section + per-model state fetch
SegovChik May 16, 2026
71369ed
fix(phases): phase 02 prompts must honor --yes non-interactive mode
SegovChik May 16, 2026
3d3d31f
fix(phases): reject keyring passwords >50 chars (cosmos-sdk bcrypt 72…
SegovChik May 16, 2026
a9b8c9c
feat(status): surface state-sync chunk progress + sync-phase diagnostics
SegovChik May 16, 2026
d3832e7
fix(status): strip ANSI codes before parsing CometBFT statesync logs
SegovChik May 16, 2026
9ac9d3e
fix(chain): parse poc-delegation response with delegate_to (not deleg…
SegovChik May 16, 2026
f8f7254
feat(cmd): collateral show/deposit/withdraw — wraps cosmos collateral…
SegovChik May 16, 2026
b432a73
fix(collateral): parse stderr from inferenced query + extract txhash …
SegovChik May 16, 2026
6b0de5e
feat(status): add Collateral section to gonka-nop status
SegovChik May 16, 2026
8a9319b
feat(config): add ImageVersions.Versiond field + parser + fallback (F…
SegovChik May 17, 2026
8b62a13
feat(docker): add AtomicReplace primitive for compose mutations (NFR-…
SegovChik May 17, 2026
8e677bd
feat(devshard): add upstream blocks + image allowlist regex (FR-M12-1…
SegovChik May 17, 2026
cab4ae9
feat(config): add DevshardConfig + Postgres setup flags (FR-M12-5/6/14)
SegovChik May 17, 2026
0229910
feat(phases): render versiond block + api/proxy env additions (FR-M12…
SegovChik May 17, 2026
6049f25
feat(phases): devshards dirs + port 9400 + versiond pull WARN (FR-M12…
SegovChik May 17, 2026
9c10fcf
feat(cmd): add update --service api --image (FR-M12-10, INV-M12-3/4)
SegovChik May 17, 2026
7d4acb8
feat(chain): add repair-permissions + permissions DiffPermissions/Gra…
SegovChik May 17, 2026
b06f903
fix(chain): permissions tx requires 2 args + gas bump (live mainnet v…
SegovChik May 17, 2026
8749886
chore: bump Go to 1.25.10 + lint/cleanup pass
SegovChik May 22, 2026
2492670
feat(ui): surface governance models + candidates discovery in status/…
SegovChik May 22, 2026
054a2af
docs: remove personal namespaces and stale tags from README
SegovChik May 22, 2026
caa9027
docs: add nop-guide.md with E2E walkthroughs per topology + README ↔ …
SegovChik May 22, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ jobs:
- name: Set up Go
uses: actions/setup-go@v5
with:
go-version: "1.25.8"
go-version: "1.25.10"

- name: Download dependencies
run: go mod download
Expand Down Expand Up @@ -48,7 +48,7 @@ jobs:
- name: Set up Go
uses: actions/setup-go@v5
with:
go-version: "1.25.8"
go-version: "1.25.10"

- name: golangci-lint
uses: golangci/golangci-lint-action@v7
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ jobs:
- name: Set up Go
uses: actions/setup-go@v5
with:
go-version: "1.25.8"
go-version: "1.25.10"

- name: Install Cosign
uses: sigstore/cosign-installer@v3
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/security.yml
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ jobs:
- name: Set up Go
uses: actions/setup-go@v5
with:
go-version: "1.25.8"
go-version: "1.25.10"

- name: Install govulncheck
run: go install golang.org/x/vuln/cmd/govulncheck@latest
Expand Down
154 changes: 11 additions & 143 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,17 +7,6 @@

One-command CLI for deploying GPU-accelerated Gonka AI inference nodes. Auto-detects GPUs, configures NVIDIA runtime, generates optimized configs, and deploys containers with security hardening.

## What's New (v0.2.1-rc5)

### Blackwell GPU support improvements
- **Auto-detect latest Blackwell image**: queries GHCR registry for the latest `-blackwell` tag (e.g., `3.0.12-post6-blackwell`) instead of using a hardcoded version. Falls back to suffix convention if registry is unreachable.
- **`--mlnode-image` flag**: override the MLNode Docker image entirely for custom builds (e.g., `ghcr.io/segovchik/gonka-b300-image:3.0.13-b300-tp1`). Takes highest priority over all auto-detection.
- **`--attention-backend` flag**: choose between `FLASHINFER` (default) and `FLASH_ATTN` (lower memory footprint for constrained setups like 2×B200 where FlashInfer workspace causes OOM).
- **Interactive prompts**: setup wizard now asks for custom image and attention backend after showing recommended config. Skipped in `--yes` mode when flags are provided.

### Bug fix
- **`gonka-nop status` broken as root**: when `state.json` loaded successfully but `AdminURL` was empty, all HTTP requests silently failed. Non-root users weren't affected because they couldn't read `state.json` (0600 permissions), triggering the default-URL fallback path. Fixed by always initializing `StatusConfig` from defaults before overriding with state values.

## Demo

[![Gonka NOP Demo](https://img.youtube.com/vi/0w6bIEROxUQ/maxresdefault.jpg)](https://youtu.be/0w6bIEROxUQ?si=FJywggqlVax90Ohn)
Expand Down Expand Up @@ -158,138 +147,6 @@ gonka-nop ml-node add
gonka-nop ml-node list
```

## GPU-Specific Deployment Guides

NOP auto-detects GPU architecture and selects optimal settings. These guides document real-world tested configurations and known issues per hardware class.

### 8× A100 SXM4 80GB (Ampere, sm_80)

Standard configuration. Works out of the box.

```bash
gonka-nop setup
```

| Setting | Value |
|---------|-------|
| Image | `mlnode:3.0.12-post6` (auto-selected) |
| TP | 4 (auto, NOP calculates PP=2 for 8 GPUs) |
| Backend | FLASHINFER |
| gpu-memory-utilization | 0.90 |
| Weight (observed) | ~860 per ML node |

### 2× B200 (Blackwell, sm_100)

Blackwell image auto-detected. **Use FLASH_ATTN** to avoid FlashInfer workspace OOM with chain-enforced `max_model_len=240000`.

```bash
gonka-nop setup --attention-backend FLASH_ATTN
```

| Setting | Value |
|---------|-------|
| Image | `mlnode:3.0.12-post6-blackwell` (auto from GHCR) |
| TP | 2 |
| Backend | **FLASH_ATTN** (FLASHINFER causes OOM on 2×B200) |
| gpu-memory-utilization | 0.88 |
| Weight (observed) | ~920 with [T,T] timeslots |

**Known issue:** The chain enforces `--max-model-len 240000` which pre-allocates a large KV cache. With TP=2 on 2×B200, FlashInfer's workspace buffer (~285MB) doesn't fit in the remaining free memory. Switching to FLASH_ATTN eliminates the workspace allocation entirely.

### 4× B200 (Blackwell, sm_100)

Works with alpha4 image or blackwell image. FLASHINFER is OK because TP=4 means less model weight per GPU → more headroom.

```bash
gonka-nop setup
```

| Setting | Value |
|---------|-------|
| Image | `mlnode:3.0.13-alpha4` or `3.0.12-post6-blackwell` |
| TP | 4 |
| Backend | FLASHINFER (enough headroom at TP=4) |
| gpu-memory-utilization | 0.90 |
| Weight (observed) | ~2,178 |

### 8× B300 SXM6 AC (Blackwell Ultra, sm_103a)

Requires custom image — standard and blackwell images lack sm_103a CUTLASS kernels. Must use `vllm/vllm-openai:v0.15.1-cu130` as base.

```bash
gonka-nop setup \
--mlnode-image ghcr.io/segovchik/gonka-b300-image:3.0.13-b300-tp1
```

| Setting | Value |
|---------|-------|
| Image | **Custom** (`ghcr.io/segovchik/gonka-b300-image:3.0.13-b300-tp1`) |
| TP | 1 (8 independent instances, one per GPU) |
| Backend | FLASHINFER |
| gpu-memory-utilization | 0.95 |
| max-model-len | 131072 |
| max-num-seqs | 128 |
| Weight (observed) | ~7,700–8,300 |
| PoC throughput | ~8,700 nonces/min |

**Why custom image?** B300 (sm_103a) needs:
1. CUDA 13.0 base (`vllm/vllm-openai:v0.15.1-cu130`) for CUTLASS kernel compatibility
2. Triton ptxas 13.0 (replacing bundled 12.8 that doesn't know sm_103a)
3. Runner patches: TP=1 for maximum PoC throughput (8 instances vs 2 with TP=4)

**Host requirement:** `cuda-compat-13-0` package must be installed if host CUDA toolkit < 13.0. Mount compat libs into container:
```yaml
# docker-compose.mlnode.yml
volumes:
- /usr/local/cuda-13.0/compat:/usr/local/cuda/compat:ro
environment:
- LD_LIBRARY_PATH=/usr/local/cuda/compat
```

See [B300 deployment report](docs/b300-mlnode-deployment.md) for the full build process.

### 8× H100/H200 SXM 80GB (Hopper, sm_90)

Standard configuration similar to A100.

```bash
gonka-nop setup
```

| Setting | Value |
|---------|-------|
| Image | `mlnode:3.0.12-post6` (auto-selected) |
| TP | 4 or 8 (NVLink full mesh enables TP=8) |
| Backend | FLASHINFER |
| gpu-memory-utilization | 0.90 |

### Changing Image on a Running Node

Use `ml-node set-image` to swap the MLNode image without full re-setup:

```bash
# Switch to blackwell image
gonka-nop ml-node set-image ghcr.io/product-science/mlnode:3.0.12-post6-blackwell

# Switch to custom B300 image
gonka-nop ml-node set-image ghcr.io/segovchik/gonka-b300-image:3.0.13-b300-tp1
```

This performs a safe rollout: disable → update compose → pull → recreate → enable.

### Spot Instance Recovery

If a spot/preemptible instance is killed and reprovisioned with the old data disk attached:

1. Mount old disk: `mount /dev/vdb4 /mnt`
2. Move Docker storage: set `data-root` in `/etc/docker/daemon.json` to mounted disk
3. Symlink gonka-node: `ln -sf /mnt/root/gonka-node /root/gonka-node`
4. Fix HF cache path: `ln -sf /mnt/mnt/shared/huggingface /mnt/shared/huggingface`
5. Install CUDA compat if needed: `apt-get install cuda-compat-13-0`
6. Start: `set -a && source config.env && set +a && docker compose up -d`

Keys, chain data, and model cache survive on the persistent disk. No re-registration needed if IP stays the same.

## Commands

| Command | Description |
Expand All @@ -307,11 +164,22 @@ Keys, chain data, and model cache survive on the persistent disk. No re-registra
| `ml-node status` | Detailed ML node status |
| `ml-node enable/disable` | Enable or disable an ML node |
| `ml-node set-image` | Change MLNode Docker image and restart (safe rollout) |
| `govmodel list` | Show this node's per-model PoC opt-in state |
| `govmodel candidates <model>` | Rank on-chain hosts running `<model>` for delegation |
| `govmodel delegate <model> <addr>` | Delegate PoC weight for `<model>` to `<addr>` |
| `govmodel refuse <model>` | Refuse participation in `<model>`'s PoC |
| `govmodel declare-intent <model>` | Bootstrap-only signal of intent to run `<model>` |
| `collateral show [addr]` | Show on-chain collateral balance |
| `collateral deposit <amount>ngonka` | Deposit collateral to back PoC weight |
| `collateral withdraw <amount>ngonka` | Withdraw collateral (1-epoch unbonding, still slashable) |
| `repair-permissions` | Re-grant missing ml-ops permissions to the warm key |
| `download-model` | Pre-download model weights before setup |
| `reset` | Stop containers and clean up |
| `cleanup` | Recover disk space |
| `version` | Print version info |

See [`docs/nop-guide.md`](docs/nop-guide.md) for full end-to-end walkthroughs per topology, including the post-epoch-180 collateral + governance opt-in flow.

## Setup Flags

| Flag | Description | Used in |
Expand Down
Loading
Loading