Repository navigation
refactor: DI fetcher, parallel node polling, 14 tests, label cardinality fixes - #1
Conversation
…ity fixes - Split fetcher.go (780 lines) into fetcher_chain/cosmos/api/admin/ml.go + interface.go; all functions promoted to HTTPFetcher methods; double-dispatch in interface.go eliminated - collector.New() accepts fetcher.Fetcher + prometheus.Registerer for DI; global package-level metric vars replaced with Metrics struct + NewMetrics(reg) - collectParticipant() decomposed into 8 private methods; epoch boundary detection logic unchanged, history/state integrity verified - FetchMaxBlockHeightFromNodes: sequential (up to 50s) → parallel WaitGroup (≤10s); goroutine accumulation eliminated - collectNetworkParticipants: Reset() on NetParticipantWeight/NetNodePocWeight before each fill to prevent unbounded label cardinality - gonka_block_time_seconds split into gonka_block_time_local_seconds (local node /status) and gonka_block_time_network_seconds (public nodes) - dashboard_union.json + dashboard.json updated for renamed metric - defaultBlockNodes reduced to 2 entries + override comment - state.go: migration TODO dated 2026-Q3; max64 replaced with builtin max - Add 14 tests: state_test.go (10) + metrics_test.go (4) - .gitignore: /exporter → root-only to stop ignoring cmd/exporter/ source dir - CI: docker.yml triggers on refactor/v2 branch → image tagged :dev
There was a problem hiding this comment.
Pull request overview
Refactors the exporter to use dependency-injected fetchers and per-registry Prometheus metrics, while improving polling performance and expanding state/metrics test coverage.
Changes:
- Split the previous monolithic fetcher into focused files and introduced a
fetcher.Fetcherinterface withHTTPFetcherimplementation for DI/mocking. - Replaced global Prometheus metric registration with
metrics.NewMetrics(reg)returning a*metrics.Metricsbundle, and updated the collector to use it. - Added new tests for state/history handling and metrics registration/value recording.
Reviewed changes
Copilot reviewed 15 out of 16 changed files in this pull request and generated 5 comments.
Show a summary per file
| File | Description |
|---|---|
| internal/state/state_test.go | Adds tests for miss-rate math and history load/save (including migration + pruning). |
| internal/state/state.go | Adds comments around old-format history migration path. |
| internal/metrics/metrics_test.go | Adds registry-based tests to ensure metrics register cleanly and record values. |
| internal/metrics/metrics.go | Refactors metrics into a struct created/registered via NewMetrics(reg); renames block-time metrics. |
| internal/fetcher/interface.go | Introduces Fetcher interface + HTTPFetcher constructor for DI. |
| internal/fetcher/fetcher_chain.go | Tendermint RPC + block-time helpers; parallel max-height polling. |
| internal/fetcher/fetcher_cosmos.go | Chain REST fetches (epoch, participant, tokenomics, PoC v2, etc.). |
| internal/fetcher/fetcher_api.go | Public API fetches (participants, pricing/models, stats, bridge, BLS epoch). |
| internal/fetcher/fetcher_admin.go | Admin API node list fetch. |
| internal/fetcher/fetcher_ml.go | ML node endpoints for GPU/service/health/driver metrics. |
| internal/fetcher/fetcher.go | Retains shared HTTP client + get() + flexInt64; removes monolith logic. |
| internal/config/config.go | Updates default public nodes list and clarifies env override. |
| internal/collector/collector.go | Switches collector to DI fetcher + metrics bundle; parallelizes node collection. |
| cmd/exporter/main.go | Wires DI: collector.New(cfg, fetcher.NewHTTPFetcher(), prometheus.DefaultRegisterer). |
| .gitignore | Fixes ignore patterns to avoid ignoring cmd/exporter/ sources. |
| .github/workflows/docker.yml | Publishes :dev on branch pushes to refactor/v2; latest only on tags. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| func (h *HTTPFetcher) FetchMaxBlockHeightFromNodes(nodes []string) (int64, string) { | ||
| sample := rand.Perm(len(nodes)) | ||
| if len(sample) > 5 { | ||
| sample = sample[:5] | ||
| } |
There was a problem hiding this comment.
FetchMaxBlockHeightFromNodes will panic when nodes is empty because rand.Perm(len(nodes)) is called with 0. This can happen if BLOCK_HEIGHT_NODES is set but contains only whitespace/commas (config will leave BlockHeightNodes empty). Add an early return when len(nodes)==0 (or fall back to defaultBlockNodes) before calling rand.Perm.
| tw, _ := strconv.ParseInt(r.EpochGroupData.TotalWeight, 10, 64) | ||
| ei, _ := strconv.ParseInt(r.EpochGroupData.EpochIndex, 10, 64) |
There was a problem hiding this comment.
FetchEpochGroupData ignores ParseInt errors for total_weight and epoch_index (tw, _ := ..., ei, _ := ...). If the API ever returns a non-numeric value, this will silently report 0s and skew downstream reward/weight metrics. Please handle the parse errors and return a descriptive error instead of discarding them.
| tw, _ := strconv.ParseInt(r.EpochGroupData.TotalWeight, 10, 64) | |
| ei, _ := strconv.ParseInt(r.EpochGroupData.EpochIndex, 10, 64) | |
| tw, err := strconv.ParseInt(r.EpochGroupData.TotalWeight, 10, 64) | |
| if err != nil { | |
| return nil, fmt.Errorf("parse current_epoch_group_data total_weight %q: %w", r.EpochGroupData.TotalWeight, err) | |
| } | |
| ei, err := strconv.ParseInt(r.EpochGroupData.EpochIndex, 10, 64) | |
| if err != nil { | |
| return nil, fmt.Errorf("parse current_epoch_group_data epoch_index %q: %w", r.EpochGroupData.EpochIndex, err) | |
| } |
| func (c *Collector) collectNodes() { | ||
| addr := c.cfg.Participant | ||
| if addr == "" { | ||
| addr = "unknown" | ||
| } | ||
|
|
||
| nodes, err := fetcher.FetchNodes(c.cfg.AdminAPIURL) | ||
| nodes, err := c.f.FetchNodes(c.cfg.AdminAPIURL) | ||
| if err != nil { | ||
| slog.Warn("fetch nodes", "err", err) | ||
| return | ||
| } | ||
|
|
There was a problem hiding this comment.
PR description mentions resetting node/GPU GaugeVecs to avoid stale series when the node set changes, but collectNodes doesn't reset any of the Node* / NodeGPU* vectors before repopulating them. Without a reset (or explicit deletion), metrics for removed nodes/devices will linger indefinitely. Consider calling Reset() on the relevant GaugeVecs at the start of collectNodes (or tracking current node IDs and deleting missing series).
| data, _ := json.Marshal(h) | ||
| os.WriteFile(path, data, 0644) |
There was a problem hiding this comment.
These tests ignore errors from json.Marshal and os.WriteFile. If either fails (e.g., unexpected marshal error or filesystem issue), the test may proceed with invalid setup and produce misleading failures. Please assert require.NoError/t.Fatalf on these errors so failures point to the real cause.
| data, _ := json.Marshal(h) | |
| os.WriteFile(path, data, 0644) | |
| data, err := json.Marshal(h) | |
| if err != nil { | |
| t.Fatalf("json.Marshal failed: %v", err) | |
| } | |
| if err := os.WriteFile(path, data, 0644); err != nil { | |
| t.Fatalf("os.WriteFile failed: %v", err) | |
| } |
| data, _ := json.Marshal(old) | ||
| os.WriteFile(path, data, 0644) | ||
|
|
There was a problem hiding this comment.
Same as above: this test ignores errors from json.Marshal / os.WriteFile, which can mask setup failures and make the assertion failures harder to interpret. Please check and fail the test immediately if these calls return an error.
- fetcher_chain: early return when nodes slice is empty - fetcher_cosmos: propagate ParseInt errors for total_weight/epoch_index - collector: Reset all node/GPU GaugeVecs at start of collectNodes - state_test: handle json.Marshal/os.WriteFile errors in test setup
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 15 out of 16 changed files in this pull request and generated no new comments.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Summary
fetcher.go(~730 lines) into 5 focused files (fetcher_chain.go,fetcher_cosmos.go,fetcher_api.go,fetcher_admin.go,fetcher_ml.go) with aFetcherDI interfacesync.WaitGroupinFetchMaxBlockHeightFromNodes— replaces sequential loop (up to 50s → ~10s with 5 nodes × 10s timeout)GaugeVec.Reset()before updating node/GPU metrics — eliminates stale series accumulation when node set changesgonka_block_time_seconds→gonka_block_time_local_seconds/gonka_block_time_network_secondsfor clarityEpochHistory(nested format + backward-compatible migration from flat format)refactor/v2branch now publishes:devimage tag.gitignore:exporter→/exporter(bare pattern was blockingcmd/exporter/source directory)internal/metrics/metrics_test.go,internal/state/state_test.goTest plan
go test ./...— 14/14 PASSgo build ./...+go vet ./...— no errors:devbuilds and publishes via GitHub Actions on push torefactor/v2:9404/metrics, healthz at:9404/healthzdashboard_union.json— panels render correctly$epochdropdown shows only numeric epoch values🤖 Generated with Claude Code