Skip to content

feat(epp): expose advisory cache and load probes - #16053

Draft
ayushag-nv wants to merge 1 commit into
mainfrom
ayushag/signal-extraction-test
Draft

ayushag-nv wants to merge 1 commit into
mainfrom
ayushag/signal-extraction-test

Conversation

@ayushag-nv

@ayushag-nv ayushag-nv commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Why

A model router needs to compare cache reuse and load before choosing a model. EPP already has the tokenizer, KV index and worker state, but does not expose that information for a request without taking the normal inference path.

What

Add an opt-in POST /v1/routing/probe endpoint to native EPP. It accepts a Chat Completions request for a concrete model and returns a candidate worker, predicted cache reuse and tracked load. It is disabled by default.

Details

How

  • Reuse EPP preprocessing and the embedded SelectionService with advisory selection. Probes do not reserve capacity or retain a selection; normal EPP routing still chooses the worker for inference.
  • Bound body size, concurrency and total probe time. Keep the listener private: it receives prompts and has no authentication.
  • Calculate GPU, CPU-inclusive and disk-inclusive hit-rate estimates using the model's own tokenized prompt length. Preserve unknown values and use cumulative tier counts without adding tiers together.
  • Record actual reuse from response usage separately. Correlate predictions and outcomes through request/session/model/worker logs; keep identities out of Prometheus labels. Actual aggregate rates use matching cached-token and prompt-token totals.

This draft supports native discovery, aggregated DP=1 workers, and text/function-tool chat requests. Unsupported modes, including salted/tenant cache namespaces, return an error that callers can treat as an unavailable signal. Lower-tier estimates remain unknown without evidence of CPU/disk residency.

Where to Start Review

Start with deploy/inference-gateway/ext-proc/routing-probe.md for the API and limits, then src/probe.rs and Router::probe in src/epp.rs. The lib/llm/src/kv_router* changes provide access to the existing advisory selection path.

Validation

Test Plan

  • EPP unit tests: 164 passed. Clippy with warnings denied, changed-file formatting, docs lint and whitespace checks passed.
  • Local two-model SGLang run through gateway, Switchyard PreProc and EPP: cold routing selected Qwen3-0.6B; warming only Qwen3-1.7B changed the selection to that model.
  • Warm predicted and actual reuse both measured 624/632 tokens. JSON and SSE usage produced a token-weighted aggregate of 1248/1589 tokens.
  • Twenty advisory calls left pending requests and tracked prefill work unchanged. Live inference raised load signals, which returned to idle afterward.
  • Verified tool history, unavailable-probe fallback, input limits, overload and timeout recovery.

Tested with the existing Dynamo 1.5.0 SGLang workers and agentgateway v1.5.0 Dynamo build. CPU/disk residency and stock agentgateway 1.0 were not validated in this run. These are functional checks, not a production latency benchmark.

Related Issues

Part of SWITCH-1627. Companion Switchyard draft PR #963 adds the probe client and SDK routing policy.

Signed-off-by: ayushag <ayushag@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 9, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions github-actions Bot added feat documentation Improvements or additions to documentation router Relates to routing, KV-aware routing, etc. labels Oct 9, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation feat router Relates to routing, KV-aware routing, etc. size/XL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant