Skip to content

feat(dynamo-preproc): route with advisory EPP signals - #963

Draft
ayushag-nv wants to merge 1 commit into
mainfrom
ayushag/signal-extraction-test
Draft

ayushag-nv wants to merge 1 commit into
mainfrom
ayushag/signal-extraction-test

Conversation

@ayushag-nv

Copy link
Copy Markdown
Contributor

Why

Switchyard PreProc can choose a model, but it currently has no request-specific view of that model's cache reuse or worker load. This adds an optional way to collect those signals from Dynamo EPP and use them in an SDK routing decision.

What

  • Add a bounded probe client to examples/dynamo-preproc. PROBE_CONFIG maps eligible models to private EPP probe endpoints; observe mode records signals and routing mode passes them to the SDK.
  • Add typed host observations and a TOML-configured cache_aware algorithm. Missing, stale or incomplete signals fall back to the first configured target.
  • Preserve request/session identity for predicted-versus-actual reuse logs. The existing StageRouter setup still works when probes are disabled.

How

PreProc sends each eligible model's original request to its EPP, then makes one SDK decision. The gateway forwards the chosen model to its normal EPP selection path.

The policy compares (effective_prefill_tokens + active_prefill_tokens) * relative_prefill_cost. Operators must choose models suitable for the task and configure their relative costs. This estimates prefill work; it does not assess model quality or predict end-to-end latency.

Probe calls share a deadline and concurrency limit. Errors remain unknown signals. Model-specific tokenization, cumulative cache-tier estimates and actual response-usage metrics belong to the companion Dynamo draft PR #16053.

Where to Start Review

Start with examples/dynamo-preproc/README.md and the two new config examples. Then review examples/dynamo-preproc/src/probe.rs, crates/libsy/src/algorithms/cache_aware.rs and its Runner configuration. Metadata::serving_observations is supplied by the host, never parsed from caller headers.

Test Plan

  • PreProc tests: 9 passed, including fresh, missing, stale and unknown-load observations through the SDK.
  • Workspace type check and PreProc Clippy with warnings denied passed; changed-file formatting and whitespace checks passed.
  • Local gateway -> PreProc -> EPP -> SGLang run: cold routing chose Qwen3-0.6B; warming only Qwen3-1.7B changed the choice to that model.
  • Verified observe mode and fallback when a probe was unavailable or rejected tenant-namespace input. Inference still completed.
  • Warm predicted and actual reuse both measured 624/632 tokens, with request/model/worker correlation checked for JSON and SSE responses.

The live check used Dynamo 1.5.0 SGLang workers and the existing agentgateway v1.5.0 Dynamo build. This is functional validation, not a production latency benchmark. The probe client requires the companion Dynamo change; CPU/disk cache residency was not exercised.

Part of SWITCH-1627.

Signed-off-by: ayushag <ayushag@nvidia.com>
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
PR Preview Action v1.8.1

🚀 View preview at
https://NVIDIA-NeMo.github.io/Switchyard/pr-preview/pr-963/

Built to branch gh-pages at 2026-10-09 22:42 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant