Skip to content

fix(translation): report Anthropic context-window stops as a token limit - #918

Open
bharadwaj-pendyala wants to merge 1 commit into
NVIDIA-NeMo:mainfrom
bharadwaj-pendyala:fix/anthropic-context-window-stop
Open

bharadwaj-pendyala wants to merge 1 commit into
NVIDIA-NeMo:mainfrom
bharadwaj-pendyala:fix/anthropic-context-window-stop

Conversation

@bharadwaj-pendyala

@bharadwaj-pendyala bharadwaj-pendyala commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

What

Treat Anthropic's model_context_window_exceeded stop reason as a token limit, like max_tokens. This updates map_anthropic_stop_reason for buffered responses and the Anthropic stream decoder's message_delta handling.

Why

Sonnet 4.5 and newer return model_context_window_exceeded without a beta header when the answer runs into the context window. Anthropic's docs say to treat it as truncated. Switchyard doesn't know the spelling, so a cut-off answer reaches other clients like this on b9e7ccce:

Path Client Gets
buffered Chat finish_reason: "stop"
buffered Responses status: "completed"
stream Chat finish_reason: "model_context_window_exceeded"
stream Responses response.completed

In the streamed Chat row, the raw Anthropic string isn't one of the values the OpenAI spec allows for finish_reason. The other three rows report a truncated answer as complete. max_tokens on the same inputs gives length and incomplete everywhere, as set up in #300.

The stream decoder change also covers streams buffered by the server. ResponseAccumulator passes the chunk reason to stop_reason_from_str in protocol/src/stream.rs, which returns Unknown for this spelling today. Emitting max_tokens instead produces MaxTokens without changing protocol. The Responses decoder already maps response.incomplete to max_tokens.

Notes for reviewers

  • Same-format Anthropic traffic keeps the original spelling. The buffered encoder returns the preserved body, and preserved stream events replay as received.
  • With preservation off, an Anthropic-to-Anthropic stream now ends with max_tokens where it used to say end_turn.
  • This doesn't handle pause_turn, which also passes through raw to Chat stream clients. It means "send the turn back to continue", and there's no Chat or Responses equivalent, so it needs its own decision.
  • feat(translation): support Bedrock Converse as a native wire format #909 adds the same spelling to stop_reason_from_str for Bedrock. This PR doesn't touch that file, so the two don't conflict.

Validation

Two regression tests, one buffered and one streamed, fail on b9e7ccce and pass on this branch. The streamed test also folds the decoded chunk through ResponseAccumulator and checks for StopReason::MaxTokens.

  • cargo fmt --all --check: clean
  • cargo clippy --workspace --all-targets --exclude switchyard-py -- -D warnings: clean
  • cargo test --workspace --exclude switchyard-py: 913 passed, 0 failed

I didn't build switchyard-py locally. Nothing in it changes here, and CI runs it.

Summary by CodeRabbit

  • Bug Fixes
    • Anthropic responses that exceed the model’s context window are now reported as token-limit stops, with consistent completion status across OpenAI Chat and Responses formats.

Signed-off-by: Bharadwaj Pendyala <bharadwajpendyala@gmail.com>
@bharadwaj-pendyala
bharadwaj-pendyala requested a review from a team as a code owner October 6, 2026 13:48
@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA-NeMo/Switchyard/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: 7c60d869-3446-4d48-b25b-8ecbaddb700e
📥 Commits

Reviewing files that changed from the base of the PR and between b9e7ccc and 3359e46.

📒 Files selected for processing (4)
  • crates/switchyard-translation/src/codecs/anthropic/buffered.rs
  • crates/switchyard-translation/src/codecs/anthropic/stream.rs
  • crates/switchyard-translation/tests/response_translation.rs
  • crates/switchyard-translation/tests/stream_translation.rs

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


Walkthrough

Changes

The translation maps Anthropic’s model_context_window_exceeded stop reason to the token-limit result in buffered and streaming paths. Tests cover OpenAI Chat, OpenAI Responses, and accumulator decoding.

Anthropic stop mapping

Layer / File(s) Summary
Normalize context-window stop reasons
crates/switchyard-translation/src/codecs/anthropic/buffered.rs, crates/switchyard-translation/src/codecs/anthropic/stream.rs, crates/switchyard-translation/tests/response_translation.rs, crates/switchyard-translation/tests/stream_translation.rs
Buffered decoding maps the stop reason to StopReason::MaxTokens. Streaming decoding normalizes it to max_tokens. Tests check the resulting OpenAI Chat and Responses outputs.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~8 minutes

Merge Risk: ⚪ Minimal · up to 3359e

The supplied review identifies no issue that needs resolution before merging.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 75.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 8 functions across 4 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly describes the main change: mapping Anthropic context-window stops to a token limit.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit checks the stop reason at dawn
The context-window label is gone
Chat reports length with care
Responses mark incomplete there
The rabbit hops on, token-limit drawn

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant