Skip to content

Commit 700a3ff

Browse files
docs(integrations): recommend anthropic_messages for Claude through an LLM gateway (#874)
* docs(integrations): add Claude gateway caching and thinking notes Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com> * docs(integrations): clarify when each Claude gateway setup applies Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com> * docs(integrations): estimate Claude cache costs and fix the thinking and key setup steps Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com> * docs(integrations): note when omit_body_fields holds and use Anthropic's published cache prices Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com> * docs(integrations): document forwarding one gateway key to OpenAI and Anthropic clients Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com> * docs(integrations): say only standalone switchyard-server forwards keys Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com> * docs(integrations): lead the Claude gateway sections with the format to use Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com> --------- Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
1 parent 611e270 commit 700a3ff

3 files changed

Lines changed: 202 additions & 4 deletions

File tree

‎docs/integrations/oh_my_pi.md‎

Lines changed: 45 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -29,7 +29,8 @@ providers:
2929
3030
- `auth: none` marks the provider as keyless. Switchyard ignores client keys unless an
3131
LLM client sets `forward_auth = true`. Without `auth: none`, `omp` refuses to send a
32-
request.
32+
request. To send your gateway key through a route that forwards it, see
33+
[Forwarded keys](#forwarded-keys).
3334
- `models[].id` must equal a route `id` from your TOML file. `contextWindow` and
3435
`maxTokens` set `omp`'s compaction limit and output cap. `reasoning: true` turns on the
3536
`--thinking` flag.
@@ -68,8 +69,8 @@ the default model.
6869
## Check the routing
6970

7071
The checks in [Use Switchyard with pi](pi.md#check-the-routing) work the same way for
71-
`omp`. On the Chat Completions API, `omp` sends no session header for a custom provider,
72-
so the routing log records `"session_id": null`. Routes with
72+
`omp`. On the Chat Completions and Responses APIs, `omp` sends no session header for a
73+
custom provider, so the routing log records `"session_id": null`. Routes with
7374
`classify_trigger = "user_turn"` or `"new_session"`, advisor budgets, and the stage
7475
router's `capable_hold_turns` then treat each request as its own session. If you need
7576
per-session routing, use `anthropic-messages`. On that API `omp` sends the
@@ -94,3 +95,44 @@ Port 4000 is also the default port for `omp`'s `litellm` provider and for
9495
as a LiteLLM proxy. In that case, run Switchyard on another port or set
9596
`LITELLM_BASE_URL`. Set `cost` on the model entry if you want `omp` to show a non-zero
9697
cost.
98+
99+
## Claude through an LLM gateway
100+
101+
The pi guide's [Claude through an LLM gateway](pi.md#claude-through-an-llm-gateway)
102+
section applies to `omp`: give Claude targets an LLM client with
103+
`format = "anthropic_messages"`. It shows how to check prompt caching and estimate the
104+
cost. This section covers what differs for `omp`, tested with Oh My Pi 18.2.11.
105+
106+
### Thinking on `anthropic-messages`
107+
108+
On `openai-completions` and `openai-responses`, Switchyard turns `omp`'s thinking level
109+
into adaptive thinking for an `anthropic_messages` target. On `anthropic-messages`,
110+
Switchyard sends `omp`'s own `thinking` object unchanged. `omp` does not
111+
recognize a route ID such as `switchyard` as a Claude model. With thinking on, it sends
112+
`thinking: {type: "enabled"}`, and Claude Opus 5.5 and Sonnet 5 return HTTP 400. Set
113+
adaptive thinking on the model entry:
114+
115+
```yaml
116+
- id: switchyard
117+
reasoning: true
118+
thinking:
119+
mode: anthropic-adaptive
120+
efforts: [low, medium, high]
121+
```
122+
123+
`omp` requires `efforts` next to `mode`. With both set, it sends
124+
`thinking: {type: "adaptive"}` and `output_config.effort`.
125+
126+
### Forwarded keys
127+
128+
To forward `omp`'s key, remove `auth: none` and set `apiKey: GATEWAY_API_KEY`, the name of
129+
the environment variable that holds your gateway key. Unlike pi, `omp` reads the name
130+
without a leading `$`. With `auth: none`, `omp` sends no key, and the gateway returns HTTP
131+
401. Only standalone `switchyard-server` forwards keys. The native Nemo Relay plugin
132+
rejects routes that use `forward_auth = true` (see
133+
[Request Handling](nemo_relay.md#request-handling)).
134+
135+
The pi guide's [Forwarded keys](pi.md#forwarded-keys) table shows which request APIs a
136+
route accepts when it forwards the key. A route that forwards the key to an OpenAI-format
137+
LLM client, such as a classifier's GPT judge, accepts only `openai-completions` and
138+
`openai-responses` requests.

‎docs/integrations/pi.md‎

Lines changed: 156 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -41,7 +41,9 @@ with route id `switchyard`.
4141

4242
- `models[].id` must equal a route `id` from your TOML file. Add one entry per route.
4343
- `apiKey` is a placeholder. Switchyard ignores client keys unless an LLM client sets
44-
`forward_auth = true`. pi still needs some value here before it lists the model.
44+
`forward_auth = true`. pi still needs some value here before it lists the model. To
45+
send your gateway key through a route that forwards it, see
46+
[Forwarded keys](#forwarded-keys).
4547
- `contextWindow` and `maxTokens` set pi's compaction limit and output cap. pi does not
4648
read these values from the server. Use the smallest context window among the route's
4749
targets.
@@ -96,3 +98,156 @@ clients keep the local id `switchyard` instead.
9698
Set `cost` on the model entry if you want pi to show a non-zero cost.
9799
[`benchmark/run-baseline.sh`](../../benchmark/README.md) runs Terminal-Bench tasks with
98100
pi through Switchyard when you pass `--agent pi`.
101+
102+
## Claude through an LLM gateway
103+
104+
An LLM gateway, such as a LiteLLM proxy, can serve Claude on three endpoints:
105+
`/v1/chat/completions`, `/v1/responses`, and `/v1/messages`. Switchyard calls the
106+
endpoint that matches the `format` of the Claude target's LLM client, whichever `api` pi
107+
uses. Give Claude targets an LLM client with `format = "anthropic_messages"`:
108+
109+
```toml
110+
[llm_clients.gateway_claude]
111+
format = "anthropic_messages"
112+
base_url = "https://gateway.example.com"
113+
api_key_env = "GATEWAY_API_KEY"
114+
115+
[targets.claude]
116+
id = "claude-sonnet-5" # the gateway's model ID
117+
llm_client = "gateway_claude"
118+
```
119+
120+
On the tested gateway, both OpenAI formats returned HTTP 400 when pi sent a thinking
121+
level, and `openai_responses` never read the prompt cache:
122+
123+
| Claude LLM client `format` | Gateway endpoint | Prompt cache | pi's thinking level |
124+
|---|---|---|---|
125+
| `anthropic_messages` | `/v1/messages` | Read on repeated prompts | Works |
126+
| `openai_chat` | `/v1/chat/completions` | Read on repeated prompts | HTTP 400 |
127+
| `openai_responses` | `/v1/responses` | Never read | HTTP 400 |
128+
129+
[Prompt caching](#prompt-caching) and [Thinking](#thinking) explain both failures. To
130+
use each developer's own gateway key instead of a key that the server holds, see
131+
[Forwarded keys](#forwarded-keys).
132+
133+
### Prompt caching
134+
135+
pi sends the whole conversation on every turn. When the gateway reads the repeated part
136+
from Claude's prompt cache, that part costs 0.1 times the input price on Claude Sonnet 5
137+
and 0.05 times on Claude Opus 5.5, according to Anthropic's
138+
[pricing page](https://platform.claude.com/docs/en/about-claude/pricing). The tested
139+
gateway never read Claude prompts from the cache on `/v1/responses`, so every turn there
140+
paid the full input price for the whole conversation.
141+
142+
To check your gateway, start the server with `--routing-log-file PATH` and send the same
143+
prompt twice. Claude does not cache short prompts, so use a prompt of at least 5,000
144+
tokens. Then read the records:
145+
146+
```bash
147+
jq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}' PATH
148+
```
149+
150+
When caching works, the first record shows the prompt in `cache_creation_tokens`, and the
151+
second shows `cached_tokens` close to `prompt_tokens`. If the second record shows
152+
`"cached_tokens": 0`, the gateway read nothing from the cache.
153+
154+
#### Estimate the cost
155+
156+
The routing log records token counts, not prices. To estimate what a request cost,
157+
multiply each count in its record by the matching price and add the results:
158+
159+
| Tokens in the record | Price |
160+
|---|---|
161+
| `prompt_tokens - cached_tokens - cache_creation_tokens` | Input |
162+
| `cached_tokens` | Cache read |
163+
| `cache_creation_tokens` | Cache write |
164+
| `completion_tokens` | Output |
165+
166+
Write the prices to `prices.json` in USD per million tokens, and key each entry by the
167+
record's `model` value. This example uses Anthropic's list prices for Claude Sonnet 5 on
168+
2026-10-02:
169+
170+
```json
171+
{
172+
"claude-sonnet-5": {"input": 2.00, "cache_read": 0.20, "cache_write": 2.50, "output": 10.00}
173+
}
174+
```
175+
176+
The example's `cache_write` price is for a 5-minute cache. A 1-hour cache costs 2 times
177+
the input price instead of 1.25 times. The routing log does not say which cache the
178+
gateway used. A gateway may also charge its own prices, so the result is an estimate, not
179+
the gateway's bill. This command prints one estimated cost per record, or a warning for a
180+
model that has no entry in `prices.json`:
181+
182+
```bash
183+
jq -r --slurpfile prices prices.json '
184+
. as $r
185+
| ($prices[0][$r.model // ""]) as $p
186+
| if $p == null then
187+
"warning: no price for model \($r.model); add it to prices.json"
188+
else
189+
((($r.prompt_tokens // 0) - ($r.cached_tokens // 0) - ($r.cache_creation_tokens // 0)) * $p.input
190+
+ ($r.cached_tokens // 0) * $p.cache_read
191+
+ ($r.cache_creation_tokens // 0) * $p.cache_write
192+
+ ($r.completion_tokens // 0) * $p.output) / 1000000
193+
| "\($r.route_id) \($r.model) estimated $\(. * 1000000 | round / 1000000)"
194+
end' PATH
195+
```
196+
197+
### Thinking
198+
199+
With `reasoning: true` on the model entry, pi sends a thinking level on every request,
200+
even when you do not pass `--thinking`. Switchyard passes that level to an OpenAI-format
201+
target in OpenAI form: `reasoning_effort` on `openai_chat` and `reasoning.effort` on
202+
`openai_responses`. The tested gateway turned either field into Anthropic's older
203+
`thinking: {type: "enabled"}`. Claude Opus 5.5 and Sonnet 5 refuse that form, so the
204+
gateway returned HTTP 400:
205+
206+
```text
207+
"thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.
208+
```
209+
210+
For an `anthropic_messages` target, Switchyard sends the level in the form that Claude
211+
accepts: `thinking: {type: "adaptive"}` with `output_config.effort`. pi's thinking level
212+
then takes effect.
213+
214+
If a Claude target must stay on an OpenAI format, remove the effort field from its
215+
requests with [`omit_body_fields`](../reference/toml_schema.md#targetsname). Claude then
216+
thinks at its default effort, and pi's thinking level has no effect on that target:
217+
218+
```toml
219+
[targets.claude]
220+
id = "claude-opus-5-5"
221+
llm_client = "gateway_chat" # format = "openai_chat"
222+
omit_body_fields = ["reasoning_effort"] # use "reasoning" on openai_responses
223+
```
224+
225+
Switchyard applies the target's `extra_body` and `reasoning_effort` after the removal, so
226+
either can set the field again.
227+
228+
### Forwarded keys
229+
230+
With `api_key_env`, every caller's requests use one gateway key that the server holds.
231+
To use each developer's own gateway key instead, set `forward_auth = true` on the LLM
232+
client. Then set pi's `apiKey` to `$` plus the name of the environment variable that
233+
holds your key: `"apiKey": "$GATEWAY_API_KEY"`. Without the `$`, pi sends the name itself
234+
as the key, and the gateway returns HTTP 401. Only standalone `switchyard-server` forwards
235+
keys. The native Nemo Relay plugin rejects routes that use `forward_auth = true` (see
236+
[Request Handling](nemo_relay.md#request-handling)).
237+
238+
Which request APIs a route accepts depends on its LLM clients that forward the key.
239+
Clients with `api_key_env` do not count:
240+
241+
| Forwarding LLM clients in the route | Request APIs the route accepts |
242+
|---|---|
243+
| None | All three |
244+
| Only OpenAI formats | `/v1/chat/completions` and `/v1/responses` |
245+
| OpenAI formats and `anthropic_messages`, all on the same scheme, host, and port | `/v1/chat/completions` and `/v1/responses` |
246+
| Only `anthropic_messages` | `/v1/messages`, which pi should not use (see [Which request API](#which-request-api)) |
247+
248+
For example, a classifier route can forward pi's key to its GPT judge on
249+
`openai_responses` and to its Claude targets on `anthropic_messages` when both LLM clients
250+
point at the same gateway. If they use different hosts, ports, or schemes, the server does
251+
not start. If a route forwards the key only to Claude targets on `anthropic_messages`, pi
252+
cannot use it. Put those targets on `openai_chat` with `omit_body_fields` (see
253+
[Thinking](#thinking)), or let the server hold the key with `api_key_env`.

‎docs/reference/toml_schema.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -149,6 +149,7 @@ such clients.
149149
| `llm_client` | Yes | — | Key under `[llm_clients]`. |
150150
| `system_prompt` | No | unset | System prompt prepended when this target serves a completion. |
151151
| `extra_body` | No | `{}` | Values merged into the upstream request when the request does not already set that key. |
152+
| `omit_body_fields` | No | `[]` | Top-level fields removed from every request body that Switchyard sends to this target. Switchyard removes them after it translates the request to the LLM client's `format`, so use that format's field names, for example `reasoning_effort` on `openai_chat` or `reasoning` on `openai_responses`. Switchyard applies `extra_body` and `reasoning_effort` after the removal, so either can set a removed field again. |
152153
| `reasoning_effort` | No | unset | Reasoning effort forced on every request to this target, replacing the value the caller sent (`reasoning.effort` on `openai_responses`, `reasoning_effort` on `openai_chat`). Rejected on `anthropic_messages` clients. Use it to run one target at a different effort than the client asked for, for example a strong tier at `max` behind a client that sends `high`. Targets with different effort settings need distinct model IDs when used within one route. Separate routes may use the same model ID with separate `llm_clients` entries (same endpoint, different name). |
153154

154155
Within one route, callable targets with the same model ID must use the same `llm_client`.

0 commit comments

Comments
 (0)