Skip to content

[Bug]: chunked compaction deadlocks on thinking-model summaries and unnumbered provider overflows / 分段压缩在思考型模型摘要与无数字超窗错误上死锁 #9878

Description

@Linearl

[Bug]: chunked compaction deadlocks on thinking-model summaries and unnumbered provider overflows / 分段压缩在思考型模型摘要与无数字超窗错误上死锁

TL;DR(中文摘要)

问题:超长会话(实测 2M tokens)的 pressure 触发压缩会进入无法收敛的失败循环,最终 turn 以 HTTP 400 超窗死亡。两条独立根因:① summarize 采集器丢弃 ChunkReasoning,思考型模型(DeepSeek vision SKUs)把整段摘要放进 reasoning_content、content 为空 → summarizer returned empty output,chunked 分块回退在 fragment N/M 上重复同一失败;② ParseContextLimitError 的正则族不识别无数字的超窗错误(智谱 GLM 1261 "Prompt exceeds max length"),AsContextLimitError 返回 nil → chunked 回退不触发 → uncapped 请求每次透传失败。

影响面:任何使用 thinking 模型(DeepSeek vision 系)或 GLM 网关的会话,一旦上下文超过单次摘要容量即进入死循环;#9572 un-deadlock follow-up 的"失败走 overflow recovery"前提被 ② 破坏。

期望行为:① 纯 reasoning-only 摘要应被 surface(对齐 #9679 对 boundedllm 的处理);② 无数字的 provider 确认超窗应被 trust 为 ContextLimitError(窗口未知),让 chunked 回退接管。

修复:PR 附上(两个定点修复 + 回归测试,已在 fork 实证)。

Background

Observed on a real 2,032,885-token session (deepseek/deepseek-v4-flash-vision-exp, mode yolo, 86 turns over 8 days). The pressure-triggered maintenance sequence was: prune OK (2,032,885 → 1,960,242) → compaction_started → compaction_done → context_maintenance failed, action=summary → turn_done failed with HTTP 400: Prompt exceeds max length. Two distinct failures, captured at different times on the same session:

Failure A — thinking-model summary is misread as empty (manual summary receipt failed-summary-1, reason context summary failed: fragment 2/14: summarizer returned empty output):

The single-shot summary fails, the chunked fallback (#9082 family) correctly kicks in and streams "compacting fragment N/M", but summarize() accumulates only ChunkText. DeepSeek vision thinking models put the whole answer in reasoning_content with an empty content block, so every fragment hits the same empty-output check and the loop never converges. This is the exact shape boundedllm learned to surface in #9679 (goal evaluator) — the compaction summarizer has the same blind spot, but the fix was never extended here. Note the intent contrast: the goal evaluator's verdict must be surfaced; a summary turn that also attempted tool calls legitimately keeps the rejection (that reasoning is private chain-of-thought, not digest material).

Failure B — unnumbered provider overflow is invisible to the chunked fallback (automatic receipt failed-summary-14, reason context summary failed: glm-cn: status 400: {"error":{"code":"1261","message":"Prompt exceeds max length"}}):

ParseContextLimitError (internal/provider/context_limit.go) matches three numeric shapes (OpenAI-style "maximum context length is N tokens...", Anthropic-style "prompt is too long: N > M", "input length and max_tokens exceed...") and a JSON field family (context_length, prompt_tokens, ...). GLM's overflow carries no numbers at all — contextLimitInvariant requires positiveToken(window) and rejects it, so AsContextLimitError returns nil. Consequences:

  1. The chunked fallback condition (errors.Is(err, errSummaryOutputTruncated) || errors.Is(err, ErrCompactionRequired) || provider.AsContextLimitError(err) != nil) does not trigger — the oversized request fails transparently on every pressure retry.
  2. The [Bug]: Compaction cannot rebuild a folded projection once invalidated — uncapped summary requests overflow shared-window models / 投影失效后压缩无法重建折叠投影:无安全前缀上限的摘要请求必然溢出共享窗口模型 #9572 un-deadlock follow-up deliberately keeps an uncapped fold when no balanced prefix remains, relying on "if the provider rejects the uncapped request, the existing overflow recovery takes over". That handoff requires the rejection to be recognized as a context-limit error; with GLM it never is, so the designed recovery path is dead on arrival for this provider.

Fix (attached PR)

Two minimal fixes + regression tests (also validated against a live 2M session):

  1. internal/agent/compact.go — accumulate ChunkReasoning alongside ChunkText; when text is empty and no tool call was attempted, surface the reasoning (clamped to 32 KiB, matching the summaryOutputMaxTokens envelope). A reasoning+tool-call turn still gets summarizer returned empty output, preserving the existing shape contract in TestSummaryCollectorRejectsEmptyAndLengthLimitedOutput.
  2. internal/provider/context_limit.go — isUnnumberedPromptTooLong(message, body) matches the GLM shape case-insensitively; ParseContextLimitError returns a trusted ContextLimitError with zero token fields. Consumers already treat window=0 as "learn nothing, fall back to the configured window" (learnContextBudget only stores window > 0), and the chunked fallback now triggers.

总结

Refs

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agentCore agent loop (internal/agent, internal/control)

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions