You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Observed on a real 2,032,885-token session (deepseek/deepseek-v4-flash-vision-exp, mode yolo, 86 turns over 8 days). The pressure-triggered maintenance sequence was: prune OK (2,032,885 → 1,960,242) → compaction_started → compaction_done → context_maintenance failed, action=summary → turn_done failed with HTTP 400: Prompt exceeds max length. Two distinct failures, captured at different times on the same session:
Failure A — thinking-model summary is misread as empty (manual summary receipt failed-summary-1, reason context summary failed: fragment 2/14: summarizer returned empty output):
The single-shot summary fails, the chunked fallback (#9082 family) correctly kicks in and streams "compacting fragment N/M", but summarize() accumulates only ChunkText. DeepSeek vision thinking models put the whole answer in reasoning_content with an empty content block, so every fragment hits the same empty-output check and the loop never converges. This is the exact shape boundedllm learned to surface in #9679 (goal evaluator) — the compaction summarizer has the same blind spot, but the fix was never extended here. Note the intent contrast: the goal evaluator's verdict must be surfaced; a summary turn that also attempted tool calls legitimately keeps the rejection (that reasoning is private chain-of-thought, not digest material).
Failure B — unnumbered provider overflow is invisible to the chunked fallback (automatic receipt failed-summary-14, reason context summary failed: glm-cn: status 400: {"error":{"code":"1261","message":"Prompt exceeds max length"}}):
ParseContextLimitError (internal/provider/context_limit.go) matches three numeric shapes (OpenAI-style "maximum context length is N tokens...", Anthropic-style "prompt is too long: N > M", "input length and max_tokens exceed...") and a JSON field family (context_length, prompt_tokens, ...). GLM's overflow carries no numbers at all — contextLimitInvariant requires positiveToken(window) and rejects it, so AsContextLimitError returns nil. Consequences:
The chunked fallback condition (errors.Is(err, errSummaryOutputTruncated) || errors.Is(err, ErrCompactionRequired) || provider.AsContextLimitError(err) != nil) does not trigger — the oversized request fails transparently on every pressure retry.
Two minimal fixes + regression tests (also validated against a live 2M session):
internal/agent/compact.go — accumulate ChunkReasoning alongside ChunkText; when text is empty and no tool call was attempted, surface the reasoning (clamped to 32 KiB, matching the summaryOutputMaxTokens envelope). A reasoning+tool-call turn still gets summarizer returned empty output, preserving the existing shape contract in TestSummaryCollectorRejectsEmptyAndLengthLimitedOutput.
internal/provider/context_limit.go — isUnnumberedPromptTooLong(message, body) matches the GLM shape case-insensitively; ParseContextLimitError returns a trusted ContextLimitError with zero token fields. Consumers already treat window=0 as "learn nothing, fall back to the configured window" (learnContextBudget only stores window > 0), and the chunked fallback now triggers.
[Bug]: chunked compaction deadlocks on thinking-model summaries and unnumbered provider overflows / 分段压缩在思考型模型摘要与无数字超窗错误上死锁
TL;DR(中文摘要)
问题:超长会话(实测 2M tokens)的 pressure 触发压缩会进入无法收敛的失败循环,最终 turn 以 HTTP 400 超窗死亡。两条独立根因:① summarize 采集器丢弃
ChunkReasoning,思考型模型(DeepSeek vision SKUs)把整段摘要放进 reasoning_content、content 为空 →summarizer returned empty output,chunked 分块回退在 fragment N/M 上重复同一失败;②ParseContextLimitError的正则族不识别无数字的超窗错误(智谱 GLM1261 "Prompt exceeds max length"),AsContextLimitError返回 nil → chunked 回退不触发 → uncapped 请求每次透传失败。影响面:任何使用 thinking 模型(DeepSeek vision 系)或 GLM 网关的会话,一旦上下文超过单次摘要容量即进入死循环;#9572 un-deadlock follow-up 的"失败走 overflow recovery"前提被 ② 破坏。
期望行为:① 纯 reasoning-only 摘要应被 surface(对齐 #9679 对 boundedllm 的处理);② 无数字的 provider 确认超窗应被 trust 为 ContextLimitError(窗口未知),让 chunked 回退接管。
修复:PR 附上(两个定点修复 + 回归测试,已在 fork 实证)。
Background
Observed on a real 2,032,885-token session (
deepseek/deepseek-v4-flash-vision-exp, mode yolo, 86 turns over 8 days). The pressure-triggered maintenance sequence was: prune OK (2,032,885 → 1,960,242) → compaction_started → compaction_done →context_maintenance failed, action=summary→ turn_done failed withHTTP 400: Prompt exceeds max length. Two distinct failures, captured at different times on the same session:Failure A — thinking-model summary is misread as empty (manual summary receipt
failed-summary-1, reasoncontext summary failed: fragment 2/14: summarizer returned empty output):The single-shot summary fails, the chunked fallback (#9082 family) correctly kicks in and streams "compacting fragment N/M", but
summarize()accumulates onlyChunkText. DeepSeek vision thinking models put the whole answer inreasoning_contentwith an empty content block, so every fragment hits the same empty-output check and the loop never converges. This is the exact shape boundedllm learned to surface in #9679 (goal evaluator) — the compaction summarizer has the same blind spot, but the fix was never extended here. Note the intent contrast: the goal evaluator's verdict must be surfaced; a summary turn that also attempted tool calls legitimately keeps the rejection (that reasoning is private chain-of-thought, not digest material).Failure B — unnumbered provider overflow is invisible to the chunked fallback (automatic receipt
failed-summary-14, reasoncontext summary failed: glm-cn: status 400: {"error":{"code":"1261","message":"Prompt exceeds max length"}}):ParseContextLimitError(internal/provider/context_limit.go) matches three numeric shapes (OpenAI-style "maximum context length is N tokens...", Anthropic-style "prompt is too long: N > M", "input length and max_tokens exceed...") and a JSON field family (context_length,prompt_tokens, ...). GLM's overflow carries no numbers at all —contextLimitInvariantrequirespositiveToken(window)and rejects it, soAsContextLimitErrorreturns nil. Consequences:errors.Is(err, errSummaryOutputTruncated) || errors.Is(err, ErrCompactionRequired) || provider.AsContextLimitError(err) != nil) does not trigger — the oversized request fails transparently on every pressure retry.Fix (attached PR)
Two minimal fixes + regression tests (also validated against a live 2M session):
internal/agent/compact.go— accumulateChunkReasoningalongsideChunkText; when text is empty and no tool call was attempted, surface the reasoning (clamped to 32 KiB, matching thesummaryOutputMaxTokensenvelope). A reasoning+tool-call turn still getssummarizer returned empty output, preserving the existing shape contract inTestSummaryCollectorRejectsEmptyAndLengthLimitedOutput.internal/provider/context_limit.go—isUnnumberedPromptTooLong(message, body)matches the GLM shape case-insensitively;ParseContextLimitErrorreturns a trustedContextLimitErrorwith zero token fields. Consumers already treat window=0 as "learn nothing, fall back to the configured window" (learnContextBudgetonly storeswindow > 0), and the chunked fallback now triggers.总结
Refs