Local-first session postmortem and improvement engine for memoryful AI agent frameworks — agents that have their own persistent identity, memory, skills, and SOP files. Transcript tooling still supports Hermes, OpenClaw, and Claude Code; the v1 desktop diagnosis loop is OpenClaw desktop only. Same shape works for any framework that records sessions as JSONL and stores its own configuration files.
Turn frustrating agent sessions into durable fixes. Read JSONL transcripts → detect failure patterns deterministically → aggregate into one finding per session → stage reviewable patches for memory, SOP, identity, tool discipline, and evals. For productized agent deployments, let the host agent run agent-doctor setup autopilot so diagnosis triggers automatically from user frustration, insults/profanity, trust-break language, hidden tool failures, and unverified completion claims. The same local signals power Agent Doctor, a small doctor-shaped desktop surface that can be summoned by the user or woken by autopilot. No network calls in the production path. No automatic edits to your agent config.
Agent Doctor is an engineering diagnosis tool. It is not therapy, HR performance management, or surveillance analytics, and it is not aimed at chat clients without their own memory or identity surface (Claude Desktop, Cursor, Cline, ChatGPT, …) — those have nothing for apply to patch.
One line. Detects pipx vs pip, installs Agent Doctor, writes skills into every detected memoryful agent framework on the machine, and invalidates host skill caches so the new skill is live on the next session — no manual restart needed where supported.
curl -fsSL https://raw.githubusercontent.com/hesong12/agent-doctor/main/install.sh | shWith extras:
# Include the MCP stdio server
curl -fsSL https://raw.githubusercontent.com/hesong12/agent-doctor/main/install.sh | sh -s -- --with-mcp
# MCP + LLM extras
curl -fsSL https://raw.githubusercontent.com/hesong12/agent-doctor/main/install.sh | sh -s -- --with-all
# Extras plus always-on autopilot sidecars
curl -fsSL https://raw.githubusercontent.com/hesong12/agent-doctor/main/install.sh | sh -s -- --with-all --with-autopilotAfter install, just say to your AI agent: "review my last session" / "diagnose this transcript" / "why does the agent keep doing X". The host's skill router will match against the SKILL.md we wrote into each detected memoryful framework's skill directory and load Agent Doctor's workflow.
For AI-agent-managed deployments, the agent should use the opinionated setup command instead of asking the user to configure paths or service managers:
agent-doctor setup autopilotThat command detects local transcript-capable hosts, installs or refreshes Agent Doctor skills, baselines existing transcripts, writes launchd/systemd user services, starts the sidecars, and installs the desktop Agent Doctor service as the default v1 user-facing surface. The OpenClaw desktop Doctor is the product loop in v1; dashboard, mobile, channel-native, Hermes delivery, reaction approval, and auto-apply flows are out of scope. It does not edit host runtime configuration.
For always-on deployments where the user should not have to remember to ask for diagnosis, run the sidecar:
# One-shot check using each platform's default transcript path
agent-doctor autopilot --platform openclaw --out ~/.agent-doctor/openclaw
agent-doctor autopilot --platform hermes --out ~/.agent-doctor/hermes
# Long-running sidecar mode. Use launchd on macOS or systemd on Linux.
agent-doctor autopilot --platform openclaw --out ~/.agent-doctor/openclaw --watch
# Install as a user service without changing OpenClaw/Hermes.
agent-doctor service install --platform openclaw --out ~/.agent-doctor/openclaw --startautopilot is outside-in: it reads existing transcript/log JSONL, keeps its own SQLite state for cooldown/de-duplication, and writes short diagnosis cards plus events.jsonl under --out. It does not require OpenClaw/Hermes runtime hooks, source changes, or platform cooperation.
If you'd rather not run a remote shell script, the manual package install path is:
pipx install git+https://github.com/hesong12/agent-doctor.git
agent-doctor bootstrap --invalidate-cache
# Optional always-on sidecars after manual install:
agent-doctor setup autopilot(On Ubuntu 24.04+ / Debian 12+ install pipx first: sudo apt-get install -y pipx && pipx ensurepath. The install.sh script handles this automatically.)
bootstrap auto-detects ~/.hermes, ~/.openclaw, and ~/.claude/skills and writes a SKILL.md with YAML frontmatter into the right place for each host. Hermes gets ~/.hermes/skills/autonomous-ai-agents/agent-doctor/SKILL.md (the categorized layout 124+ existing Hermes skills use); OpenClaw and Claude Code get ~/.<host>/skills/agent-doctor/SKILL.md. It also prints the MCP configuration snippet you can paste into any MCP-aware host.
For local development:
git clone https://github.com/hesong12/agent-doctor.git
cd agent-doctor
python3 -m pip install -e ".[dev]"
python3 -m pytest -qThe package exposes the agent-doctor console script and can also be run with python3 -m agent_doctor.cli.
Read this section if you are an AI agent inside a memoryful framework (Hermes, OpenClaw, Claude Code, …) deciding whether to invoke Agent Doctor.
When to invoke:
| User intent | Command |
|---|---|
| "review / diagnose / postmortem my last session" | agent-doctor scan --hermes --format markdown --out ./postmortem (or --openclaw, or --path <jsonl-or-dir>) |
| "why does the agent keep doing X" | same scan; look for findings with count >= 3 and severity high |
| "fix the patterns you found" | agent-doctor apply --findings ./postmortem --out ./staging --target <live-config-dir> |
| "enable proactive diagnosis / install Agent Doctor autopilot" | agent-doctor setup autopilot |
| "help right now / user is angry in this turn" | agent-doctor pet --message "<current user message>" |
| "is the detector accurate / measure improvement" | agent-doctor eval generate → eval bench → eval replay |
Operating rules (these mirror the SKILL.md bootstrap installs):
- Local-only.
scan,apply,bootstrap, andmcp servemake no network calls. The only commands that contact a remote LLM areeval generate --llmandeval replay, both gated onANTHROPIC_API_KEYand the[llm]extra. - Treat patch output as dry-run.
applywrites to a staging directory; live host-agent config is never modified. Always ask the user before copying staged patches into memory / identity / SOPs / skills / permissions / routing / evals. - Autopilot setup is reversible host-side setup, not a host runtime edit.
setup autopilotmay install Agent Doctor skills, local state, the desktop Agent Doctor service, and user-level launchd/systemd services. It must not edit OpenClaw/Hermes runtime config. - Never paste full transcripts to a remote LLM unless the user explicitly approves that disclosure.
- Cite evidence. Findings include file paths, line numbers, role, and quoted excerpts. Prefer those over broad claims about the user or the agent.
MCP-native invocation (if the host speaks MCP and the [mcp] extra is installed):
The server exposes scan, list_findings, read_finding, bench, stage_patches, generate_corpus, doctor_pet_status, and doctor_pet_intervene. All write tools restrict writes to caller-supplied staging_dir / out_dir. No tool calls a remote LLM. See MCP server below.
One-shot golden flow an agent can run end-to-end:
agent-doctor scan --path ./sessions --format markdown --out ./postmortem
agent-doctor apply --findings ./postmortem --out ./staging --target ~/.hermes/skills --min-severity medium
# Then summarize ./staging/sop.md, ./staging/memory.md, and ./staging/DIFF.txt
# back to the user and ask which sections to copy into the live config.agent-doctor scan --path ./sessions --format markdown --out ./agent-doctor-report
agent-doctor scan --hermes --format json --out ./agent-doctor-hermes
agent-doctor scan --openclaw --format markdown --out ./agent-doctor-openclawscan writes three files:
report.md— human-readable summary with redacted evidence quotes.findings.json— machine-readable findings.eval-cases.yaml— starter eval cases.
Findings are aggregated per (failure_mode, session_id). A session with 20 user complaints of the same kind becomes one high-severity finding with all 20 evidence quotes attached, not 20 separate medium findings. Severity escalates by count.
Add --strict to fail on malformed JSONL lines instead of silently skipping them; the default behavior surfaces the skipped count in the summary.
agent-doctor apply --findings ./agent-doctor-report --out ./staging
agent-doctor apply --findings ./agent-doctor-report --out ./staging --target ~/.hermes/skillsapply reads findings.json and writes a staging directory:
staging/
memory.md # one block per memory candidate
sop.md # SOP guidance, grouped by failure mode
identity.md # identity / communication-style guidance
tool-discipline.md # tool-discipline guards
eval/<id>.yaml # one starter eval case per finding
MANIFEST.json # finding → emitted file mapping
DIFF.txt # unified diff vs --target (if given)
Nothing is applied automatically. Real Hermes / OpenClaw configuration is never touched. The staging directory is a curated copy-paste source plus a unified diff against your live config so you can see exactly what would change. Use --min-severity and --min-count to filter noise.
agent-doctor doctor # environment + privacy info
agent-doctor pet --message "Why are you so dumb?" # manually summon Agent Doctor
agent-doctor pet --path ./sessions --out ./doctor-pet # write pet-status.json/card
agent-doctor pet-display --status-file ./doctor-pet/pet-status.json
agent-doctor setup autopilot # install skills, sidecars, and desktop Agent Doctor
agent-doctor autopilot --platform openclaw --out ~/.agent-doctor/openclaw
agent-doctor autopilot --platform hermes --out ~/.agent-doctor/hermes --watch
agent-doctor service install --platform openclaw --out ~/.agent-doctor/openclaw --start
agent-doctor bootstrap # auto-detect hosts and install skills into each
agent-doctor bootstrap --dry-run # preview without writing
agent-doctor bootstrap --target claude-code --force # force into a specific host
agent-doctor install-skill --target hermes --out ./skills # write a single skill file by hand
agent-doctor mcp # print MCP server metadata + tool list
pip install 'agent-doctor[mcp]' && agent-doctor mcp serve # run the stdio MCP serverSupported --target values: hermes, openclaw, claude-code, generic. Hermes, OpenClaw, and Claude Code get a SKILL.md with YAML frontmatter in the host's expected skill location; generic writes a flat Markdown SOP file.
autopilot is the no-runtime-modification product path. It is intended to run as a local daemon, not as a dashboard and not as a cron-only batch job:
agent-doctor autopilot --platform openclaw --out ~/.agent-doctor/openclaw --watch --interval 2
agent-doctor autopilot --platform hermes --out ~/.agent-doctor/hermes --watch --interval 2
agent-doctor autopilot --platform generic --path ./sessions --out ./doctor-autopilotThe default desktop setup uses short polling because Agent Doctor must notice live OpenClaw transcript changes quickly enough to be useful in the same user flow. The watch daemon scans only changed JSONL files after its first baseline, and the desktop pet reloads a small status JSON file; this keeps the 2-second daemon interval and 1-second pet refresh practical for local desktop use. If a host later exposes file-system events or a native session-change stream, that should replace polling for lower CPU and I/O overhead.
For an AI agent configuring Agent Doctor on the user's behalf, prefer:
agent-doctor setup autopilotThis is the zero-touch setup flow: it detects OpenClaw/Hermes, runs bootstrap,
best-effort invalidates host skill caches, installs launchd/systemd user
services, baselines existing transcripts, starts the services, and installs the
desktop Agent Doctor service. Sidecars write host-local artifacts plus shared Agent Doctor
status under ~/.agent-doctor/pet; they do not send system notifications or
inject host messages by default. Use --dry-run to preview, --platform openclaw / --platform hermes to limit scope, --no-start to only write
service files, --no-desktop-pet to skip the desktop surface, and --force
when provisioning a host home before the platform has created its root
directory. --notify-command <cmd> remains available as an explicit legacy hook.
Current automatic triggers:
- user frustration signals, including direct complaints like "not useful", "no value", "not thinking", direct dumb-feedback such as "Why are you so dumb?", "Are you stupid?", direct insults/profanity, trust-break language, trust-degradation phrases like "you are getting worse" / "越来越笨" / "你最近怎么越来越笨了", and common Chinese equivalents such as "不够聪明", "没价值", "废物", "每次都这样", and "你怎么这么笨的?".
- trust-degradation episodes: when two or more trust-eroding signals cluster in one session within a small window, the autopilot emits a high-confidence episode event with an explicit "Required Acknowledgement" section in both the diagnosis card and inbox advisory.
- assistant completion claims without nearby verification evidence (tracked also as the
unsupported_completion_claimfailure mode). - hidden or unacknowledged tool failures surfaced by the deterministic detectors.
High-severity frustration emits an intervene event. Agent Doctor renders that as a visible desktop intervention with a small panel, evidence summary, and recovery action instead of relying on passive OS notifications.
Intervention cards tell the host agent to pause the normal success path, identify the concrete failure, cite evidence, and provide a short corrective action instead of defending itself or writing a long apology. Trust-degradation phrases are also escalated to intervene because they describe cumulative quality loss, not a one-off complaint.
Each emitted high-severity user-frustration / trust-degradation event also appends a regression entry to <out>/regressions/frustration-regressions.jsonl, pinning the exact phrase that tripped the detector. This is the missed-phrase regression library the bench harness can replay against future detector edits to ensure phrases like 越来越笨 are never silently lost.
Agent Doctor is the user-facing state model for those intervention moments. It is intentionally local and small: a doctor persona, a state (idle, watching, concerned, or intervening), redacted evidence, and 2-3 action options. It is the default desktop entry point after agent-doctor setup autopilot, and it is also usable through CLI and MCP:
agent-doctor pet --message "This is useless. You keep doing this." --format markdown
agent-doctor pet --path ./sessions --format json --out ./doctor-pet
agent-doctor pet --message "This is useless." --display
agent-doctor pet-display --status-file ./doctor-pet/pet-status.jsonManual summon (--message) is for the current turn. Transcript mode (--path, --hermes, or --openclaw) uses the same ingestion, detectors, and autopilot event selection as the sidecar. Optional artifacts are written as pet-status.json and pet-card.md under --out with 0600 permissions and redacted transcript strings.
In autopilot mode, Agent Doctor is always displayable by default: every sidecar pass writes the current pet-status.json and pet-card.md under the autopilot --out directory and, when setup installed the desktop service, also refreshes shared status under ~/.agent-doctor/pet. The desktop surface uses a packaged chibi doctor sprite with state-specific motion: idle breathing, watching scan, concerned diagnostic pulse, and intervening alert. Drag it to move it, and click it to open the single status/action panel. Right-click (Control-click on macOS) the pet to pick "Change sprite…" and replace the doctor with any image you like.
You can replace the desktop pet image with your own:
pip install agent-doctor[sprite] # one-time: install Pillow
agent-doctor pet-set-sprite ~/Downloads/my-cat.jpg # crop, resize, transparent bg
agent-doctor pet-set-sprite ~/Downloads/my-cat.png --no-bg-removal # skip floodfillThe pipeline center-square crops the source, resizes it to 512×512 with LANCZOS, and runs a corner floodfill (thresh=28, soft 0.6px alpha blur) to drop a uniform/cream background. The result is written atomically to ~/.agent-doctor/pet/sprite.png and the running pet hot-reloads within ~2 seconds — no kill/relaunch. The packaged default doctor sprite is left untouched; delete ~/.agent-doctor/pet/sprite.png to revert. Pillow stays an optional extra; the base install does not require it. Healthy idle is passive: it has no setup/start button and no user action requirement. The panel keeps explicit user controls in one place. For actionable incidents, v1 exposes Dismiss and, when the incident is routable, Tell Current Agent. Tell Current Agent attempts to inject a structured intervention payload into the current OpenClaw system-event stream; if routing or delivery is unavailable, the panel shows a degraded/failure result instead of pretending success. Agent Doctor never auto-applies config, SOP, or memory changes in v1. The desktop service is not KeepAlive, so stopping the service keeps it closed until the next login or explicit service start.
Watch mode automatically runs a full first pass, then switches to changed-file
scanning using JSONL path, mtime, and size state in SQLite. With
--changed-only, OpenClaw first scans are bounded to the most recent
ordinary session JSONL files and then snapshot the rest, so live monitoring does
not replay the whole historical transcript directory on startup.
Artifacts:
~/.agent-doctor/openclaw/
state.sqlite3 # local de-dupe / cooldown state
events.jsonl # machine-readable emitted interventions
latest.md # most recent short diagnosis card
pet-status.json # always-present Agent Doctor state for desktop/UI shells
pet-card.md # always-present human-readable Agent Doctor card
cards/<event>.md # one card per emitted event
regressions/frustration-regressions.jsonl # missed-phrase regression library
This is the Agent Doctor "self-healing layer" boundary: observe from the outside, diagnose locally, update the desktop state, and stage durable fixes. It does not block host runtime execution or patch live configuration.
Legacy delivery options stay outside the host runtime and are explicit opt-ins. They are not used by setup autopilot defaults:
agent-doctor autopilot --platform openclaw --out ~/.agent-doctor/openclaw \
--inbox-dir ~/.agent-doctor/inbox/openclaw \
--notify-command "/usr/local/bin/send-agent-doctor-card"--inbox-dirwrites a per-session advisory file that a memoryful agent can read on its next turn or heartbeat.--notify-commandruns a local command after a card is emitted. Metadata is passed throughAGENT_DOCTOR_*environment variables such asAGENT_DOCTOR_CARD,AGENT_DOCTOR_TRIGGER,AGENT_DOCTOR_ACTION,AGENT_DOCTOR_SEVERITY, andAGENT_DOCTOR_SESSION_ID.- Delivery failures are recorded in
delivery-errors.jsonl; diagnosis itself still succeeds, but failed interventions are not marked handled in SQLite, so the next watch pass can retry instead of hiding the recovery moment behind cooldown. agent-doctor notify openclaw-system-eventis the legacy OpenClaw delivery adapter. It reads the sameAGENT_DOCTOR_*environment, skips non-interveneevents by default, resolves OpenClaw from host command paths such as/opt/homebrew/binunder launchd/systemd, and enqueues a local OpenClaw system event without changing OpenClaw configuration.
Install as a background user service:
# macOS: writes ~/Library/LaunchAgents/com.agentdoctor.openclaw.plist
# Linux: writes ~/.config/systemd/user/agent-doctor-openclaw.service
agent-doctor service install --platform openclaw --out ~/.agent-doctor/openclaw \
--startService installation baselines existing transcript files before starting by
default and starts the service with changed-file scanning enabled, so a fresh
sidecar does not surface historical findings through Agent Doctor. Pass
--no-baseline-existing when you intentionally want the service to scan old
sessions as soon as it starts. Generated launchd/systemd services also include
a host command PATH (/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin)
so optional legacy hooks do not depend on an interactive shell profile.
Agent Doctor's comfort copy uses the user's configured OpenClaw text model by
default. Set AGENT_DOCTOR_COMFORT_MODEL=<provider/model> before installing or
starting the service only when you intentionally want a dedicated model for the
desktop comfort surface.
The installer also supports this as an opt-in:
curl -fsSL https://raw.githubusercontent.com/hesong12/agent-doctor/main/install.sh | sh -s -- --with-autopilot# 1. Generate synthetic transcripts with ground-truth labels.
agent-doctor eval generate --cards tests/fixtures/cards --out ./corpus
# 2. Benchmark detector P/R/F1 against the labels.
agent-doctor eval bench --corpus ./corpus --out ./bench \
--gate-precision 0.95 --gate-recall 0.85
# 3. Closed-loop replay: apply patches, re-run user turns through a patched agent.
ANTHROPIC_API_KEY=... agent-doctor eval replay \
--transcript ./sessions/frustrating.jsonl \
--patches ./staging \
--out ./replayThe eval pipeline is deliberately separated from the production scan path so the local-first guarantee is preserved. See docs/evaluation.md for the full framework, including scenario card schema, distractor kinds, the LLM-backed generator, and CI gating.
Voice → optimized-prompt clipboard pipeline. Press a hotkey, speak, paste an LLM-rewritten prompt into any AI app. macOS-first; transcription runs locally via faster-whisper or whisper.cpp.
agent-doctor dictate start [--mode chat|coding|research|raw]
agent-doctor dictate stop # transcribe + enhance + copy to clipboard
agent-doctor dictate toggle # one-shot for hotkey binding
agent-doctor dictate status # JSON: recording? pid, mode, elapsed seconds
agent-doctor dictate cancel # discard the in-flight recording
agent-doctor dictate history # show recent transcripts + final promptsSee docs/dictate.md for backends, modes, LLM configuration, the Karabiner example, and the full flag reference.
# Show catalog + which models are installed
agent-doctor dictate models list
# Download an authorized model (Hugging Face/ggerganov)
agent-doctor dictate models download ggml-large-v3-turbo
# Make it the default for new recordings
agent-doctor dictate models set ggml-large-v3-turbo
# What's currently active?
agent-doctor dictate models current
# Re-verify SHA-256 of every installed model
agent-doctor dictate models doctorThe catalog is allow-listed to huggingface.co/ggerganov/whisper.cpp/resolve/main/. Each download is SHA-256-verified and installed atomically into ~/.agent-doctor/models/whisper/.
# Which providers can we see?
agent-doctor dictate llm probe
# Point at LM Studio with a specific model
agent-doctor dictate llm set --provider lm_studio --model qwen2.5-7b-instruct
# Or Ollama
agent-doctor dictate llm set --provider ollama --model llama3.1:8b
# Custom OpenAI-compatible endpoint
agent-doctor dictate llm set --provider custom --url http://localhost:8000/v1 --model whatever
# Show the active config
agent-doctor dictate llm current
# Round-trip a canned transcript
agent-doctor dictate llm test "rewrite this as a clean prompt"Phase 2 collapses the old --mode chat|coding|research flags into a single "optimize for any downstream LLM" prompt. The flags continue to parse but emit a deprecation warning and behave as --mode optimize.
When you trigger a dictation, the pet sprite shows a pulsing cyan ring (listening) while audio is being captured/transcribed, then three orbiting amber dots (thinking) while the LLM rewrites your transcript. The autopilot-driven state (e.g. intervening) is restored as soon as the pipeline finishes — dictation never clobbers an open critical alert.
Disable per-state via ~/.agent-doctor/dictate.json (pet.animate_listening, pet.animate_thinking).
# One-time install: compiles the Swift helper and registers a LaunchAgent.
agent-doctor dictate hotkey install
# Change the chord.
agent-doctor dictate hotkey set "ctrl+option+space"
# Show status + binding.
agent-doctor dictate hotkey show
# Stop and remove the daemon.
agent-doctor dictate hotkey uninstallThe global hotkey now defaults to Right Command (hold) — the
Handy-style "single modifier hold to talk" trigger. Open
agent-doctor dictate preferences → Hotkey tab to pick a different key
or to switch to a chord (⌃⌥Space-style). Modifier-only bindings are
push-to-talk only; chord bindings can be either push-to-talk or toggle.
The capture overlay live-previews each key event; press for ≥ 400ms and
release to commit a single modifier, or press a chord and click "Use
this chord".
The helper is a ~150 LOC Swift binary compiled with swiftc (requires Xcode Command Line Tools). It reads ~/.agent-doctor/dictate.json on launch and on SIGHUP, so set updates take effect without re-installing.
You will be prompted to grant Input Monitoring in System Settings -> Privacy & Security the first time the daemon registers an NSEvent monitor.
By default agent-doctor copies the optimised prompt to the clipboard and stops there. To have it land at the focused cursor automatically:
# One-time: runs a synthetic Cmd+V test and asks you to grant Accessibility permission.
agent-doctor dictate paste enable
# Re-test after enabling.
agent-doctor dictate paste test
# Stop auto-pasting.
agent-doctor dictate paste disableIf the keystroke fails (e.g. permission revoked), the text is still on the clipboard and a notification fires explaining the fallback.
The primary entry point is the desktop pet: right-click the sprite and choose Preferences…. The same window can also be launched from the terminal:
agent-doctor dictate preferencesNote: the Preferences window uses Tk. If your Python was built without Tcl/Tk (common for
pyenvinstalls), launching from the pet shows a "missing Tk support" alert and the CLI exits withModuleNotFoundError: No module named '_tkinter'. Fix withbrew install python-tk@<version>matching the Python that runsagent-doctor, or rebuild pyenv with Tcl/Tk linked.
Five tabs:
- Dictation — model, language, extra recording buffer.
- LLM — provider (LM Studio / Ollama / Custom), base URL, model, test connection.
- Hotkey — chord, push-to-talk vs toggle, install / uninstall daemon.
- Paste — auto-paste toggle, delay, permission test.
- Pet — listening / thinking animation toggles, sprite picker.
All changes save immediately. No Save button.
Agent Doctor is local-only by design.
- The
scan,apply, andeval generate(without--llm) commands make no network calls and do not call remote LLMs. - Default operation is read-only against agent state and transcript inputs.
applywrites patches into a staging directory only; it never edits your live agent config.- Reports redact common secrets, API keys, bearer tokens, and passwords by default.
- All artifacts are written with
0o600permissions. - The generated host-agent SOP explicitly tells agents not to paste full transcripts to a remote LLM unless the user approves that disclosure.
eval generate --llmandeval replayare the only commands that contact a remote LLM, and only when anANTHROPIC_API_KEYis present and the[llm]extra is installed. They live in theagent_doctor.evalssubpackage to keep the production path dependency-free.
| Failure mode | Signal | Patch targets |
|---|---|---|
repeated_user_correction |
"I already told you", "not what I asked", "X again", "我已经说过" | memory, SOP |
execution_discipline |
promised action without observed tool execution; "don't just plan" | SOP, eval |
verification_failure |
"did you test", "without verifying", "not actually tested" | SOP, eval |
memory_failure |
"you forgot", imperative "remember", "last time", "I told you" | memory |
tool_failure_or_hidden_error |
tool emits error/timeout/401/500/traceback; assistant claims success without acknowledging | SOP, tool discipline |
communication_mismatch |
"too verbose", "stop explaining" | memory (with overfit warning), identity |
user_frustration_signal |
user shows anger, direct insult/profanity, direct quality complaints, repeated correction, trust-degradation phrases ("you are getting worse", "越来越笨", "你最近怎么越来越笨了"), or trust-break language | identity, SOP, eval |
trust_degradation_episode |
meta-finding: 2+ trust-eroding signals cluster in one session within a small window | identity, SOP, eval |
missed_core_question |
"you didn't answer my question", "answer my actual question", "你没回答", "答非所问" | SOP, eval |
instruction_drift |
"I didn't ask you to…", "stop adding extra", "我没让你…" | SOP, eval |
over_process_response |
"stop narrating", "just give me the answer", or long assistant messages dominated by step-by-step planning tokens | identity, SOP |
unsupported_completion_claim |
assistant claims completion with no recent tool action / verification keyword; user pushback escalates severity | SOP, eval |
Distractors that are deliberately not flagged:
- "Just so I remember the timeline ..." (informational
remember). - Tool output containing "0 errors" / "no failures".
- Identifier-like terms such as
error_handler.pyorerror.log. - Technical Chinese terms such as "笨重" (cumbersome/heavy), which are not user frustration.
- Assistant offers like "I can run ... if you want" (capability, not promise).
The bench corpus under tests/fixtures/cards/ includes a distractor-only scenario; CI fails if any of these false-positives reappear.
The MVP ingests JSONL files. It auto-detects Hermes-ish, OpenClaw-ish, and generic event shapes using common fields such as role, actor, speaker, payload, data, entry, content, message, text, output, and error. Nested message.role and message.content values are normalized under common message, data, entry, and payload containers.
Default paths:
- Hermes:
~/.hermes/sessions - OpenClaw:
~/.openclaw/agents/main/sessions
Per-message content is capped at 8,000 characters so a single huge tool stdout cannot dominate downstream detection.
python3 -m pytest -qSmoke commands:
python3 -m agent_doctor.cli doctor
python3 -m agent_doctor.cli scan --path tests/fixtures --out /tmp/agent-doctor-smoke --format markdown
python3 -m agent_doctor.cli apply --findings /tmp/agent-doctor-smoke --out /tmp/agent-doctor-staging
python3 -m agent_doctor.cli eval generate --cards tests/fixtures/cards --out /tmp/ad-corpus
python3 -m agent_doctor.cli eval bench --corpus /tmp/ad-corpus --out /tmp/ad-benchAgent Doctor ships a stdio MCP server that exposes the same diagnosis surface as the CLI. Install the optional extra and let any memoryful MCP-aware agent framework (Hermes, OpenClaw, Claude Code, or any framework that supports MCP tools and has its own memory / identity files) call it mid-session:
pip install 'agent-doctor[mcp]'
agent-doctor mcp serve # stdio server; configure your host with the snippet from `bootstrap`Tools exposed:
| Tool | Reads | Writes |
|---|---|---|
scan |
JSONL transcripts | report artifacts under out_dir |
list_findings |
findings.json |
— |
read_finding |
findings.json |
— |
bench |
corpus dir | bench.json, bench.md under out_dir |
stage_patches |
findings.json (+ optional read-only target_dir) |
staging_dir only |
generate_corpus |
scenario cards | corpus under out_dir |
doctor_pet_status |
JSONL transcripts or current message | optional pet-status.json / pet-card.md under out_dir |
doctor_pet_intervene |
JSONL transcripts or current message | optional pet-status.json / pet-card.md under out_dir |
The trust boundary matches the CLI: write tools never touch live host-agent configuration, only staging_dir / out_dir. No tool calls a remote LLM — the LLM-augmented generator is a CLI-only path on purpose.
agent-doctor mcp (no subcommand) prints the metadata and tool list as JSON.
- Cross-session aggregation for stronger memory candidates.
- LLM-augmented detection layer (opt-in second pass for sarcasm / indirect signals).
- Annotation UI for the real-data half of the golden corpus.
- Judge-LLM recommendation rubric runner (described in
docs/evaluation.md). - Calibrated confidence model and severity from labeled data.