Skip to content

Refactor and clean up unused code, enhance IR and optimizer - #91

Merged
krypticmouse merged 75 commits into
mainfrom
v1-program-unification
Aug 21, 2026
Merged

krypticmouse merged 75 commits into
mainfrom
v1-program-unification

Conversation

@krypticmouse

Copy link
Copy Markdown
Owner

No description provided.

krypticmouse and others added 30 commits August 19, 2026 22:00
Zero callers outside its own test. Removes src/trace/export/,
tests/test_trace_export.rs, and the Otel*/Rl* re-exports from
trace::mod and lib.rs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Deletes Trace::absorb, Trace::redact, Trace::for_component_id, the
reserved SpanEvent::Chunk variant, and the Span.links sequential-
approximation field (plus its capture-side bookkeeping). Updates the
tracing examples and trace tests accordingly; attach_program tests
are kept.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ParetoFrontier had zero callers outside its own test; GEPA uses the
engine's ScoreMatrix/ParetoView directly. ParetoStatistics (the one
type GEPA reports) moves into engine.rs next to ParetoView::statistics,
which is now the single statistics implementation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ctural grammar)

New workspace member crates/dsrs-syntax: the shared syntax layer for the
.dsrs text format. src/lex.rs is dspy-rs/src/ir/text/lex.rs moved verbatim
(pub instead of pub(crate)); src/structure.rs is the include_program!
structural checker rebuilt on that one shared lexer; ParseError (line/col/
message) moves here as the shared error type. Leaf on purpose: serde_json
only — no dspy-rs, dsrs-macros, facet, or rig, so the proc-macro crate can
depend on it without a cycle.

Fixture copies of the golden .dsrs artifacts live in tests/fixtures with
the parity test (checker accepts everything the full parser accepts) —
no more cross-crate relative-path reaching.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tax lexer

Delete src/ir/text/lex.rs (moved verbatim to dsrs-syntax) and re-export
dsrs_syntax::ParseError from ir::text so dspy_rs::ir::ParseError keeps its
path and shape. parse.rs now pulls Lexed/Lexer/Span/Tok from dsrs_syntax::lex
— the full parser and the macro's structural checker can no longer disagree
about tokens.

Golden behavior proven unchanged: Program::from_dsrs/to_dsrs round-trips on
tests/fixtures/*.dsrs are byte-identical to the pre-refactor output and
program hashes match (canonical text is the program-hash preimage).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
EngineCheckpoint, EvalEngine::checkpoint(), and EvalEngine::resume()
had no callers outside their own tests. Drops the feature, its tests,
and the now-stale checkpoint mentions in Spend/RolloutCache docs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…uses

Deletes ResponseCheck, ConstraintOutcome, evaluate_constraints, and the
Constraint::new_check/new_assert constructors (zero callers). Keeps
Constraint, ConstraintKind/ConstraintLevel, and the two expression
evaluators used by the chat adapter and ir/text parser. lib.rs and
typesys re-export lists updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ete the shadow parser

Remove dsrs-macros/src/dsrs_syntax.rs (923 lines) — the hand-maintained
mirror of the .dsrs lexer + structural grammar. include_program! now calls
dsrs_syntax::check from the shared leaf crate, so a grammar change is made
once instead of in two lexers and two parsers. Error strings are unchanged
(the trybuild UI expectations pass as-is), and the macro crate drops its
serde_json dependency, which only the shadow lexer used.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…endencies]

Every dep used by 2+ member crates now has a single pin at the workspace
root (anyhow, async-trait, futures, reqwest, serde, serde_json, tempfile,
thiserror, tokio, plus the rig-core/minijinja git pins and the dspy-rs/
dsrs-tools path deps); members declare 'workspace = true' and add only
their own features. Pins match what Cargo.lock already resolves — verified
with 'cargo metadata --locked' — so nothing was upgraded. The facet
[patch.crates-io] fork pin stays as-is (load-bearing; see TODO to unpin).

Also move rstest from [dependencies] to [dev-dependencies] in dspy-rs:
it is only used by tests, and downstream users should not compile a test
framework. tempfile stays a runtime dep — utils/cache.rs owns a TempDir
for the disk cache tier.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ive/

Both are superseded historical documents; each now carries an archival
note pointing at docs/v1-vision-report.md and docs/rfcs/. Nothing in the
repo referenced their old root paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e cache-wide mutex

Three defects in ResponseCache:

1. Tempdir lifetime: new() created a tempfile::TempDir, pointed
   FsDeviceBuilder at it, and dropped the guard at end of scope —
   deleting the directory the disk tier had just opened, so the disk
   cache only worked through already-open handles, intermittently.
   The TempDir now lives in the struct (Arc<TempDir>, shared by
   clones) for the cache's whole lifetime.

2. Panicking init: the four unwraps in new() could take down the
   process from LM construction. Disk-tier setup now returns a Result;
   on failure we log a tracing warning and fall back to a memory-only
   foyer cache (the storage phase's default noop engine cannot fail).

3. Lock contention: insert_entry took &mut self only to maintain the
   debug history ring, forcing Arc<Mutex<ResponseCache>> in LM so
   every call in the process serialized on one async mutex — with an
   O(n) Vec::insert(0, ..) memmove per insert. The foyer handle is
   internally synchronized, so all methods now take &self; the history
   ring is an isolated Mutex<VecDeque> (push_back/pop_front, cap 100)
   that never blocks cache lookups. LM::cache_handler is now a plain
   Arc<ResponseCache>. get_history awaited nothing and is now sync
   (its only callers were inspect_history and one test).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…race

TRACING_INITIALIZED was checked before set_global_default but only set
after, so two concurrent init_tracing() calls could both pass the check
and one would get a spurious SetGlobalDefault error, defeating the
documented idempotency. The OnceLock is now claimed up front: exactly
one caller wins set(()) and performs initialization; racing losers
return Ok immediately.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The stub unconditionally returned Vec::new(). FieldDef.constraints and
ClassDef.constraints stay: the .dsrs text format reads and writes
field-level constraints (ir/text parse.rs / print.rs).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
dataloader.rs and utils.rs move up one level; the v1 forwarding module
is gone. crate::data::DataLoader and friends resolve unchanged; the
crate::data::v1:: path alias had zero users.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…entifier

A capability becomes a globalThis property, so a tool named JSON produced
globalThis.JSON = <shim> and broke every later capability shim (each depends
on JSON.stringify/parse); reserved words like `class` were installable but
uncallable from script code.

- Add a RESERVED_JS_NAMES blacklist: ECMAScript reserved words/literals plus
  the ambient globals the sandbox depends on (JSON, Object, Promise, Error,
  Array, String, Number, Boolean, Math, Reflect, Proxy, Symbol, globalThis,
  undefined, null, typed arrays, URI/number parsing functions, ...).
- js_identifier now suffixes colliding names with `_tool` (JSON -> JSON_tool),
  following the existing collision-mangling style.
- Capability::validate_name rejects reserved names at install time.
- Document that capability calls are bounded by the remaining sandbox
  deadline (enforced executor-side in the follow-up commit).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s, bound bytecode cache

Four defects from the security/robustness review of the QuickJS sandbox:

1. Capability-name JS injection: the shim was built with format! and eval'd,
   and only add_capability validated names — run_script (Code Mode, dspy-rs
   interpreter) installed caller-supplied capabilities unvalidated, so
   Capability::new("x; globalThis.leak = 1; //", ...) injected code into the
   sandbox bootstrap. install_capabilities (the single choke point every
   sandbox goes through) now re-validates every name, and no name is ever
   spliced into evaluated JS: a constant SHIM_FACTORY is eval'd once and the
   shim is installed via globals.set(name, shim), where the name is only a
   property key. JSON.parse/stringify are captured at bootstrap so sandboxed
   code cannot corrupt later capability calls by clobbering globalThis.JSON.

2. Deadline did not cover capability calls: the engine interrupt handler
   cannot fire while a host future runs, so in Code Mode the deadline bounded
   only JS arithmetic. Each capability handler is now wrapped in
   tokio::time::timeout bounded by the *remaining* sandbox deadline; on
   expiry the sandbox timed_out flag is set (classifies as ExecError::Timeout)
   and a clear, catchable error is thrown into JS.

3. Handle::block_on inside spawn_blocking deadlocked on current_thread
   runtimes. The hook now spawns the handler onto the host runtime and parks
   the sandbox thread on an mpsc channel (recv_timeout with a grace period as
   a loud failure path if the runtime is not being driven). Capabilities now
   work on both runtime flavors.

4. Bytecode cache never evicted: deregister removed the tool but leaked its
   bytecode, and the map was unbounded. deregister now evicts the tool's
   entry unless another registered tool shares the hash, and the cache is
   bounded at 128 entries with a documented cap-and-clear policy (registered
   tools hold their own Arc, so eviction only costs a recompile on miss).

No public API changes: register/execute/tool/run_script signatures untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Capability-name injection and reserved-global names rejected on both the
  builder path and the run_script path (which previously skipped validation).
- A tool named JSON is installed as JSON_tool, leaves globalThis.JSON intact,
  and other capabilities keep working in the same sandbox.
- Capability calls are bounded by the wall-clock deadline: a stalling handler
  surfaces as a typed timeout, both via Executor::execute and end-to-end
  through CodeModeTool.
- Capabilities work on a current_thread runtime (no Handle::block_on).
- deregister evicts bytecode unless another tool shares the source hash, and
  the cache stays within its 128-entry bound while evicted-but-registered
  tools keep executing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…etry race

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The alias existed only to keep pre-Schema spellings compiling after the
vendored BAML integration was removed; it expanded to the identical facet +
serde derives. Update the remaining tests and example 16 to spell it
#[Schema]. Historical mentions in docs/adr and docs/specs are left as
records.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rsers

Chat::from_json (which took &self and ignored it) had no production
callers; Message::from_json_value and parse_content_block existed only
to serve it, including the legacy string-content and type-tagged
formats. Chat::to_json stays. Round-trip/legacy-parse tests that
exercised the deleted API are removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ants

LmError::{Network, RateLimit, InvalidResponse, Timeout} and
ErrorClass::{NotFound, Forbidden} were never produced anywhere — every
provider failure arrives as LmError::Provider. Simplifies class() /
is_retryable() and fixes the stale retryability docs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ReAct exists (modules/react.rs) and the IR program graph ships
default-on behind the ir feature — the docs claimed the opposite.
Adds ir to the crate-organization list.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…drop BamlType alias

Resolved manifest conflict: dsrs-syntax promoted to [workspace.dependencies].

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…dback_helpers, checkpoint, dead typesys/errors, data/v1 flatten)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The structural mutation half of the IR: a typed, validated graph-edit
calculus so optimizers can change program STRUCTURE, not just parameter
values (the harness-optimization enabler).

New `ir::edit`:
- `Edit` (serde values — inspectable, diffable, replayable): AugmentSig
  (the CoT move; copy-on-write SigId), SwapLeaf (Predict <-> AgentLoop,
  context slot minted/dropped), WrapRetry (parent rewire + downstream
  port redirect), Remove (Seq step; dangling refs surface as
  validate.rs's own error), AddTool/RemoveTool (existence + cap ceiling;
  RemoveTool clears stop_tools), SetStop, SetInstructionDefault.
- `Program::edited(&[Edit])` — pure: clone arenas, apply in order,
  GC dead nodes/newly-orphaned sigs/orphaned params, rebuild the
  ParamPath index, validate(), stamp lineage.parent like bake(), seal a
  new content hash. `edited(&[])` is a hash no-op.
- `Program::legal_edits(NodeId) -> Vec<EditKind>` — the structural menu
  for an LLM proposer, per-tool add/remove entries included.
- `migrate_overlay(parent, overlay, child)` — re-mints tuned values by
  ParamPath when the owning leaf's signature still carries (inputs
  identical, outputs may widen), so instruction/demos survive the CoT
  move; ModelRefs re-mint by model name.
- `Program::leaf_id(name)` — stable relocation across edits.

Tests: apply/validate failure modes, GC, text + JSON round trips,
overlay migration (25 new tests; all IR suites green).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…its, migrate_overlay)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ReAct reimplements the tool loop with a string trajectory, substring-based
tool dispatch, and a type-erasing single-blob output; the IR AgentLoopNode
plus the #[agent] macro is its replacement. Delete the module, its builder
tests, and both examples that exercised it (19-react, 93-smoke-slice4).

Map/AndThen/ModuleExt::{map,and_then} wrap transforms in
#[facet(opaque, skip)] closures — invisible to optimizers, pure sugar that
adds two types to the reflection story. module_ext.rs had no other items,
so the module goes with them, along with test_module_ext.rs and the
Map/AndThen shape tests in test_module_facet_shapes.rs.

Also drop forward_all's hardcoded kdam progress bar (trivially separable:
one local bar + one .inspect link) and the now-unused kdam dependency.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
krypticmouse and others added 29 commits August 20, 2026 11:01
Optimizer is object-safe: compile(target: &mut OptimizeTarget, engine:
&mut Engine) -> Report, with compile_module typed sugar. One Engine
replaces the EvalEngine/ProgramEvalEngine split; candidates are name-keyed
Candidates injected ambiently (no mutation during evaluation; single
install of the winner). Deleted API scrubbed: apply_candidate/
CandidateUndo, DynPredictor, checkpoint/resume, ParetoFrontier wrapper,
feedback helpers, MIPRO's PromptCandidate/select_best_traces/
create_prompt_candidates; GEPA's compile_with_valset is now
compile_module_with_valset. Examples declare leaves via predictors!.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
dspy_rs::prelude re-exports the curated core surface: Signature (trait
+ derive), Example derive, Predict/ChainOfThought/Predicted/Module,
Predictors + the predictors! macro, configure/LM/LMConfig, Demo,
DataLoader/TypedLoadOptions, TypedMetric/evaluate_trainset/Eval,
Optimizer + the five strategies + Engine/Candidate/OptimizeTarget,
capture/replay entry points, ir::{Program, Interpreter, Overlay, Edit},
and init_tracing.

Additive: the crate-root glob re-exports are untouched. lib.rs docs now
recommend the prelude, and quickstart examples 01-05 import only
dspy_rs::prelude::* as proof it suffices for the basic path. Also
scrubs stale 'ir feature (default-on)' doc mentions left from the
feature collapse.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ta::v1

Traces drops the OTel/RL export sections (trace/export is deleted); state
documents the Predictors-keyed ModuleState (&M from_module) and the
PredictorInfo::load_state install seam; fx documents clear_instruction/
set_demos/bind/from_overlay and the ambient-candidate role; data drops the
data::v1 versioning story and EvalEngine row; evaluation drops the deleted
feedback_helpers section; signatures notes #[BamlType] alias removal;
tools-and-agents/quickstart/examples stop referencing ReAct and deleted
examples; code-mode documents reserved-name mangling, capability deadline
bounding, and the bounded bytecode cache; runtime adds run_collecting/
LeafOutcome and the dsrs-syntax expansion-time check in include_program!;
utils fixes ResponseCache signatures (&self, sync get_history).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One page for the structural mutation half of the IR: the Edit enum,
Program::edited semantics (positional NodeIds, batch validation, identity-
preserving GC, cot re-sugaring), EditError/ApplyError, the legal_edits
proposer menu, and migrate_overlay. Registered in the Programs as Data nav
group and the index components table; program-and-nodes cross-links it and
drops its stale checkpointing mention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s-syntax

gen_api.py gains the dsrs-syntax companion crate; all api/*.mdx pages are
regenerated from fresh rustdoc JSON at the merged tree (deleted items —
ReAct, ModuleExt, DynPredictor, EvalEngine/ProgramEvalEngine, trace
exports, pareto wrapper, feedback helpers, BamlType — drop out of the
inventory by construction). dsrs-syntax registered under Companion crates
in the nav; README and the API overview page document the fourth rustdoc
command.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The old overview described the V5/V6 milestone: DynPredictor handles, the
facet walker, #[derive(Module)], ReAct, ModuleExt::map, and a planned
ProgramGraph/registry layer. Rewritten against current truth: predictors!
declaration, the ambient-candidate optimizer contract over OptimizeTarget/
Engine, Predict-as-1-node-program, and the IR edit calculus as the concrete
form of structural optimization. dsrs-format.md checked against the
current grammar (reserved list, node forms, ports) — no drift found.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The standalone fn used to build a bare Predict::with_tools predictor,
silently ignoring max_turns, budget, context, stop_tools, and
until_parse — the docs even admitted loop options applied to the
module-lowered form only.

Predict now carries an optional AgentLoopSpec (builder:
with_agent_spec), applied when the 1-node agent program is built:
stop_tools/max_turns/until_parse land in the node's StopSpec, budget in
its NodeBudget, context in its ContextPolicy. The #[agent] standalone
fn feeds its parsed AgentStepOpts through that seam, so both lanes now
execute the same honored configuration.

model = "..." cannot be honored standalone (model refs bind only in a
module program's model table), so it is now a compile error on that
path: with model set no standalone fn is generated, and calling it
fails with 'expected function, found module'. max_turns = 0 is rejected
at macro expansion. New behavioral tests prove stop_tools and max_turns
pass-through standalone; trybuild UI tests pin both compile errors;
macro and component docs updated to drop the admission.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
RFC 0004 collects the deliberate post-unification remainder in one
place: the conversation surface (TODO(dsrs-phase4-conversation)), the
caller-managed tool loop still on the LM-layer path
(TODO(dsrs-phase4-caller-managed)), the retired shared-ptr-policy
marker (moot since the facet walker was deleted), whole-rollout credit
assignment in harvest.rs, tool membership as a ParamSlot (the ToolSet
gene), and structural optimizers over ir::Edit — one paragraph each on
what, why deferred, and the suggested shape.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…de, #[agent] options honored, seams ledger

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…r prelude page

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…r ir::Edit

Closes seam 6 of RFC 0004: the edit calculus was fully shipped but no
optimizer proposed edits. Structural is a GEPA-style hill-climbing loop
over the shared Engine, program lane only:

- gathers Program::legal_edits for every leaf, filtered to the kinds it
  can materialize without free text (AugmentSig as the CoT move,
  SwapToAgent/SwapToPredict, WrapRetry, Remove, AddTool/RemoveTool;
  SetStop and SetInstructionDefault stay value-level work);
- a reflection LM (prompt_model) reads the canonical .dsrs text, the
  serialized menu (one JSON object per line with an option number), and
  the incumbent's per-example feedback, and answers with one option;
  no prompt model, an unparseable reply, or an out-of-range option
  degrades to a seeded-uniform pick;
- applies the choice via Program::edited, migrates the incumbent
  overlay with migrate_overlay, loads the child through a
  caller-supplied RuntimeEnv factory (only the host knows the live
  bindings), and accepts through the engine's minibatch gate: a strict
  win on the shared minibatch promotes to a full-set evaluation and
  the child becomes the new incumbent;
- every rejection path (EditError, LoadError, gate loss) is recorded
  in the step and skipped, never a panic;
- budgets are engine budgets: every child is a fresh program whose
  hash keys fresh rollout-cache rows, so max_rollouts/max_lm_calls are
  the real control surface; the baseline pass seeds the cache so
  parent minibatch reads are free.

Entry points are compile_program / compile_program_with_overlay rather
than the object-safe Optimizer trait: the trait's currency is one
lane-erased target, and a structural child is a new program needing a
new interpreter, which only the caller's RuntimeEnv can load. The
winner is returned (program + migrated overlay) for Program::bake,
never installed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ural-optimizer claims

- docs/optimizers/structural.mdx: how-to page mirroring the gepa.mdx
  shape (overview, quick start, config table, report tables, the loop,
  cost model, troubleshooting), registered in docs.json next to the
  other optimizer pages;
- components/optimizers.mdx: six strategies, comparison row, the
  compile_program exception note, and full Structural config/step/
  report tables;
- components/edit-calculus.mdx: the proposer-menu section now points
  at the shipped optimizer instead of a hypothetical one;
- components/optimizer-engine.mdx and the shared comparison snippet
  pick up the sixth strategy.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…::Edit

Closes RFC 0004 seam 6: reflection over legal_edits menus, Program::edited,
migrate_overlay, minibatch-gated hill climb inside the existing engine.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…004 seam 4)

Demo harvesting scored every span in a rollout with the rollout's single
metric score, so a good final answer vouched for every intermediate
predictor call — including ones a later step recovered from. Close the
seam with span-level evals:

- Span gains an optional eval field (serde default + skip-when-None):
  additive under RFC 0001 §5.1, no format version bump; eval-free traces
  serialize byte-identically to before.
- TypedMetric gains evaluate_spans, a per-trace hook called once per
  traced rollout after evaluate, returning (SpanId, Eval) pairs that
  rollout_traced stamps onto the trace. The default body returns no span
  scores, so every existing metric compiles and behaves unchanged.
- harvest::collect_demo_candidates gates and ranks each span on its
  effective score — the span's own eval when present, the whole-rollout
  score otherwise — so a scored-down span stays out of the demo pool even
  from a winning rollout, and a scored-up span qualifies from a losing
  one. Bootstrap, MIPROv2, and SIMBA inherit this through the shared join.

Tests: harvest unit tests (baseline parity without span evals, override in
both directions, ranking); JSONL round-trip with the field absent and
present; end-to-end draft/refine bootstrap pair showing the whole-rollout
metric harvests recovered-from drafts and a span-aware metric keeps them
out. Docs: evaluation, traces, optimizers, optimizer-engine, miprov2 pages
updated per the docs register.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes RFC 0004 seam 4: TypedMetric::evaluate_spans hook, Span.eval additive
field, harvest prefers a span's own eval over the rollout score.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… seam 5)

Which tools an AgentLoop carries becomes an optimizable slot. Declaration
stays structural (AgentLoopNode::tools is the capability footprint and the
gene's legal alphabet); selection is ParamKind::ToolSet at
"<leaf>.tool_set", defaulting to the full declared table so existing
programs print, hash, and run unchanged.

- params: ParamKind::ToolSet, ParamValue::ToolSet { tools }, ToolSetK kind
  tag, Overlay::set_tool_set; Overlay::set refuses values naming a tool the
  owning agent does not declare (OverlayError::ToolSetUndeclared), so the
  serde load path (from_named) is guarded too.
- graph/builder: AgentLoopNode::tool_set: ParamId; the builder mints the
  slot per agent node (NodeSpec::tool_set seeds a restricted default).
- validate: load-time only — id range checks for ToolSet values, kind/owner
  check for the slot, and subset + duplicate-free checks against the node's
  declared table (ToolSetUndeclared / ToolSetDuplicate).
- interp: eval_agent resolves the gene through the overlay and assembles
  definitions, by_name dispatch, sandbox code, and the Code Mode surface
  from the selection; absent slot = full declared set. A deselected tool
  cannot execute even if the model hallucinates its name.
- text: `tool_set [a b]` agent option; canonical print elides it when the
  default equals the declared list, so pre-ToolSet hashes are unchanged.
- edit: SwapLeaf to agent mints the slot (predict direction collects it);
  AddTool/RemoveTool keep the default in sync with the declaration;
  migrate_overlay re-mints ToolSet entries by tool name and intersects with
  the child's declared table.
- bake folds ToolSet entries into defaults like every other kind; trace
  attach_program now joins the new leaf-owned slot.
- docs: program-and-nodes, dsrs-file, tools-and-agents, edit-calculus.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes RFC 0004 seam 5: ParamKind::ToolSet per agent node, subset-of-declared
validation at load time, interpreter honors the selection, edit/migration and
text-format wiring included.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ams 4-6 as landed-unreviewed

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…erges

Answers the §3 review questions (hash stability, goldens, harvest parity),
records per-seam sharp edges, flags in-flight seams 1+2 and deferred work.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…FC 0004 seams 1-2)

Seam 1, conversation-in/conversation-out:
- Interpreter::run_conversation(chat, input, overlay, budget) -> (RunOutput, Chat):
  one turn with the program's single leaf over a caller-owned Chat. Opening
  turns (empty chat + input) render system+demos+input through the same
  overlay-resolved path as map-in runs, so prompts are byte-identical and
  request_hash matches across paths. A turn is not a run: one span per turn,
  seq increments per turn, replay serves turn by turn.
- Interpreter::conversation_opening(input, overlay) renders the opening Chat
  without calling anything.
- Predict::build_chat / call_and_parse are now thin wrappers over these
  entries. The static prompt-prefix cache and the LM-layer conversation path
  (call_and_parse_with_input, serve_recorded_span) are deleted; build_chat is
  now async. TODO(dsrs-phase4-conversation) markers removed.

Seam 2, caller-managed tool loop as suspension:
- Interpreter::run_conversation_caller_managed runs the same AgentLoop in
  suspending mode: on tool calls it returns
  ConversationTurn::Suspended(ToolSuspension) carrying the pending calls, the
  open span guard, and the loop/run meters; resume_conversation feeds results
  back (same ToolRun events, batched tool-result turn, context-policy
  clipping) and re-enters the loop at the saved turn cursor. Stop tools
  complete the turn in both modes, budget refusals are identical, replay
  serves whole turns so suspensions never occur under a replay scope, and
  Code Mode does not apply (the caller executes the tools). Dropping a
  suspension closes its span as Cancelled.
  TODO(dsrs-phase4-caller-managed) marker removed.
- LM-layer ToolLoopMode::CallerManaged stays: it is the one-exchange
  primitive the interpreter's own loop is built on (lm_call_toolset), and
  external callers still use LM::call_with_tool_loop_mode directly.

Tests: tests/test_interp_conversation.rs (11 tests) covers multi-turn chat
growth with per-turn spans, opening/map-in render parity, typed
continuations, multi-node and empty-turn refusals, suspend/resume
round-trip, dropped-suspension cancellation, dispatch/suspend span parity,
stop-tool and budget parity across modes, and turn-by-turn plus
caller-managed replay. Wrapper equivalence (request_hash equality with the
typed call) added to test_predict_conversation.rs.

Docs: components/predict.mdx, runtime.mdx, tools-and-agents.mdx updated to
the interpreter-native conversation surface.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…spend/resume

Closes RFC 0004 seams 1-2: run_conversation / run_conversation_caller_managed /
resume_conversation on the Interpreter; Predict::build_chat / call_and_parse are
thin wrappers; static prompt-prefix cache and LM-layer compat path deleted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

# Conflicts:
#	crates/dspy-rs/src/ir/interp.rs
#	docs/docs/components/tools-and-agents.mdx
…2 ledger, regenerate API reference

All five open seams now carry landed annotations (seam 3 was already retired).
API pages regenerated against the post-seam surface via docs/scripts/gen_api.py.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@krypticmouse
krypticmouse merged commit aec48fe into main Aug 21, 2026
2 of 9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant