Repository navigation
Refactor and clean up unused code, enhance IR and optimizer - #91
Merged
Merged
Conversation
Zero callers outside its own test. Removes src/trace/export/, tests/test_trace_export.rs, and the Otel*/Rl* re-exports from trace::mod and lib.rs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Deletes Trace::absorb, Trace::redact, Trace::for_component_id, the reserved SpanEvent::Chunk variant, and the Span.links sequential- approximation field (plus its capture-side bookkeeping). Updates the tracing examples and trace tests accordingly; attach_program tests are kept. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ParetoFrontier had zero callers outside its own test; GEPA uses the engine's ScoreMatrix/ParetoView directly. ParetoStatistics (the one type GEPA reports) moves into engine.rs next to ParetoView::statistics, which is now the single statistics implementation. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ctural grammar) New workspace member crates/dsrs-syntax: the shared syntax layer for the .dsrs text format. src/lex.rs is dspy-rs/src/ir/text/lex.rs moved verbatim (pub instead of pub(crate)); src/structure.rs is the include_program! structural checker rebuilt on that one shared lexer; ParseError (line/col/ message) moves here as the shared error type. Leaf on purpose: serde_json only — no dspy-rs, dsrs-macros, facet, or rig, so the proc-macro crate can depend on it without a cycle. Fixture copies of the golden .dsrs artifacts live in tests/fixtures with the parity test (checker accepts everything the full parser accepts) — no more cross-crate relative-path reaching. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tax lexer Delete src/ir/text/lex.rs (moved verbatim to dsrs-syntax) and re-export dsrs_syntax::ParseError from ir::text so dspy_rs::ir::ParseError keeps its path and shape. parse.rs now pulls Lexed/Lexer/Span/Tok from dsrs_syntax::lex — the full parser and the macro's structural checker can no longer disagree about tokens. Golden behavior proven unchanged: Program::from_dsrs/to_dsrs round-trips on tests/fixtures/*.dsrs are byte-identical to the pre-refactor output and program hashes match (canonical text is the program-hash preimage). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
EngineCheckpoint, EvalEngine::checkpoint(), and EvalEngine::resume() had no callers outside their own tests. Drops the feature, its tests, and the now-stale checkpoint mentions in Spend/RolloutCache docs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…uses Deletes ResponseCheck, ConstraintOutcome, evaluate_constraints, and the Constraint::new_check/new_assert constructors (zero callers). Keeps Constraint, ConstraintKind/ConstraintLevel, and the two expression evaluators used by the chat adapter and ir/text parser. lib.rs and typesys re-export lists updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ete the shadow parser Remove dsrs-macros/src/dsrs_syntax.rs (923 lines) — the hand-maintained mirror of the .dsrs lexer + structural grammar. include_program! now calls dsrs_syntax::check from the shared leaf crate, so a grammar change is made once instead of in two lexers and two parsers. Error strings are unchanged (the trybuild UI expectations pass as-is), and the macro crate drops its serde_json dependency, which only the shadow lexer used. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…endencies] Every dep used by 2+ member crates now has a single pin at the workspace root (anyhow, async-trait, futures, reqwest, serde, serde_json, tempfile, thiserror, tokio, plus the rig-core/minijinja git pins and the dspy-rs/ dsrs-tools path deps); members declare 'workspace = true' and add only their own features. Pins match what Cargo.lock already resolves — verified with 'cargo metadata --locked' — so nothing was upgraded. The facet [patch.crates-io] fork pin stays as-is (load-bearing; see TODO to unpin). Also move rstest from [dependencies] to [dev-dependencies] in dspy-rs: it is only used by tests, and downstream users should not compile a test framework. tempfile stays a runtime dep — utils/cache.rs owns a TempDir for the disk cache tier. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ive/ Both are superseded historical documents; each now carries an archival note pointing at docs/v1-vision-report.md and docs/rfcs/. Nothing in the repo referenced their old root paths. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e cache-wide mutex Three defects in ResponseCache: 1. Tempdir lifetime: new() created a tempfile::TempDir, pointed FsDeviceBuilder at it, and dropped the guard at end of scope — deleting the directory the disk tier had just opened, so the disk cache only worked through already-open handles, intermittently. The TempDir now lives in the struct (Arc<TempDir>, shared by clones) for the cache's whole lifetime. 2. Panicking init: the four unwraps in new() could take down the process from LM construction. Disk-tier setup now returns a Result; on failure we log a tracing warning and fall back to a memory-only foyer cache (the storage phase's default noop engine cannot fail). 3. Lock contention: insert_entry took &mut self only to maintain the debug history ring, forcing Arc<Mutex<ResponseCache>> in LM so every call in the process serialized on one async mutex — with an O(n) Vec::insert(0, ..) memmove per insert. The foyer handle is internally synchronized, so all methods now take &self; the history ring is an isolated Mutex<VecDeque> (push_back/pop_front, cap 100) that never blocks cache lookups. LM::cache_handler is now a plain Arc<ResponseCache>. get_history awaited nothing and is now sync (its only callers were inspect_history and one test). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…race TRACING_INITIALIZED was checked before set_global_default but only set after, so two concurrent init_tracing() calls could both pass the check and one would get a spurious SetGlobalDefault error, defeating the documented idempotency. The OnceLock is now claimed up front: exactly one caller wins set(()) and performs initialization; racing losers return Ok immediately. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The stub unconditionally returned Vec::new(). FieldDef.constraints and ClassDef.constraints stay: the .dsrs text format reads and writes field-level constraints (ir/text parse.rs / print.rs). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
dataloader.rs and utils.rs move up one level; the v1 forwarding module is gone. crate::data::DataLoader and friends resolve unchanged; the crate::data::v1:: path alias had zero users. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…entifier A capability becomes a globalThis property, so a tool named JSON produced globalThis.JSON = <shim> and broke every later capability shim (each depends on JSON.stringify/parse); reserved words like `class` were installable but uncallable from script code. - Add a RESERVED_JS_NAMES blacklist: ECMAScript reserved words/literals plus the ambient globals the sandbox depends on (JSON, Object, Promise, Error, Array, String, Number, Boolean, Math, Reflect, Proxy, Symbol, globalThis, undefined, null, typed arrays, URI/number parsing functions, ...). - js_identifier now suffixes colliding names with `_tool` (JSON -> JSON_tool), following the existing collision-mangling style. - Capability::validate_name rejects reserved names at install time. - Document that capability calls are bounded by the remaining sandbox deadline (enforced executor-side in the follow-up commit). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s, bound bytecode cache
Four defects from the security/robustness review of the QuickJS sandbox:
1. Capability-name JS injection: the shim was built with format! and eval'd,
and only add_capability validated names — run_script (Code Mode, dspy-rs
interpreter) installed caller-supplied capabilities unvalidated, so
Capability::new("x; globalThis.leak = 1; //", ...) injected code into the
sandbox bootstrap. install_capabilities (the single choke point every
sandbox goes through) now re-validates every name, and no name is ever
spliced into evaluated JS: a constant SHIM_FACTORY is eval'd once and the
shim is installed via globals.set(name, shim), where the name is only a
property key. JSON.parse/stringify are captured at bootstrap so sandboxed
code cannot corrupt later capability calls by clobbering globalThis.JSON.
2. Deadline did not cover capability calls: the engine interrupt handler
cannot fire while a host future runs, so in Code Mode the deadline bounded
only JS arithmetic. Each capability handler is now wrapped in
tokio::time::timeout bounded by the *remaining* sandbox deadline; on
expiry the sandbox timed_out flag is set (classifies as ExecError::Timeout)
and a clear, catchable error is thrown into JS.
3. Handle::block_on inside spawn_blocking deadlocked on current_thread
runtimes. The hook now spawns the handler onto the host runtime and parks
the sandbox thread on an mpsc channel (recv_timeout with a grace period as
a loud failure path if the runtime is not being driven). Capabilities now
work on both runtime flavors.
4. Bytecode cache never evicted: deregister removed the tool but leaked its
bytecode, and the map was unbounded. deregister now evicts the tool's
entry unless another registered tool shares the hash, and the cache is
bounded at 128 entries with a documented cap-and-clear policy (registered
tools hold their own Arc, so eviction only costs a recompile on miss).
No public API changes: register/execute/tool/run_script signatures untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Capability-name injection and reserved-global names rejected on both the builder path and the run_script path (which previously skipped validation). - A tool named JSON is installed as JSON_tool, leaves globalThis.JSON intact, and other capabilities keep working in the same sandbox. - Capability calls are bounded by the wall-clock deadline: a stalling handler surfaces as a typed timeout, both via Executor::execute and end-to-end through CodeModeTool. - Capabilities work on a current_thread runtime (no Handle::block_on). - deregister evicts bytecode unless another tool shares the source hash, and the cache stays within its 128-entry bound while evicted-but-registered tools keep executing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…etry race Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The alias existed only to keep pre-Schema spellings compiling after the vendored BAML integration was removed; it expanded to the identical facet + serde derives. Update the remaining tests and example 16 to spell it #[Schema]. Historical mentions in docs/adr and docs/specs are left as records. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rsers Chat::from_json (which took &self and ignored it) had no production callers; Message::from_json_value and parse_content_block existed only to serve it, including the legacy string-content and type-tagged formats. Chat::to_json stays. Round-trip/legacy-parse tests that exercised the deleted API are removed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ants
LmError::{Network, RateLimit, InvalidResponse, Timeout} and
ErrorClass::{NotFound, Forbidden} were never produced anywhere — every
provider failure arrives as LmError::Provider. Simplifies class() /
is_retryable() and fixes the stale retryability docs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ReAct exists (modules/react.rs) and the IR program graph ships default-on behind the ir feature — the docs claimed the opposite. Adds ir to the crate-organization list. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…drop BamlType alias Resolved manifest conflict: dsrs-syntax promoted to [workspace.dependencies]. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…dback_helpers, checkpoint, dead typesys/errors, data/v1 flatten) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The structural mutation half of the IR: a typed, validated graph-edit calculus so optimizers can change program STRUCTURE, not just parameter values (the harness-optimization enabler). New `ir::edit`: - `Edit` (serde values — inspectable, diffable, replayable): AugmentSig (the CoT move; copy-on-write SigId), SwapLeaf (Predict <-> AgentLoop, context slot minted/dropped), WrapRetry (parent rewire + downstream port redirect), Remove (Seq step; dangling refs surface as validate.rs's own error), AddTool/RemoveTool (existence + cap ceiling; RemoveTool clears stop_tools), SetStop, SetInstructionDefault. - `Program::edited(&[Edit])` — pure: clone arenas, apply in order, GC dead nodes/newly-orphaned sigs/orphaned params, rebuild the ParamPath index, validate(), stamp lineage.parent like bake(), seal a new content hash. `edited(&[])` is a hash no-op. - `Program::legal_edits(NodeId) -> Vec<EditKind>` — the structural menu for an LLM proposer, per-tool add/remove entries included. - `migrate_overlay(parent, overlay, child)` — re-mints tuned values by ParamPath when the owning leaf's signature still carries (inputs identical, outputs may widen), so instruction/demos survive the CoT move; ModelRefs re-mint by model name. - `Program::leaf_id(name)` — stable relocation across edits. Tests: apply/validate failure modes, GC, text + JSON round trips, overlay migration (25 new tests; all IR suites green). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…its, migrate_overlay) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ReAct reimplements the tool loop with a string trajectory, substring-based
tool dispatch, and a type-erasing single-blob output; the IR AgentLoopNode
plus the #[agent] macro is its replacement. Delete the module, its builder
tests, and both examples that exercised it (19-react, 93-smoke-slice4).
Map/AndThen/ModuleExt::{map,and_then} wrap transforms in
#[facet(opaque, skip)] closures — invisible to optimizers, pure sugar that
adds two types to the reflection story. module_ext.rs had no other items,
so the module goes with them, along with test_module_ext.rs and the
Map/AndThen shape tests in test_module_facet_shapes.rs.
Also drop forward_all's hardcoded kdam progress bar (trivially separable:
one local bar + one .inspect link) and the now-unused kdam dependency.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Optimizer is object-safe: compile(target: &mut OptimizeTarget, engine: &mut Engine) -> Report, with compile_module typed sugar. One Engine replaces the EvalEngine/ProgramEvalEngine split; candidates are name-keyed Candidates injected ambiently (no mutation during evaluation; single install of the winner). Deleted API scrubbed: apply_candidate/ CandidateUndo, DynPredictor, checkpoint/resume, ParetoFrontier wrapper, feedback helpers, MIPRO's PromptCandidate/select_best_traces/ create_prompt_candidates; GEPA's compile_with_valset is now compile_module_with_valset. Examples declare leaves via predictors!. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
dspy_rs::prelude re-exports the curated core surface: Signature (trait
+ derive), Example derive, Predict/ChainOfThought/Predicted/Module,
Predictors + the predictors! macro, configure/LM/LMConfig, Demo,
DataLoader/TypedLoadOptions, TypedMetric/evaluate_trainset/Eval,
Optimizer + the five strategies + Engine/Candidate/OptimizeTarget,
capture/replay entry points, ir::{Program, Interpreter, Overlay, Edit},
and init_tracing.
Additive: the crate-root glob re-exports are untouched. lib.rs docs now
recommend the prelude, and quickstart examples 01-05 import only
dspy_rs::prelude::* as proof it suffices for the basic path. Also
scrubs stale 'ir feature (default-on)' doc mentions left from the
feature collapse.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ta::v1 Traces drops the OTel/RL export sections (trace/export is deleted); state documents the Predictors-keyed ModuleState (&M from_module) and the PredictorInfo::load_state install seam; fx documents clear_instruction/ set_demos/bind/from_overlay and the ambient-candidate role; data drops the data::v1 versioning story and EvalEngine row; evaluation drops the deleted feedback_helpers section; signatures notes #[BamlType] alias removal; tools-and-agents/quickstart/examples stop referencing ReAct and deleted examples; code-mode documents reserved-name mangling, capability deadline bounding, and the bounded bytecode cache; runtime adds run_collecting/ LeafOutcome and the dsrs-syntax expansion-time check in include_program!; utils fixes ResponseCache signatures (&self, sync get_history). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One page for the structural mutation half of the IR: the Edit enum, Program::edited semantics (positional NodeIds, batch validation, identity- preserving GC, cot re-sugaring), EditError/ApplyError, the legal_edits proposer menu, and migrate_overlay. Registered in the Programs as Data nav group and the index components table; program-and-nodes cross-links it and drops its stale checkpointing mention. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s-syntax gen_api.py gains the dsrs-syntax companion crate; all api/*.mdx pages are regenerated from fresh rustdoc JSON at the merged tree (deleted items — ReAct, ModuleExt, DynPredictor, EvalEngine/ProgramEvalEngine, trace exports, pareto wrapper, feedback helpers, BamlType — drop out of the inventory by construction). dsrs-syntax registered under Companion crates in the nav; README and the API overview page document the fourth rustdoc command. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The old overview described the V5/V6 milestone: DynPredictor handles, the facet walker, #[derive(Module)], ReAct, ModuleExt::map, and a planned ProgramGraph/registry layer. Rewritten against current truth: predictors! declaration, the ambient-candidate optimizer contract over OptimizeTarget/ Engine, Predict-as-1-node-program, and the IR edit calculus as the concrete form of structural optimization. dsrs-format.md checked against the current grammar (reserved list, node forms, ports) — no drift found. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The standalone fn used to build a bare Predict::with_tools predictor, silently ignoring max_turns, budget, context, stop_tools, and until_parse — the docs even admitted loop options applied to the module-lowered form only. Predict now carries an optional AgentLoopSpec (builder: with_agent_spec), applied when the 1-node agent program is built: stop_tools/max_turns/until_parse land in the node's StopSpec, budget in its NodeBudget, context in its ContextPolicy. The #[agent] standalone fn feeds its parsed AgentStepOpts through that seam, so both lanes now execute the same honored configuration. model = "..." cannot be honored standalone (model refs bind only in a module program's model table), so it is now a compile error on that path: with model set no standalone fn is generated, and calling it fails with 'expected function, found module'. max_turns = 0 is rejected at macro expansion. New behavioral tests prove stop_tools and max_turns pass-through standalone; trybuild UI tests pin both compile errors; macro and component docs updated to drop the admission. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
RFC 0004 collects the deliberate post-unification remainder in one place: the conversation surface (TODO(dsrs-phase4-conversation)), the caller-managed tool loop still on the LM-layer path (TODO(dsrs-phase4-caller-managed)), the retired shared-ptr-policy marker (moot since the facet walker was deleted), whole-rollout credit assignment in harvest.rs, tool membership as a ParamSlot (the ToolSet gene), and structural optimizers over ir::Edit — one paragraph each on what, why deferred, and the suggested shape. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…de, #[agent] options honored, seams ledger Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…r prelude page Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…r ir::Edit Closes seam 6 of RFC 0004: the edit calculus was fully shipped but no optimizer proposed edits. Structural is a GEPA-style hill-climbing loop over the shared Engine, program lane only: - gathers Program::legal_edits for every leaf, filtered to the kinds it can materialize without free text (AugmentSig as the CoT move, SwapToAgent/SwapToPredict, WrapRetry, Remove, AddTool/RemoveTool; SetStop and SetInstructionDefault stay value-level work); - a reflection LM (prompt_model) reads the canonical .dsrs text, the serialized menu (one JSON object per line with an option number), and the incumbent's per-example feedback, and answers with one option; no prompt model, an unparseable reply, or an out-of-range option degrades to a seeded-uniform pick; - applies the choice via Program::edited, migrates the incumbent overlay with migrate_overlay, loads the child through a caller-supplied RuntimeEnv factory (only the host knows the live bindings), and accepts through the engine's minibatch gate: a strict win on the shared minibatch promotes to a full-set evaluation and the child becomes the new incumbent; - every rejection path (EditError, LoadError, gate loss) is recorded in the step and skipped, never a panic; - budgets are engine budgets: every child is a fresh program whose hash keys fresh rollout-cache rows, so max_rollouts/max_lm_calls are the real control surface; the baseline pass seeds the cache so parent minibatch reads are free. Entry points are compile_program / compile_program_with_overlay rather than the object-safe Optimizer trait: the trait's currency is one lane-erased target, and a structural child is a new program needing a new interpreter, which only the caller's RuntimeEnv can load. The winner is returned (program + migrated overlay) for Program::bake, never installed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ural-optimizer claims - docs/optimizers/structural.mdx: how-to page mirroring the gepa.mdx shape (overview, quick start, config table, report tables, the loop, cost model, troubleshooting), registered in docs.json next to the other optimizer pages; - components/optimizers.mdx: six strategies, comparison row, the compile_program exception note, and full Structural config/step/ report tables; - components/edit-calculus.mdx: the proposer-menu section now points at the shipped optimizer instead of a hypothetical one; - components/optimizer-engine.mdx and the shared comparison snippet pick up the sixth strategy. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…::Edit Closes RFC 0004 seam 6: reflection over legal_edits menus, Program::edited, migrate_overlay, minibatch-gated hill climb inside the existing engine. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…004 seam 4) Demo harvesting scored every span in a rollout with the rollout's single metric score, so a good final answer vouched for every intermediate predictor call — including ones a later step recovered from. Close the seam with span-level evals: - Span gains an optional eval field (serde default + skip-when-None): additive under RFC 0001 §5.1, no format version bump; eval-free traces serialize byte-identically to before. - TypedMetric gains evaluate_spans, a per-trace hook called once per traced rollout after evaluate, returning (SpanId, Eval) pairs that rollout_traced stamps onto the trace. The default body returns no span scores, so every existing metric compiles and behaves unchanged. - harvest::collect_demo_candidates gates and ranks each span on its effective score — the span's own eval when present, the whole-rollout score otherwise — so a scored-down span stays out of the demo pool even from a winning rollout, and a scored-up span qualifies from a losing one. Bootstrap, MIPROv2, and SIMBA inherit this through the shared join. Tests: harvest unit tests (baseline parity without span evals, override in both directions, ranking); JSONL round-trip with the field absent and present; end-to-end draft/refine bootstrap pair showing the whole-rollout metric harvests recovered-from drafts and a span-aware metric keeps them out. Docs: evaluation, traces, optimizers, optimizer-engine, miprov2 pages updated per the docs register. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes RFC 0004 seam 4: TypedMetric::evaluate_spans hook, Span.eval additive field, harvest prefers a span's own eval over the rollout score. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… seam 5)
Which tools an AgentLoop carries becomes an optimizable slot. Declaration
stays structural (AgentLoopNode::tools is the capability footprint and the
gene's legal alphabet); selection is ParamKind::ToolSet at
"<leaf>.tool_set", defaulting to the full declared table so existing
programs print, hash, and run unchanged.
- params: ParamKind::ToolSet, ParamValue::ToolSet { tools }, ToolSetK kind
tag, Overlay::set_tool_set; Overlay::set refuses values naming a tool the
owning agent does not declare (OverlayError::ToolSetUndeclared), so the
serde load path (from_named) is guarded too.
- graph/builder: AgentLoopNode::tool_set: ParamId; the builder mints the
slot per agent node (NodeSpec::tool_set seeds a restricted default).
- validate: load-time only — id range checks for ToolSet values, kind/owner
check for the slot, and subset + duplicate-free checks against the node's
declared table (ToolSetUndeclared / ToolSetDuplicate).
- interp: eval_agent resolves the gene through the overlay and assembles
definitions, by_name dispatch, sandbox code, and the Code Mode surface
from the selection; absent slot = full declared set. A deselected tool
cannot execute even if the model hallucinates its name.
- text: `tool_set [a b]` agent option; canonical print elides it when the
default equals the declared list, so pre-ToolSet hashes are unchanged.
- edit: SwapLeaf to agent mints the slot (predict direction collects it);
AddTool/RemoveTool keep the default in sync with the declaration;
migrate_overlay re-mints ToolSet entries by tool name and intersects with
the child's declared table.
- bake folds ToolSet entries into defaults like every other kind; trace
attach_program now joins the new leaf-owned slot.
- docs: program-and-nodes, dsrs-file, tools-and-agents, edit-calculus.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes RFC 0004 seam 5: ParamKind::ToolSet per agent node, subset-of-declared validation at load time, interpreter honors the selection, edit/migration and text-format wiring included. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ams 4-6 as landed-unreviewed Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…erges Answers the §3 review questions (hash stability, goldens, harvest parity), records per-seam sharp edges, flags in-flight seams 1+2 and deferred work. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…FC 0004 seams 1-2) Seam 1, conversation-in/conversation-out: - Interpreter::run_conversation(chat, input, overlay, budget) -> (RunOutput, Chat): one turn with the program's single leaf over a caller-owned Chat. Opening turns (empty chat + input) render system+demos+input through the same overlay-resolved path as map-in runs, so prompts are byte-identical and request_hash matches across paths. A turn is not a run: one span per turn, seq increments per turn, replay serves turn by turn. - Interpreter::conversation_opening(input, overlay) renders the opening Chat without calling anything. - Predict::build_chat / call_and_parse are now thin wrappers over these entries. The static prompt-prefix cache and the LM-layer conversation path (call_and_parse_with_input, serve_recorded_span) are deleted; build_chat is now async. TODO(dsrs-phase4-conversation) markers removed. Seam 2, caller-managed tool loop as suspension: - Interpreter::run_conversation_caller_managed runs the same AgentLoop in suspending mode: on tool calls it returns ConversationTurn::Suspended(ToolSuspension) carrying the pending calls, the open span guard, and the loop/run meters; resume_conversation feeds results back (same ToolRun events, batched tool-result turn, context-policy clipping) and re-enters the loop at the saved turn cursor. Stop tools complete the turn in both modes, budget refusals are identical, replay serves whole turns so suspensions never occur under a replay scope, and Code Mode does not apply (the caller executes the tools). Dropping a suspension closes its span as Cancelled. TODO(dsrs-phase4-caller-managed) marker removed. - LM-layer ToolLoopMode::CallerManaged stays: it is the one-exchange primitive the interpreter's own loop is built on (lm_call_toolset), and external callers still use LM::call_with_tool_loop_mode directly. Tests: tests/test_interp_conversation.rs (11 tests) covers multi-turn chat growth with per-turn spans, opening/map-in render parity, typed continuations, multi-node and empty-turn refusals, suspend/resume round-trip, dropped-suspension cancellation, dispatch/suspend span parity, stop-tool and budget parity across modes, and turn-by-turn plus caller-managed replay. Wrapper equivalence (request_hash equality with the typed call) added to test_predict_conversation.rs. Docs: components/predict.mdx, runtime.mdx, tools-and-agents.mdx updated to the interpreter-native conversation surface. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…spend/resume Closes RFC 0004 seams 1-2: run_conversation / run_conversation_caller_managed / resume_conversation on the Interpreter; Predict::build_chat / call_and_parse are thin wrappers; static prompt-prefix cache and LM-layer compat path deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> # Conflicts: # crates/dspy-rs/src/ir/interp.rs # docs/docs/components/tools-and-agents.mdx
…2 ledger, regenerate API reference All five open seams now carry landed annotations (seam 3 was already retired). API pages regenerated against the post-seam surface via docs/scripts/gen_api.py. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.