Skip to content

v1 release: revamp the crate and add functional utilities and graph tracing - #90

Merged
krypticmouse merged 78 commits into
mainfrom
v1-revamp
Aug 19, 2026
Merged

krypticmouse merged 78 commits into
mainfrom
v1-revamp

Conversation

@krypticmouse

Copy link
Copy Markdown
Owner

No description provided.

krypticmouse and others added 30 commits July 28, 2026 16:38
The function is the signature: params become inputs, the return type the
output field, the doc comment the instruction, and the fn name the fx
params/trace slot. Plus fx microbenches and Copy-usage cleanups.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Removed, per the v1 vision report audit:
- empty Adapter marker trait; configure(lm) no longer takes an adapter
- core/specials.rs placeholders (History, empty ToolCall shadowing rig's, NoTool)
- BamlValue alias — serde_json::Value everywhere
- DummyLM (superseded by TestCompletionModel)
- field! macro + the schemars dependency it existed for
- generated but never-used {Sig}All helper structs
- empty evaluate/metrics.rs
- non-functional trace::Executor (replay comes back designed, with the
  phase-2 trace format)

Net −699 lines, zero behavior change; full suite green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
LMConfig is the data half — every generation parameter, serde-round-trippable,
PartialEq for candidate comparison, api_key #[serde(skip)] so secrets never
serialize (resolved from provider env vars at client init instead). LM is the
live half: config + initialized client + response cache, built via
LM::from_config or the unchanged LM::builder() fluent path.

Phase-1 prerequisite from the v1 vision report (§5.6): program artifacts can
now record their model configuration.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One Span per Predict invocation (tool-loop round-trips as inner events),
per-trace interning of components/prefixes/model configs, task-local capture,
request-hash determinism for strict + counterfactual replay, JSONL wire format.
Kills trace::Graph, mipro::Trace, and the already-dead evaluate::ExecutionTrace.
Carrier-decoupled: spans store JsonMap, so the 5→2 carrier collapse can't
deadlock with it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…fThought is now pure sugar

Four ways to get CoT existed: the ChainOfThought<S> wrapper struct (with its
own parallel builder + call/forward/Module surface), a manual `reasoning`
output field, Augmented<S, Reasoning> directly, and #[cot]. The canonical
mechanism is the augmentation; everything else now rides it:

- ChainOfThought<S> = type alias for Predict<Augmented<S, Reasoning>> —
  same constructor, builder, call, and Module impl as any Predict; the
  wrapper struct, ChainOfThoughtBuilder, and with_predict are gone
- #[cot] already expanded to Augmented<Sig, Reasoning> — unchanged
- manual `reasoning` fields remain a user choice, not a code path

Integrator note: modules embedding ChainOfThought lose the synthetic
`.predictor` path segment — the field itself is now the Predict leaf, so
optimizer dotted paths and saved ModuleState keys shorten accordingly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Five paths wrote predictor state independently: PredictBuilder fields,
four DynPredictor setters (set_instruction / restore_instruction /
set_demos_from_examples / load_state, each resetting the prompt cache on
its own), triplicated save->set->eval->restore blocks in COPRO, GEPA, and
MIPROv2, ModuleState::apply, and the fx::Params overlay.

Now there is one seam, documented on the trait:

- DynPredictor::apply_update(StateUpdate) is the only mutator; StateUpdate
  is a partial overlay (instruction and/or demos), load_state is a default
  method delegating to it, and ModuleState / fx::Params ride load_state
- Predict::apply_state is the single typed applicator that writes fields
  and invalidates the prompt prefix; the builder and apply_update both
  funnel into it
- the per-optimizer candidate blocks collapse into one shared
  optimizer::evaluate_with_instruction (save, apply, eval, restore --
  restore runs on both success and failure)

No optimizer behavior change; instruction-only candidates still skip the
demo serialization round-trip. Full suite green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
New workspace crate implementing the v1 vision report's Tier-1 tool
runtime (docs/v1-vision-report.md 4.2/5.5):

- `Executor` trait: async execute + validate/register lifecycle, the
  escape-hatch seam for future Wasmtime/subprocess/remote executors.
- `QuickJsExecutor`: fresh in-process quickjs-ng sandbox per call with
  per-call memory limit, interrupt-driven wall-clock deadline (runaway
  `while(true)` is killed as a typed Timeout), and no ambient authority
  (no fs/network/env/module loader).
- `Capability`: explicit host-capability injection - async Rust fns
  exposed as JS globals via a JSON shim (Code Mode falls out for free).
- LATM validate-then-register lifecycle: shape check -> compile ->
  instantiate -> in-sandbox self-test; failures are typed, serializable
  errors an LLM can repair against.
- BLAKE3 content-hash bytecode cache: compile once per unique source.
- `SandboxTool`: registered tools implement `rig::tool::ToolDyn`, so
  they slot into every DSRs surface that already accepts tools.

28 integration tests cover happy paths, deadline/memory kills,
ambient-authority denial, capability round trips and failure
attribution, self-test gating, cache-hit behavior, and the rig bridge.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Verifies `Arc<dyn ToolDyn + Send + Sync>` from `register_rig` coerces to
the plain `Arc<dyn ToolDyn>` that dspy-rs' Predict/CoT/ReAct accept.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…:forward

Eight-plus ways to invoke a predictor: Predict::call and Predict::forward
(inherent duplicates), Predict::forward_continue and Predict::call_and_parse
(one was a pure alias of the other), Module::forward / Module::call,
forward_all and forward_all_with_progress, ReAct's inherent call/forward
mirroring its Module impl, ChainOfThought's ditto (already gone via the
alias), plus fx::predict.

Now:
- Predict::call is the typed direct call (the implementation, tracing
  span and all); inherent forward is gone; Module::forward delegates to it
- call_and_parse is the one chat-level conversation seam (build_chat ->
  call_and_parse per turn); forward_continue deleted
- forward_all absorbs forward_all_with_progress (progress bar stays; the
  bool-parameter twin had no callers)
- ReAct is invoked only through the Module trait; its inherent
  call/forward duplicates are gone
- fx::predict now rides Predict::call

Module::call remains the trait-side caller entry (default delegating to
forward) — that pair is one path, kept for future middleware. Examples
updated where they used the deleted inherent forward.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rmat

SignatureDef as owned value bridging the derive statics; cranelift-entity
arenas with a closed 9-variant Node enum (tool use = AgentLoop only);
field-level dataflow; Overlay as the candidate type unifying fx::Params /
ModuleState / apply_update; one addressing story across params and trace;
23-keyword LL(1) text format with capability declarations; three artifacts
(.dsrs / state / trace) joined by hashes; 7-stage migration plan.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…SONL

RFC 0001 PR-1. One Span per Predict invocation (tool loops are ordered
events inside one span), one Trace per rollout with interned component
names, prompt prefixes, and redacted model configs.

- trace/span.rs: Span, SpanEvent, SpanError, Trace + intern tables,
  Eval, TraceMeta/TraceOutcome, for_component/prompt/successes/absorb
- trace/capture.rs: task-local capture()/capture_with_meta() scope,
  two-phase begin_span()/SpanGuard::finish() with Cancelled-on-drop
- trace/serialize.rs: JSONL header/spans/footer (v1), skip-unknown
  event tags, serialization-time truncation with content-hash markers
- utils/hash.rs: stable FNV-1a hasher; request_hash and the LM cache
  key both use it (DefaultHasher is not stable across toolchains)
- LMResponse.events: ordered per-round-trip Exchange/ToolRun record
  built inside the tool loop; cache hits synthesize one Exchange
- Predict::call_and_parse_with_input opens a span before the LM call
  (failed calls stay visible: prompt+input recorded, output absent);
  parse failures keep raw_output; dual-writes to the legacy graph
  until PR-3 removes it

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
RFC 0001 PR-2. The metric position becomes Eval { score: f64, feedback:
Option<String> } with trace access; optimizers consume rollout traces
instead of pointer-keyed graphs.

- TypedMetric::evaluate gains trace: Option<&Trace> and returns Eval;
  MetricOutcome, FeedbackMetric, and the never-consumed ExecutionTrace
  (+ builder) are deleted; feedback_helpers return Eval with metadata
  folded into the feedback text; average_score is f64
- the evaluation loop captures every rollout (metric runs outside the
  scope, keeping LM-as-judge calls out of the trace) and records the
  eval into Trace::outcome
- optimizers get a naming pass: predictor_names assigns each Predict
  leaf its dotted path as span component name via the new
  DynPredictor::set_trace_name, so traces join to predictors by name
- GEPA reads per-component sub-traces (for_component) into the
  reflection prompt: per-invocation input/output-or-error + tool runs
- MIPROv2 demo harvesting moves off mipro::Trace<S> + the raw-pointer
  instance_keys join onto capture() + Trace::successes(); mipro::Trace
  is deleted, PromptCandidate/min_demo_score widen to f64

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
RFC 0001 PR-3. One trace system remains.

- delete trace/{dag,context,value}.rs: Graph/Node/NodeType, the
  trace()/record_node/record_output scope, TrackedValue/IntoTracked
- Predict records only into the capture sink; the RawExample input
  serialization and pointer instance_key are gone
- CallMetadata.node_id: Option<usize> becomes span_id: Option<SpanId>;
  Prediction and data::Example lose their node_id fields
  (Prediction::get_tracked dies with TrackedValue)
- predictor_instance_keys (raw-pointer join) deleted - span component
  names are the join
- examples 12/14 move to trace::capture(): scoped capture, spans
  addressed by param names, params<->trace name join, JSONL output
- test_fx/test_phase_upgrades trace tests rewritten against spans
  (prefix interning and per-slot seq replace instance-key assertions)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ked TypeId caches

RFC 0002 IR-1 groundwork:
- FieldType, ClassDef, EnumDef, EnumValueDef, FieldDef, Constraint, and
  ConstraintKind gain Serialize/Deserialize (purely additive, §1.3).
- New typesys::TypeTable — the owned class/enum registry a program owns in
  the dynamic lane. OutputSchema is now target + TypeTable.
- render::type_name/schema_block and coerce::coerce take &TypeTable.
- Leaked cache #2 (blanket Schema::output_schema TypeId cache) deleted:
  the trait now returns an owned OutputSchema; the only repeat consumer
  (SignatureSchema) caches it in its own entry.
- Leaked cache #3 (internal_name_for_shape string intern) deleted: the
  name is a trivial format!, returned owned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…che (RFC 0002 IR-1)

- ir::sig::SignatureDef/FieldDef/ConstraintDef/RenderSpec: a signature as an
  owned, serde-derivable runtime value; SignatureBuilder validates at finish()
  with the same rejections as #[derive(Signature)] (empty side, duplicate
  aliased names, bad format value, invalid jinja, check-without-label,
  malformed constraint expression, non-string map keys).
- SignatureDef::of::<S>() bridges the derive statics (FieldSchema -> FieldDef,
  flattened fields keep leaf names); types_of::<S>() exposes the class/enum
  TypeTable from the same entry; matches::<S>() is the structural-equality
  check include_program! will use; augmented_with() is the value-lane
  Augmented<S, A>.
- The remaining two leaked TypeId caches collapse into the single
  StaticSigCache: SignatureSchema::of now delegates to it (leak #1), and the
  per-derive OnceLock schema() fast path is no longer emitted (leak #4).
  Loaded programs own their arenas; nothing dynamic touches the cache.
- Golden tests: derive-built == builder-built for representative signatures
  (plain, aliases+constraints+format+docs, class/enum tables, flatten,
  generic Augmented), serde round-trips, builder rejections, pointer-stable
  cache.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…, Pareto matrix (§5.4)

The evaluation core every optimizer becomes a thin strategy over:

- Candidate as data: named overlay set (predictor name -> instruction/demos)
  with a stable order-insensitive content hash; applied/restored in exactly
  one place (apply_candidate/restore_candidate) through the DynPredictor
  apply_update mutation seam. Pre-IR form of the §5.4 overlay contract.
- Bounded-concurrency async fan-out over (candidate x examples) with
  per-rollout trace capture (buffer_unordered, no unbounded spawns).
  Example-level parallelism per candidate; candidate-level parallelism is a
  documented seam awaiting render-time overlays (IR).
- Rollout cache keyed (baseline hash, candidate hash, example uid, salt):
  re-evaluation of a seen (candidate, example) returns the cached Eval with
  zero LM and zero metric calls. Baseline hash invalidates entries when the
  module skeleton changes mid-run; salt is the sampling-params seam.
- Budget metering (max metric calls / max LM calls / max tokens): the engine
  stops cleanly before an unaffordable batch and reports Spend, including
  exact span counts and token totals from traces.
- Minibatch gating (evaluate_gated): the GEPA acceptance pattern as engine
  API — promote to full eval only if the minibatch mean beats a threshold.
- ScoreMatrix (candidates x examples) with ParetoView win/frontier/statistics
  bookkeeping, generalizing GEPA's ParetoFrontier for any strategy, including
  column-restricted views for split train/val sets.
- Checkpoint/resume: engine state (candidates, matrix, spend, cache) as JSON;
  a resumed run skips completed rollouts via the cache.

Tests: concurrency-proof fan-out (barrier), cache hit accounting, budget
stop, gate promote/reject, matrix/Pareto bookkeeping, checkpoint->resume
equivalence, candidate seam reversibility, canonical-hash stability.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e engine contract

Teacher pass over the trainset under trace capture (via the shared engine),
demo harvest by pure trace name-join (successful spans of rollouts scoring
>= min_demo_score become demo rows for the predictor that recorded them),
demos attached as a Candidate overlay, engine-evaluated against the baseline,
installed through the single candidate seam only if strictly better.

Harvesting (demo_from_json / collect_demo_candidates / select_demos with
top-k + input dedup) is extracted to optimizer/harvest.rs so MIPROv2's
existing harvesting can reuse it when it moves onto the engine.

Tests: end-to-end adopt-when-better on canned responses (demo contents
verified per-predictor by name join), keep-baseline-when-worse, no-demos
short-circuit (no second pass), max_demos cap.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- The four system-prompt builders are now lane-neutral view functions
  (FieldView) shared by both paths, so a derive-bridged SignatureDef renders
  byte-identical prompt sections to the static SignatureSchema path — by
  construction, and proven by parity tests.
- New ChatAdapter methods with no 'static requirement anywhere:
  build_system_def, format_input_def (JsonMap in), parse_output_def
  (JsonMap out, per-field FieldMeta, constraint enforcement).
- Dynamic-lane Jinja templates and constraint expressions compile per call
  (typesys::evaluate_expression) — the process-global template/expression
  caches stay keyed on the derive's &'static strs only, so loaded defs
  never leak.
- Proof tests: runtime-only signature (hand-built TypeTable with class +
  enum, aliased output, assert constraint) formats system + input and
  round-trips a canned LM response through parse; lane-parity for system,
  input (incl. jinja), and parse.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Reflective mutation, minibatch acceptance, and Pareto sampling become a
strategy over engine primitives (CandidateEval, ParetoView); gepa.rs's own
eval/budget/pareto plumbing deleted, pareto.rs slimmed to the shared matrix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ic cache

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ets, charge, versioning, baseline-hash and salt cache keys

The engine itself was complete against the §5.4 checklist; these tests pin
the behaviors that had no coverage: the fan-out never exceeds
EngineConfig::concurrency, evaluate_gated reports budget exhaustion for both
the minibatch and the promotion leg, charge() meters auxiliary spend into the
budget and survives checkpoints, unknown checkpoint versions fail resume,
permanent installs invalidate cached rollouts via the baseline hash, and the
cache salt partitions the rollout cache across resumes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Each candidate instruction becomes an overlay Candidate evaluated through
EvalEngine's cached bounded-concurrency fan-out; the per-round winner is
installed through the one candidate seam (apply_candidate), whose baseline-
hash change correctly invalidates stale cache entries between rounds. The
bespoke save/set/eval/restore loop (score_candidate / set_instruction over
evaluate_with_instruction) is gone; candidate generation is unchanged and
COPRO's public surface and behavior are preserved.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The teacher pass is now the engine's traced baseline candidate, demo
bootstrapping goes through the shared harvest name-join (mipro's private
demo_from_json/select_demos duplicates are gone), demos and per-predictor
instruction winners install through the one candidate seam, and trial
evaluation is the engine's cached minibatch fan-out — mipro now holds
strategy logic only. One documented behavior change: teacher-pass LM/metric
failures now propagate (the engine contract) instead of being skipped with a
warning; that tolerant path had no test coverage and every covered behavior
is preserved.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… engine

Each step samples a seeded minibatch, contrasts the current program's best
and worst rollouts (served from engine bookkeeping — the current program
always has full-trainset coverage, so introspection costs no rollouts), and
proposes exactly one move: append-demo (one demo per predictor harvested
from the best rollout via the shared trace name-join, capped and deduped) or
append-rule (a rule distilled by a reflection LM from the contrasting
rollouts — metric-feedback concatenation without one — appended to the
predictor with the most spans in the worst rollout). Acceptance is the
engine's minibatch gate; promotions refresh the rollout store from the full
pass and the winner installs through the one candidate seam. Reflection
calls are charged against the engine budget.

Deterministic tests on TestCompletionModel canned responses (reflection LM
canned too) cover both move types, the no-prompt-model fallback, gate
rejection leaving the module untouched, and clean budget stops.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
krypticmouse and others added 29 commits August 14, 2026 09:41
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…image hash (RFC 0002 IR-5)

The text form is now the wire form of a program:

- Hand-rolled recursive-descent parser (LL(1) after keyword dispatch, zero
  new dependencies): .dsrs text -> validated Program, reusing the builder
  frontend for lowering so both frontends construct the same runtime value
  by construction. Errors carry line/column plus expected-token context —
  these programs are model-generated and parse errors are the feedback
  signal; a shadow tree maps post-lowering validation errors back to
  source positions, and capability-ceiling violations are caught at parse
  time with the exact cap span.
- Deterministic canonical printer with documented ordering rules
  (print::canonical form rules 1-10): parse(print(p)) prints identically,
  print(parse(t)) is the canonical form of any valid t. cot re-sugars from
  the augmented signature; tool interfaces print inline; container steps
  get stable synthesized names (_0, _1, ...).
- Program::compute_hash preimage rebased from canonical JSON to the
  canonical printed text minus the lineage block (the RFC 0002 §5 rule).
  Load paths now validate before sealing so hostile arenas fail the load,
  never a print; overlay base-hash guards work unchanged across both
  frontends, and the serde-JSON projection agrees with text on the hash.
- Public API: Program::from_dsrs / to_dsrs / load_dsrs / save_dsrs
  (text-only artifact — binary files rejected), exported via dspy_rs::ir.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…, parse-error quality

- qa.dsrs: the RFC 0002 §4.3 worked example (cot -> agent loop -> hole) as
  canonical text; parsed program == builder program (same canonical print,
  same program hash), and the fixture is byte-for-byte the builder's print.
- qa_scrambled.dsrs: shuffled declarations, collapsed whitespace, explicit
  defaults, split out-steps — canonicalizes to the golden with an unmoved
  content hash.
- kitchen.dsrs: the whole closed node vocabulary (predict/cot/agent/hole/
  seq/fork/route/retry/refine/loop) plus classes, enums, exotic types,
  sandboxed tools, demos, budgets, context policy, and lineage; lineage
  round-trips but stays out of the hash preimage.
- parse . print = id and hash stability over every fixture; serde-JSON and
  text agree on the hash; overlay base guards hold across frontends; code
  gene hashes recompute from source; save/load round-trips and binary
  artifacts are rejected.
- Parse-error quality: missing capability declaration, unknown keyword,
  unbound node reference, binding type mismatch, unbound input, reserved
  names, duplicate names, unbounded loop, unknown sig/model/tool, wrong
  format major — each asserted to name the offending line and the problem.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The paste-into-a-system-prompt reference for program generation: file
skeleton, type syntax, all ten node forms with their option blocks, port
forms, the hard rules the compiler enforces, and the worked example —
generated from the landed parser/printer, not the RFC sketch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…age hash

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Cross-branch reconciliation: IR-6 added overlay provenance to Lineage;
the IR-5 text format now prints and parses it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Auto-wrap any rig ToolDyn as a Capability: name/description fetched from
the tool's definition at wrap time, JS args -> JSON -> ToolDyn::call ->
JSON result (string fallback). Tool errors surface as catchable JS
exceptions attributed to the original tool name, and as typed
ExecError::Capability when uncaught. js_identifier documents the name
mangling rule; Capability::from_toolset batch-wraps and errors on
identifier collisions instead of silently shadowing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… IR-7, §6.1)

Embeds a .dsrs artifact as a module (file stem) with SOURCE, program(),
try_program(), and a generated #[cfg(test)] validation test.

Layering (documented in the macro rustdoc): the full parser lives in
dspy-rs, which depends on this crate, so calling it here would be a
dependency cycle. Build time therefore gates SYNTAX ONLY — a standalone
structural check against the same surface grammar (pragma, top-level
keyword vocabulary, declaration shapes, balanced delimiters, strings/
numbers/fences), deliberately more permissive inside blocks than the
real parser. Semantics (types, dataflow, caps) validate at first use
via LazyLock<Program::from_dsrs> and at CI time via the generated test
— the sqlx-offline analogue.

Path resolution: CARGO_MANIFEST_DIR (sqlx rule), invoking file's dir as
fallback (Span::local_file). include_str! keeps rustc rebuild tracking.

Tests: syntax-checker unit suite incl. parity over the golden fixtures
(everything Program::from_dsrs accepts must pass the gate), happy-path
integration test (embedded qa.dsrs loads, hashes agree, generated test
runs), trybuild ui: missing file, non-literal arg, syntax error with
line/column.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… calls

The Cloudflare Code Mode / CodeAct pattern: CodeModeTool::new(tools,
config) wraps a tool set as sandbox capabilities and exposes a single
rig tool named run_js whose auto-generated description lists the
injected JS APIs with token-compact argument schemas (a plain string fn
— the future optimizable ToolDesc default). Scripts run via the new
run_script primitive: async-IIFE body in a fresh no-ambient-authority
sandbox, deadline/memory limits enforced, results settled on the
microtask queue. Model-repairable failures (JS errors, tool failures,
limit kills, bad args) return as typed-JSON tool results so outer tool
loops feed them back instead of aborting; only ExecError::Internal is a
hard error.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…R-7, §6.2)

New crate crates/dsrs-cli (binary `dsrs`); axum/clap/tokio confined
here — the library crates gain no deps.

- `dsrs check p.dsrs`: full parse+validate via Program::load_dsrs; the
  parser's line/column "expected ..." messages go to stderr, exit 1 —
  the LLM regeneration signal and the human pre-commit gate.
- `dsrs fmt p.dsrs [--write]`: canonical print (Program::to_dsrs);
  --write rewrites in place only when bytes differ.
- `dsrs serve p.dsrs [--host --port] [--overlay o.json] [--allow cap]`:
  fail-fast load before binding the port — named-form overlay verified
  against program hash + slot kinds, grants checked (caps ⊄ grants
  refused printing the missing set, per the RFC sketch's --allow),
  models bound from env-held secrets, QuickJS sandbox attached, holes
  registered. Endpoints per the IR-7 brief: POST /run (?trace=1 adds
  the RFC 0001 trace JSONL with param_ids attached), GET /schema (main
  SignatureDef + TypeTable serde forms), GET /program (canonical
  text), GET /healthz. Input errors 400, run failures 500.
- Host-tool programs are refused with a hint (embed via
  include_program! instead) — a CLI cannot bind ToolDyn values.

Everything is a library function; main.rs is arg parsing + exit codes.
Tests drive check/fmt directly and the exact production router on an
ephemeral port with TestCompletionModel pre-bound through RuntimeEnv
(the honest injection point — no bespoke --canned flag): health,
schema, program text, run, traced run, missing-field 400, non-object
400, overlay read-through asserted in the rendered prompt, caps
refusal + --allow, stale-overlay refusal.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…lane

ToolSet::code_mode(tools, config) builds a ToolSet containing just the
CodeModeTool, usable in Predict/LM tool loops today: the model writes JS
against the tools instead of emitting N tool-call JSONs. Gated behind a
new code-mode feature (default-on: ir already pulls dsrs-tools, so it
costs nothing; --no-default-features stays dep-light). Re-exports
Capability/CodeModeTool/SandboxConfig/RUN_JS_TOOL_NAME at the root.
Integration test: canned LM chains two tools in one run_js execution;
script errors come back as typed tool results and the loop survives.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Covers the path users write (dspy_rs::include_program!) and the
proc-macro-crate FoundCrate::Itself branch: the embedded program and
Program::load_dsrs agree on name, content hash, and canonical text.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
RuntimeEnv::with_code_mode(config) makes every AgentLoop collapse its
non-stop tools into one sandboxed run_js definition; the model writes JS
against them as globals (host tools call through ToolDyn, sandboxed
tools route through the executor under their registered names). Chosen
as a RuntimeEnv option, NOT a ToolKind variant: code mode is host
presentation strategy, not program semantics — same artifact, same
genes, same program hash either way; ToolKind is per-tool while code
mode collapses a whole loop's surface; and the closed enum / .dsrs text
format stays untouched. Overlay-resolved ToolDesc genes still flow into
the generated run_js description. Stop tools stay individual (the loop
must see their calls by name). JS-identifier collisions refuse the LOAD.
run_js executions record ToolRun span events like any tool; script
failures stay conversational.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…le + IR lanes

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The bare 'reference/' pattern in .gitignore (meant for upstream source
copies at the repo root) also matched docs/docs/reference/, so all five
Reference pages were untracked and the published site shipped dead nav
entries. Anchor the pattern to the repo root and track the pages.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Concepts essays, how-to guides, four tutorials, nav config, and the
reworked quickstart were living untracked/modified on this branch.
Snapshot them as the baseline the Field Guide revamp builds on.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Learn becomes a 15-chapter book (placeholders for now), Guides keeps the
task surface plus the optimizer recipes, Reference grows to 19 pages.
Old slugs redirect. Blank pages (data/prediction, tutorials/overview)
and the orphan community stub are gone; adapter/lm move under reference;
traces-and-cli becomes traces ahead of the CLI split. New landing is the
book's Prologue.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The book begins: Prologue through Give the Model Hands, all recast onto
the Front Desk running example with the map-and-territory thread. Steps,
signatures, predictors, the adapter interlude, struct/fx lanes, the
#[module] keystone, holes, agents, and the Code Mode collapse.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… retired

Nothing To Declare through Running in the Wild: capabilities, traces and
replay, evaluation, the optimizer family, GEPA with an LLM judge, bake/
check/serve/include_program, and the production flywheel coda. The
concepts essays, four tutorials, building-blocks pages, and data pages
they absorbed are deleted (redirects already in docs.json); quickstart
repointed at the new surfaces and de-BAMLed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s 18-26

Every chapter's Rust now mirrors a runnable example (cargo check green):
code-mode and ReAct demos, plus seven Front Desk milestones spanning
contracts, the struct/fx lanes, #[module] with a hole, agents and caps,
replay, evaluation, and overlay tuning. Chapter corrections that fell
out of verification: real to_dsrs()/OPACITY output quoted verbatim, the
actual CLI refusal messages, Option<&Trace> metric signatures, the
two-arg Eval::with_feedback constructor, struct-literal FrontDesk
construction, and an honest paragraph on binding named model handles.
Also places the seven hand-drawn hero SVGs (theme-aware inline JSX).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ples index

The four optimizer pages now match the real API: one-arg configure, the
three-arg TypedMetric returning Eval, live builder fields only (unused
knobs called out), MIPROv2's actual four phases, GEPA intro rewritten to
concrete claims. Their duplicated comparison tables collapse into one
snippet mirrored from optimizer/mod.rs, which also finally lists SIMBA
and BootstrapFewShot. New examples index covers 01-17 plus the Front
Desk set with chapter cross-links.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Bare angle-bracket generics in an ADR and RFC 0002 hard-stopped
mint broken-links before it could scan the site. Backticks only,
no content changes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Feedback on the narrative book: too casual, and splitting a component
across a chapter, a reference page, and a guide re-scattered the docs.
This collapses everything into a single Documentation tab where each
component has exactly one self-contained page: short declarative
opening, working code early, complete API tables on the same page.

The 15-chapter book, the reference tab, and five recipe guides are
deleted; their verified substance (API tables, captured .dsrs/OPACITY
output, real error messages, diagrams, compile-checked snippets) is
merged into 22 component pages plus a single How DSRs Thinks concepts
page. gepa-llm-judge merges into the GEPA page. All old slugs redirect.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
New tab indexing every public item across dspy-rs, dsrs-tools, and
dsrs-macros: one page per top-level module, tables of structs, enums,
traits, functions, macros, and re-exports with compiler-sourced doc
summaries, each item deep-linked to docs.rs for full signatures.
Generated by docs/scripts/gen_api.py from rustdoc JSON (stable
toolchain via RUSTC_BOOTSTRAP); pages stamp the crate version and
commit they were built from. Regen commands documented in docs/README.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…IR module build)

Pre-existing working-tree state committed as-is ahead of the example-rows
refactor; not part of it. (Macro registrations for these land with the next
commit's dsrs-macros lib.rs, so this commit is not standalone-buildable.)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…Demo<S> + ToInput/ToOutput

The signature-bound pair Example<S> is gone. Trainsets are Vec<E> for any
row struct, connected to modules by ToInput<M::Input> (and ToOutput<O> for
seeding labeled demos):

- #[derive(Example)] with #[example(Sig)] targets and #[input]/#[output]/
  #[meta] field marks generates the projections by field name, checked at
  compile time; default partition is everything-not-input-is-output, any
  explicit #[output] switches to explicit mode, and an empty output set
  skips ToOutput (gold-only rows against reasoning-style signatures)
- (input, output) tuples and identity impls work out of the box
- TypedMetric<E, M> receives the full row, so gold data no longer has to
  fit the module's output type (metric-only fields like supporting facts)
- Optimizer::compile / evaluate_trainset / EvalEngine are row-generic and
  the compile turbofish is gone (E infers from the trainset argument)
- Demo<S> replaces Example<S> on Predict, the one place signature
  coupling is load-bearing (schema-driven prompt rendering)
- DataLoader is row-generic (DeserializeOwned + Facet), keeping field_map,
  unknown-field policy, and shape-driven scalar coercion; Option<_> row
  fields tolerate missing source columns

Migrates 16 test files (+ new test_example_rows.rs reference suite),
14 examples, the component/optimizer docs pages, and regenerates the API
tab. Workspace check clean across all targets; 418 tests passing.

Note: lib.rs (both crates), predict.rs, and test_trace_capture.rs also
carry small in-flight M1 edits that were already in the working tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tributes

Follow-up to the row refactor, per review: #[derive(Example)] carried two
kinds of redundancy. It named its target signature in #[example(Sig)] while
the call site already knows the target type, and it needed a shopping list of
companion derives to satisfy the loader/optimizer bounds.

The derive now takes no arguments and no field marks. It generates blanket
ToInput<I>/ToOutput<O> impls that project a row into any target type by field
name through serde, so one row type serves every signature whose input it can
fill and extra columns are metric-only by construction — no #[input],
#[output], or #[meta] to keep in sync with the signature.

The trade is compile-time field checking for runtime: ToInput/ToOutput are
now fallible (anyhow::Result), and a missing or mismatched field returns an
error naming both the row and the target type, propagated as an eval/compile
error. Tuple rows remain the compile-checked path. The identity impl
(T: Clone -> ToInput<T>) is dropped: it collided with the blanket impls.

Docs and API tab regenerated; workspace clean, 418 tests + doctests passing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@krypticmouse
krypticmouse merged commit 5682d46 into main Aug 19, 2026
3 of 10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant