Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,8 @@ source lives under `plugins/railyard/`; everything else is documentation.

## Always

- Before every push, satisfy the [whole-candidate review gate](plugins/railyard/references/whole-candidate-review.md), including repair and release pushes. This is a project publishing requirement for CE; a delta review or test pass alone does not satisfy it.

- Choose model and reasoning effort deliberately for each assignment;
use `railyard:model-routing` when resolving or changing an allocation.
Routine work runs natively; select CE workflows when they help, and use
Expand Down
4 changes: 3 additions & 1 deletion docs/agents/release-coupling.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,5 +22,7 @@ Never treat an installed plugin cache as the source repository.

Documentation-only changes (`docs/**`, `README.md`, `UPSTREAM.md`) need no
version bump, no marketplace repin, and no fleet redeploy/convergence pass —
commit and push them directly. Only changes under `plugins/` couple to the
they still require the [whole-candidate review gate](../../plugins/railyard/references/whole-candidate-review.md)
before every push. The exemption removes release machinery, not publication
review. Only changes under `plugins/` couple to the
release machinery above.
27 changes: 27 additions & 0 deletions docs/whole-candidate-review-baseline.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# Whole-candidate review baseline

PR18 in `clairernovotny/dotfiles` provides an observed baseline on September 29,
2026: six unique actionable findings escaped local review and were identified
by the external Codex reviewer. This is one migration PR, not an estimated
fleet-wide escape rate. The following comment IDs identify the source evidence:

| Finding | Source comment |
| --- | --- |
| Tilde-based OpenCodex/Codex home paths | https://github.com/clairernovotny/dotfiles/pull/18#discussion_r4138976317 |
| Valid agents-table TOML spellings | https://github.com/clairernovotny/dotfiles/pull/18#discussion_r4139039498 |
| Quoted root review_model keys | https://github.com/clairernovotny/dotfiles/pull/18#discussion_r4139066047 |
| Duplicate unrelated disabled-model entries | https://github.com/clairernovotny/dotfiles/pull/18#discussion_r4139066053 |
| Skipped catalog sync reported on stderr | https://github.com/clairernovotny/dotfiles/pull/18#discussion_r4139276692 |
| Portable writes before Sol prerequisite checks | https://github.com/clairernovotny/dotfiles/pull/18#discussion_r4139276698 |

Initial review findings, local catches, successful push count, external revision
round count, and final settlement elapsed are unknown in this bounded baseline.
The six comments do not establish six revision rounds. Do not convert missing
values into zeros or assert a 98% improvement.

For subsequent deliveries, use the compact ledger in
[the shipped review gate](../plugins/railyard/references/whole-candidate-review.md).
Record receipts, unique local catches, escaped actionable findings, successful
pushes, external revision rounds, and settlement timestamps. Evaluate fewer
external revision rounds only after comparable completed deliveries supply
those measurements; this instruction change alone proves no outcome improvement.
2 changes: 1 addition & 1 deletion plugins/railyard/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "railyard",
"version": "0.12.6",
"version": "0.12.7",
"description": "Lean automatic workflow selection, deliberate model and reasoning effort, native delegation, selected CE delivery, and explicit fleet orchestration.",
"author": {
"name": "Claire Novotny LLC",
Expand Down
2 changes: 1 addition & 1 deletion plugins/railyard/.codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "railyard",
"version": "0.12.6",
"version": "0.12.7",
"description": "Lean automatic workflow selection, deliberate model and reasoning effort, native delegation, selected CE delivery, and explicit fleet orchestration.",
"author": {
"name": "Claire Novotny LLC",
Expand Down
3 changes: 2 additions & 1 deletion plugins/railyard/hooks/routing-charter.js
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ const lines = [
"Railyard routing:",
"- Ordinary work uses native tools and subagents. Pick Compound Engineering stages when they help:",
" ce-debug, ce-plan, ce-code-review, or lfg. Open PRs with compound-engineering:ce-commit-push-pr.",
"- Before every push, load Railyard's references/whole-candidate-review.md gate, including direct CE/LFG publishing calls.",
"- CE owns review settlement and CI watching. An explicit Deliver request (railyard:deliver) runs",
" through merge, required release or deployment, and consumer verification unless narrowed; honor",
" plan-only, local-only, and PR-only stops. Merges go through deliver's CE snapshot handoff.",
Expand All @@ -17,7 +18,7 @@ const lines = [
];
// Presence enables advice, not an inference call or a credential disclosure.
if (process.env.TYPESAFE_API_KEY?.trim()) {
lines.push("- Jev is configured: consult railyard:jev when a model, effort, or workflow choice is genuinely open; skip explicit choices, clear defaults, and small dispatches.");
lines.push("- Jev configured: consult railyard:jev for open model, effort, or workflow choices; skip explicit choices, clear defaults, and small dispatches.");
}
process.stdout.write(lines.join("\n") + "\n");

Expand Down
7 changes: 7 additions & 0 deletions plugins/railyard/hooks/routing-charter.test.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -107,6 +107,13 @@ test("startup keeps CE ownership, the deliver endpoint, and explicit-only orches
assert.match(out, /completion events rather than polling/);
});

test("startup supplies the shared publication gate to direct CE and LFG callers", (t) => {
const out = flat(run(fixture(t)).out);
assert.match(out, /Before every push, load Railyard's references\/whole-candidate-review\.md gate/);
assert.match(out, /including direct CE\/LFG publishing calls/);
assert.ok(readFileSync(path.join(path.dirname(script), "../references/whole-candidate-review.md"), "utf8").includes("complete cumulative change-set"));
});

test("configured Jev adds one line without exposing credentials or startup content", (t) => {
const home = fixture(t);
const base = run(home).out;
Expand Down
106 changes: 106 additions & 0 deletions plugins/railyard/references/whole-candidate-review.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,106 @@
# Whole-candidate review before publication

The delivery owner applies this gate before every push: initial publication,
feedback repairs, rebases, release/version changes, stack updates, and pushes
inside CE or LFG. It is a project publishing requirement supplied to the CE
publisher, not a second settlement watcher. Review does not authorize a push
outside the user's endpoint. A focused implementation assignment or a narrow
child handoff does not narrow the integration owner's review scope.

## Candidate and investigation

Resolve the intended target repository, branch, base commit, and exact candidate
commit/tree. Review the complete cumulative change-set against the integration
base, never only the latest commit or last few changed lines. For stacks, cover
each pushed head against its intended base plus interactions across the stack;
record those bases and heads. If a base or target is unresolved, stop publication.
Include every intended published file, including manifests, version bumps,
generated artifacts, deletions, renames, and executable/mode changes. Do not
include or disturb unrelated work. A provisional index/worktree review is valid
only if the final published tree matches it and the final commit and base are
bound in the receipt. Changed commit metadata requires renewed identity checks;
changed tree, base, requirements, or dependencies requires renewed review.

Use the existing Thermos correctness/security and maintainability passes alongside
Codex review, using the allocation selected by model-routing. Their completed
reports can fulfill this gate; do not add a third review tree or watcher. Supply both reviewers
with the complete candidate, user requirements, and affected lifecycle. Trace
unchanged dependencies and real entrypoints through producer, packaging,
publication, installation, startup, and consumer contracts where applicable.
Check host-owned state preservation and legal existing input forms. Investigate
failure paths, prerequisite ordering before writers, partial application,
recovery/idempotence, and false success from stale state or diagnostic streams.
Diff scope bounds reported defects, not investigation: report defects introduced
or exposed by the candidate, including interactions with unchanged code, rather
than unrelated historical issues. Tests, matching bytes, and receipts prove only
the behaviors and inputs they actually cover.

Before each push, revalidate the complete candidate and coverage map. Reuse
prior review/proof only for unchanged inputs and behaviors: compare their hashes,
requirements, dependency/runtime identities, assumptions, and affected
interactions. Recheck changed interactions and uncovered boundaries; a new receipt
must explain reused evidence. A previous clean verdict is not transferable to
new code merely because the newest patch is small. Collect every required
reviewer's completed report; progress or spawn acknowledgements are insufficient.

## Hash-bound receipt and publication decision

Keep a compact receipt in a user-owned local evidence location outside the
candidate payload to avoid self-referential hashes. Bind the base SHA, exact
candidate commit/tree, and SHA-256 of the cumulative binary diff. Link completed
broad review reports and applicable test results, coverage, and findings
dispositions. Hash pertinent dependencies or record runtime identities when they
matter to the behavior or reused proof; a deterministic manifest of every external
input is not required, particularly for trivial documentation changes. Existing
reports and test logs are sufficient evidence when their scope and identity match;
do not duplicate them or rerun passed tests without invalidated proof. Record:

- Target repository/branch; base SHA, candidate SHA/tree, cumulative diff SHA-256;
completed review report and applicable test references, reviewer identities and
observed model/effort (unknown when unavailable), and completion times.
- Coverage mapping each requirement and affected lifecycle boundary to inspected
inputs and behavior evidence, including reused proof and why it remains valid.
- Known unverified boundaries with reason, risk, disposition, and owner. Separate
required pre-push checks from future post-merge release, installation, and UI or
consumer verification. Required pre-push behavior missing evidence blocks
publication. Future consumer verification that depends on publication belongs
to the delivery tail: mark it pending with a named owner, never passed, and do
not force premature deployment to satisfy this gate. Use disposable fixtures
for relevant failure cases; do not perform unauthorized live mutations.
- Deduplicated findings, severity, evidence, and disposition: fixed with proof,
rejected with rationale, or explicitly accepted by the authorized owner.
Unresolved actionable findings and missing required pre-push behavior evidence
block publication unless the user explicitly accepts the stated risk. Stale
candidate identity and incomplete required broad reviews also block publication.
Missing formal receipt detail is not itself a code bug; seek only the compact
identity, review coverage, and disposition evidence needed for this decision.

Immediately before the publishing command, the CE publisher verifies the final
commit/tree, base, pertinent inputs, and receipt still match; stops if they do not; and
records the successful remote head afterward. Failed or unknown push results
must be read back before retrying and are not counted as successful publication.
Keep receipt references in the delivery handoff. This is an instruction gate;
Git itself does not enforce it automatically.

## Compact delivery ledger

The delivery owner starts one ledger per PR/change and updates it after each
review, push, external feedback batch, and settlement. Link receipts and source
comments, deduplicate the same defect across reviewers, and keep these fields:

| Field | Definition |
| --- | --- |
| Initial review findings | Unique actionable findings in the first whole-candidate local review, before repairs |
| Local catches | Unique actionable defects caught locally before their first publication, across all candidates |
| Escaped actionable external findings | Unique valid defects first identified externally after publication; exclude duplicates, rejected comments, and newly requested scope |
| Pushes | Successful published head updates, including initial publication; distinguish failed attempts |
| External revision rounds | External actionable feedback batches that require a published repair; multiple pushes for one batch remain one round |
| Settlement elapsed | UTC time from first successful publication to CE's settled disposition; pending while unsettled |

Record candidate IDs, timestamps, and findings dispositions with each event so
counts can be audited. Retrospective unavailable values stay unknown, not zero.
Compare escaped findings and external revision rounds across comparable completed
changes; report sample size, scope, and settlement elapsed. A local catch rate,
if useful, is local catches / (local catches + escaped findings), deduplicated
by defect and only when both counts are known. Measure whether revision rounds
fall; do not promise a percentage or claim improvement before outcome data.
18 changes: 16 additions & 2 deletions plugins/railyard/skills/deliver/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,18 @@ never patch its source or plugin cache. If a selected stage is missing, report
it, complete independent work, and install it through the supported manager
only within existing authorization.

## Before every push

Complete the [whole-candidate review gate](../../references/whole-candidate-review.md)
before every push, including the first publication, focused feedback fixes,
release/version changes, and pushes inside LFG or CE continuations. Pass this
requirement and its exact-candidate receipt to the CE publishing owner as a
project publishing requirement before invoking a stage that can push. Keep
ownership of the gate across handoffs; CE remains the only settlement owner.
Revalidate the complete cumulative candidate each time, reusing prior proof
only for unchanged inputs and behaviors. Do not substitute the last few
changed lines. Keep the gate's delivery ledger through settlement.

## One review and CI owner

CE alone owns review settlement and CI/PR monitoring. When LFG already runs
Expand All @@ -74,8 +86,10 @@ user action.

## End-of-PR review

Once the PR's change is complete, run these two in parallel against its base
and hand the combined findings to the CE owner before settlement:
For each publication candidate, run these two in parallel against its base
and hand the combined findings and coverage receipt to the CE owner before
pushing and before settlement. Reuse a pre-push review at settlement only
when its complete candidate and relevant inputs still match:

- `railyard:thermos`
- `codex review --base <base>`, with the model and effort that
Expand Down
10 changes: 10 additions & 0 deletions plugins/railyard/skills/deliver/references/ce-call-adapter.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,16 @@ execute the applicable procedure, and return to the requested delivery
boundary. Reading the file alone does not execute the workflow. A CE stage can
run in the current session or in a native subagent.

Before calling any stage or child that can publish, supply the
[whole-candidate review gate](../../../references/whole-candidate-review.md)
as a project publishing requirement, with the candidate receipt and an owner
for its refresh. This applies to LFG internal stages and every CE feedback
repair continuation. The publisher must stop before every push until the full
cumulative candidate is covered and the receipt matches; an earlier review or
a stage's completed result cannot retroactively satisfy the gate. If a selected
workflow cannot honor this requirement, stop before its external write and
report the integration limitation.

LFG owns its internal stages, including its CE review and babysitting loop;
do not duplicate them around it. Deliver consumes its result and continues the
authorized merge, release/deployment, and consumer-verification tail. For an
Expand Down
6 changes: 5 additions & 1 deletion plugins/railyard/skills/orchestrate/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,7 +81,11 @@ orchestration work and reclassifies the operation.
bounded handoff does not complete the caller's delivery. Use `compound-engineering:ce-commit-push-pr`
to create a PR or push user-requested commits to one, and `gh-stack` for
dependent PRs. Name an owner for integration and for each affected repository,
and publish only within the requested boundary.
and publish only within the requested boundary. Every publishing lane inherits
the [whole-candidate review gate](../../references/whole-candidate-review.md)
before every push, including CE/LFG repair continuations. The integration owner
checks the cumulative candidate receipt and maintains the delivery ledger;
a child's bounded review does not establish integration coverage.

CE alone owns review settlement and CI/PR monitoring. Reuse the lane's CE
watcher, including LFG's, and route reviewer findings to it; the orchestrator
Expand Down
Loading
Loading