Skip to content

tracking: HAR capture and scrub -- controller queue, ordering, and in-flight work #368

Description

@MarkMichaelis

Tracking issue for the HAR capture-and-scrub work being run by a controller
session across two repos. Not a design document — every design decision
lives on the issue that owns it. This exists so the queue, its ORDERING, and the
reasons for that ordering survive a session crash or a handover.

  • UPSTREAM IntelliTect-Samples/IntelliSDLC.ai
  • CONSUMER MarkMichaelis/CodiwomplerSocialMedia

origin in this repo points at the old org name IntelliTect-Dev and redirects.
Always pass --repo IntelliTect-Samples/IntelliSDLC.ai.


The rule that orders the queue

A WIDENING change lands after the NARROWING change that constrains it.

IntelliTect-Samples/IntelliSDLC.ai#360 identifierFields alignment (narrowing)  ->  IntelliTect-Samples/web-api-discovery#57, IntelliTect-Samples/IntelliSDLC.ai#344 (widening)  ->  IntelliTect-Samples/web-api-discovery#55

Widening first puts a newly-visible population on the REPLACE path with nothing
suppressing the false positives among it — the 20-of-22 object-id corruption
measured in IntelliTect-Samples/web-api-discovery#56, applied to a fresh node class. This is beat 6 of
docs/designs/297-detector-predicates.md, and it is the single most important
constraint on this queue.


In flight

Item PR Author Reviewer State
#360 scrubber stops replacing at declared identifier fields #364 Opus Sonnet in review
#358 gate a request body that was REPLACED, not shortened #362 Opus Sonnet re-review after 2 accepted Important findings
#367 + #366 capture durability, and --describe required — Opus Sonnet authoring

#360 is the gate on the consumer audit. audit-scrub-drift.js's own
provisionalReasons() prints that its CLEAN and CORRUPTED verdicts are
provisional until this lands. UNADJUDICABLE is unaffected and is trustworthy
today.

One decision still open, on #364

When the same value sits at a declared identifier field in one place and a plain
field in another, the PR promotes it back and REPLACES it everywhere.
audit-scrub-drift.js's verdictFor calls that same case UNADJUDICABLE and
refuses to guess, "because picking one would be a guess with a repair behind
it." Beat 2 says an uncertain predicate on a replace path should fail toward a
MISS. Under review; the controller decides, not the author.


Done

  • #766's three orphaned Important findings (consumer): already filed as
    MarkMichaelis/CodiwomplerSocialMedia#809, #810, #811 on 2026-08-29 — two days
    before the issue asserting they were never filed. Verified still true of the
    code, then recorded on #862. No duplicates created.

Deferred, deliberately — not dropped

  • #355, IntelliTect-Samples/IntelliSDLC.ai#351 block on IntelliTect-Samples/web-api-discovery#59 Stage 10's provenance stamp. The contract is
    agreed and firm: log._scrubPolicy carrying
    { version, scrubbedAt, disabledIdentityClasses }, written LAST so presence
    means a completed scrub, replaced rather than appended, carrying no values.
    Refuse on PRESENCE, never on content. Stage 10 is unstarted and may be
    unowned.
  • #330, IntelliTect-Samples/IntelliSDLC.ai#344 — widening. Queued behind fix(web-api-discovery): the scrubber replaces at declared identifier fields the gate agreed to allow #360 by the ordering rule above.
  • #339, IntelliTect-Samples/web-api-discovery#58, IntelliTect-Samples/IntelliSDLC.ai#341, IntelliTect-Samples/web-api-discovery#54, IntelliTect-Samples/web-api-discovery#50, IntelliTect-Samples/web-api-discovery#48 — unassigned, none blocking.

Consumer queue — CodiwomplerSocialMedia#862

Held until the upstream items above merge, and in this order:

  1. Sync IntelliSDLC (roughly 70 commits behind). Deliberately not started yet
    — syncing before fix(web-api-discovery): the scrubber replaces at declared identifier fields the gate agreed to allow #360/fix(web-api-discovery): no gate detects a reference whose request body was wholly replaced by a placeholder #358/fix(web-api-discovery): recording from a worktree writes raw captures where worktree cleanup destroys them #367 merge would pull a tree without them and
    force a second sync.
  2. Run audit-scrub-drift.js over docs/har-reference/.
  3. Re-order the rest of #862 against what the audit finds. #862 says
    explicitly not to plan off its numbers without re-running.

Its defect 1 (a request body that is only a redaction sentinel) is being fixed
upstream as #358 on purpose, so the consumer inherits it by sync instead of
maintaining a second guard that drifts.

Non-scrub work on #862 — the Facebook and Instagram conformance reviews — does
not block on upstream but is held until the audit re-orders the list.


Incident: a raw capture was permanently destroyed

71 MB, 666 entries, recoveredFromSnapshot: true, no description. It sat in
a consuming project's worktree at .worktrees/<issue>-<name>/.har-captures/ and
was deleted by a routine git worktree remove from a concurrent session.
git worktree remove deletes outright — nothing reached the Recycle Bin and no
copy existed. Not recoverable.

Two upstream defects made it possible, both now filed:

Mitigation already in place, outside these repos: the full 8.42 GB store is
mirrored to C:\temp\har-capture\har-captures-mirror (257 files, verified
file-by-file), with a CAPTURE-INDEX.md carrying each capture's description,
size and SHA-256 — the hash being the durable link that survives a rename or
a move. A sweep of every worktree in both repos found no other at-risk capture.

The operational lesson, recorded because it is the reusable part: preserve
first, ask afterwards. The capture was identified as at-risk and then lost while
a question about how to LABEL it was outstanding. Copying is cheap and
reversible; deliberation is not free when another session holds the delete.


Invariants enforced on every worker

  • REVIEWER != AUTHOR, MECHANICALLY. A subagent spawned without an explicit
    model override INHERITS the authoring model, which is self-review in costume.
    Both models are assigned explicitly and the reviewer's model is recorded per
    PR. Opus authors, Sonnet reviews; the invariant is difference, not those names.
  • ABLATE EVERY ASSERTION. Break what it checks, watch it fail. Six tests in
    this subsystem shipped satisfied by something other than what they named —
    including one asserting typeof fieldTypeFor === 'function' with the message
    "so a project cannot extend it", which passed FIVE review rounds while nothing
    ever passed a policy. Mutation found all six; re-reading found none.
  • Falsifiers and guards labelled separately. A test that passes on arrival is
    a guard, not evidence.
  • A generator that cannot express a shape cannot falsify it. Seed it with the
    shapes ADJACENT to the one being fixed.
  • Measure the residue, not the delta. 2,841 -> 176 and 1,224 -> 94 are both
    true and both misled.
  • A review is ADVISORY. Triage each finding; accept and fix, or reject with a
    rationale validated against the code. Never auto-apply. Any accepted
    Critical/Important routes back to TDD and the reviewer re-reads before merge.
  • Never print a detected value — finding, log, error, test failure, PR body,
    dry-run report. Kind, class, key path, entry index, count, non-reversible
    fingerprint only.
  • Nothing under .har-captures/ is deleted, except by an operator's explicit
    -RemoveSource on a VERIFIED scrub.
  • No new command-line options without explicit human approval, including in
    autopilot. (fix(web-api-discovery): a capture can be recorded with no description, leaving it unidentifiable in a shared store #366 makes an EXISTING flag required, which was approved.)
  • Worktree + feature branch, never main. Never git add -A — a reviewer's
    scratch file was once swept into a commit that way and would have shipped to
    every consuming project.
  • Green before merge; rebase merges only.

Environment notes that cost time to rediscover

  • Node tests reach CI ONLY via a Pester wrapper in .github/agents/tests/;
    node-test-coverage.Tests.ps1 fails when one is missing. Local
    Invoke-Pester -Path .\.github\agents\tests\ is ~47 files, NARROWER than CI's
    .\.github at ~51 — always name which scope a result came from.
  • node --test <dir> fails on Node 26; pass an explicit file list.
  • gh on Windows garbles Unicode and collapses newlines in inline body text.
    Always --body-file.
  • Never verify against a worktree another agent owns — it may be mid-edit and
    git log will not warn you. Pin with git show <sha>:path.
  • Peer sessions work these repos concurrently. Check git log origin/main and
    the session list before dispatching; treat any stage list as a claim to
    verify, not a fact.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions