Conversation
This was referenced Sep 15, 2026
sirtimid
added this pull request to stack #1108
September 15, 2026 22:38
sirtimid
force-pushed
the
sirtimid/restart-vat-run-queue-item
branch
from
October 1, 2026 12:17
1f752ce to
2b78778
Compare
Contributor
Coverage Report
File Coverage
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
sirtimid
force-pushed
the
sirtimid/restart-vat-run-queue-item
branch
from
October 1, 2026 16:38
2b78778 to
ec916be
Compare
Bugbot needs on-demand usage enabledBugbot uses usage-based billing for this team and requires on-demand usage to be enabled. A team admin can enable on-demand usage in the Cursor dashboard. |
sirtimid
force-pushed
the
sirtimid/restart-vat-run-queue-item
branch
from
October 2, 2026 09:03
59c175f to
3f26187
Compare
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit c9b7cb9. Configure here.
A restart keeps the vat's c-list while taking the vat itself out of the kernel's reach for as long as launching a worker and negotiating with it takes. Done where it is asked for, with the run loop free to run cranks throughout, a crank landing in that window reads a live vat as a dead one. `restartVat` now queues a `restartVat` run queue item and waits on a RAM waiter keyed by the vat; the run loop carries the request out in a crank of its own, where nothing else can reach the vat. Two callers for one vat share one item, because the first one's crank produces exactly what the second asked for. A request nobody is waiting for outlived the process that queued it, and is dropped. A relaunch that fails retires the vat and kills the worker it left behind, without throwing: the run loop's catch would roll the crank back, undoing those records and restoring the request, so every later start would replay the same failing restart. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Skipping the enqueue on the strength of someone else's waiter made a lost item permanent: a request rolled away by an aborting crank left a waiter nothing would consume, and every later request for that vat then skipped the queue too and hung with a healthy run loop. Every request queues an item; the crank that arrives first settles the whole list and the rest are dropped as stale, which the same path already did for an item that outlived its process. The enqueue also moves ahead of the waiter. A run loop already dead rejects the waiter the moment it is registered and then refuses the enqueue, so the throw left by way of the `try` and nothing ever awaited the rejected promise — an unhandled rejection, which Node's default turns into a dead process. An old worker whose channel will not close no longer costs the vat its relaunch: the handle is off the books and the worker killed either way, so retiring it for that was retiring a vat that was merely untidy. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e test #1084 landed with a router built positionally, so the logger took the slot the new parameter added and the `already settled` assertions saw nothing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The crank completed without aborting, so whatever the failed `initVat` buffered was flushed after the vat had been retired: sends from a vat that no longer existed went out, and a notify addressed to it could kill the run loop. `performVatRestart` now reports an abort and a termination, the shape a failed delivery already has. The run loop rolls the crank back and then retires the vat; the rollback restores the request, which finds no waiters and is dropped. Callers are answered once the crank ends, so they wake to a committed termination. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…not say Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…nk ends Written into an open crank, the request was rolled back with it if the crank aborted, and restartVat waited on an item that no longer existed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A caller now wakes after the crank ends, by which time the next crank may be another restart that has already taken the handle away, and looking the vat up then threw VatNotFoundError for a restart that succeeded. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
sirtimid
force-pushed
the
sirtimid/restart-vat-run-queue-item
branch
from
October 2, 2026 18:43
c9b7cb9 to
b2477a9
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Stacked on #1093. Diff:
git diff sirtimid/retire-vat-synchronously...sirtimid/restart-vat-run-queue-item.A restart takes a vat out of the kernel's reach while a new worker launches. Done outside the run loop, a crank landing in that window reads a live vat as dead.
restartVatnow queues arestartVatrun queue item and waits, and the run loop carries it out in a crank of its own, like SwingSet'supgrade-vat.{ abort, terminate }, like a failed delivery. The run loop rolls back what the failedinitVatbuffered, then retires the vat. Callers are answered once the crank ends.Known gaps
restartVat,queueMessage) during a crank that aborts are rolled away, and their callers wait until the run loop dies. This is pre-existing for any aborting delivery; a failed restart is a longer such crank, since it waits out the handshake. It closes when control-plane writes move onto the run loop.reset,clearStorageandstopcan abandon a queued restart request. Left for the PR that stops the run loop first.Testing
VatManager.restart-crank.test.ts: real run loop and store. A failed relaunch's buffered send or notify is discarded and the vat terminated. Fails without the abort.VatManager.test.ts: waiter coalescing, stale and leftover items, a vat gone before its crank, relaunch failure, and an old worker that will not shut down cleanly.kernel-test: a vat restarted twice reachesstart count: 3in baggage.@metamask/ocap-kernelsuite green.Carries the findings from #1065, #1068 and #1081.
🤖 Generated with Claude Code
Note
Medium Risk
Changes vat lifecycle and run-loop crank/commit semantics; failed restarts now terminate vats, but behavior is heavily tested and aligns with existing abort/terminate paths.
Overview
restartVatno longer stops and relaunches a worker inline. It enqueues a newrestartVatrun-queue item, registers per-vat waiters, and resolves only after the run loop finishes the restart crank—so no delivery can see a persisted vat with no live worker in between.The run loop routes that item through
KernelRoutertoVatManager.performVatRestart, which terminates the old worker, relaunches from config (including vats already between workers), and coalesces multiple concurrent restart callers on one crank.KernelQueueaddsenqueueRestartVat, defers enqueue while a crank is open (so an aborted crank does not roll away the request), andonRunLoopDeathrejects restart waiters if the loop dies first.If relaunch fails,
performVatRestartreturns{ abort, terminate }so the crank rolls back bufferedinitVatoutput before the vat is retired—matching failed-delivery semantics instead of leaving a persisted vat with no worker.Integration and unit coverage include a
kernel-testdouble-restart baggage check,VatManager.restart-crank.test.tswith a real store/queue, and expandedVatManager/KernelQueuetests;Kernel.testnow asserts enqueue + run-loop-death behavior rather than full inline restart.Reviewed by Cursor Bugbot for commit b2477a9. Bugbot is set up for automated code reviews on this repo. Configure here.