Found in a second reading of #1023. The safety net that PR adds fires only on a drain rejection; two common worker deaths produce no rejection.
Two gaps
-
A clean stream close resolves. BaseDuplexStream.drain is a for await, so if the far side sends StreamDoneSymbol or reader.return() is called, the drain resolves and onCriticalFailure never runs.
-
A Node worker that crashes after startup emits nothing. Verified in packages/kernel-node-runtime/src/kernel/PlatformServices.ts: the error and exit handlers are registered only for the pre-'online' window, and on 'online' both are removed —
127 worker.once('online', () => {
129 worker.removeAllListeners('error');
130 worker.removeAllListeners('exit');
So a worker that dies once running produces no message, no error and no stream end. The reader never yields, the drain never settles, the in-flight rpcClient.call is never rejected (RpcClient has no timeout), and the crank hangs with ctx.inCrank still true.
Why the hang is worse than one stuck caller
With inCrank stuck true, beginOutOfCrank never resolves — so terminateVat (via #trackFlux → withStoreOutOfCrank), inbound remote messages and waitForCrank all block forever. The operator cannot kill the vat to break the wedge. #runLoopState stays 'running', so getRunLoopStatus() reports a healthy kernel and onRunLoopDeath never fires.
#1023 moves worker launch into a crank (performVatRestart), which widens this: a relaunched worker that never answers initVat wedges the whole kernel, and restart is the natural operator response to a misbehaving vat.
Suggested fix
Keep an exit/error listener alive past 'online' and route it to onCriticalFailure, and/or put a deadline on sendVatCommand. #987 already tracks a delivery timeout; the listener removal is the separable half and is a small change.
Consequence for the current comments: VatManager's "when its worker dies … the RPC client has no timeout" describes a wider fix than the code delivers. What is actually closed is "when the stream reports a read error."
Found in a second reading of #1023. The safety net that PR adds fires only on a drain rejection; two common worker deaths produce no rejection.
Two gaps
A clean stream close resolves.
BaseDuplexStream.drainis afor await, so if the far side sendsStreamDoneSymbolorreader.return()is called, the drain resolves andonCriticalFailurenever runs.A Node worker that crashes after startup emits nothing. Verified in
packages/kernel-node-runtime/src/kernel/PlatformServices.ts: theerrorandexithandlers are registered only for the pre-'online'window, and on'online'both are removed —So a worker that dies once running produces no message, no error and no stream end. The reader never yields, the drain never settles, the in-flight
rpcClient.callis never rejected (RpcClienthas no timeout), and the crank hangs withctx.inCrankstill true.Why the hang is worse than one stuck caller
With
inCrankstuck true,beginOutOfCranknever resolves — soterminateVat(via#trackFlux→withStoreOutOfCrank), inbound remote messages andwaitForCrankall block forever. The operator cannot kill the vat to break the wedge.#runLoopStatestays'running', sogetRunLoopStatus()reports a healthy kernel andonRunLoopDeathnever fires.#1023 moves worker launch into a crank (
performVatRestart), which widens this: a relaunched worker that never answersinitVatwedges the whole kernel, and restart is the natural operator response to a misbehaving vat.Suggested fix
Keep an
exit/errorlistener alive past'online'and route it toonCriticalFailure, and/or put a deadline onsendVatCommand. #987 already tracks a delivery timeout; the listener removal is the separable half and is a small change.Consequence for the current comments:
VatManager's "when its worker dies … the RPC client has no timeout" describes a wider fix than the code delivers. What is actually closed is "when the stream reports a read error."