Skip to content

A vat worker that dies without a stream error produces no signal at all, and the crank hangs forever #1078

Description

@sirtimid

Found in a second reading of #1023. The safety net that PR adds fires only on a drain rejection; two common worker deaths produce no rejection.

Two gaps

  1. A clean stream close resolves. BaseDuplexStream.drain is a for await, so if the far side sends StreamDoneSymbol or reader.return() is called, the drain resolves and onCriticalFailure never runs.

  2. A Node worker that crashes after startup emits nothing. Verified in packages/kernel-node-runtime/src/kernel/PlatformServices.ts: the error and exit handlers are registered only for the pre-'online' window, and on 'online' both are removed —

    127  worker.once('online', () => {
    129    worker.removeAllListeners('error');
    130    worker.removeAllListeners('exit');
    

    So a worker that dies once running produces no message, no error and no stream end. The reader never yields, the drain never settles, the in-flight rpcClient.call is never rejected (RpcClient has no timeout), and the crank hangs with ctx.inCrank still true.

Why the hang is worse than one stuck caller

With inCrank stuck true, beginOutOfCrank never resolves — so terminateVat (via #trackFlux → withStoreOutOfCrank), inbound remote messages and waitForCrank all block forever. The operator cannot kill the vat to break the wedge. #runLoopState stays 'running', so getRunLoopStatus() reports a healthy kernel and onRunLoopDeath never fires.

#1023 moves worker launch into a crank (performVatRestart), which widens this: a relaunched worker that never answers initVat wedges the whole kernel, and restart is the natural operator response to a misbehaving vat.

Suggested fix

Keep an exit/error listener alive past 'online' and route it to onCriticalFailure, and/or put a deadline on sendVatCommand. #987 already tracks a delivery timeout; the listener removal is the separable half and is a small change.

Consequence for the current comments: VatManager's "when its worker dies … the RPC client has no timeout" describes a wider fix than the code delivers. What is actually closed is "when the stream reports a read error."

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions