Skip to content

core: an unhandled consumer-delete timeout in disarmMembershipWatch kills the whole observer process #1047

Description

@davidfarah2003

What happened

cotal web --space netcup --server nats://100.98.140.41:4222 --detach on a Mac observing the space over a VPN (RTT ~90ms) crashed during startup with an uncaught TimeoutError, taking the whole dashboard process down:

TimeoutError: timeout
    at NatsConnectionImpl.request (@nats-io/nats-core/lib/nats.js:356:23)
    at ConsumerAPIImpl.delete (@nats-io/jetstream/lib/jsmconsumer_api.js:164:30)
    at PushConsumerImpl.delete (@nats-io/jetstream/lib/pushconsumer.js:364:31)
    at CotalEndpoint.disarmMembershipWatch (@cotal-ai/core/dist/endpoint.js:1912:54)
    at @cotal-ai/core/dist/endpoint.js:1811:32

cotal-ai 0.34.0, Node 26.7.0, 2026-08-30 ~02:25Z. The detach parent then reported web dashboard exited before becoming ready (pid 25579). A retry one minute later started fine, so the trigger is a transient: the JS-API CONSUMER.DELETE request behind disarmMembershipWatch timed out over the VPN (the same overlay had just recovered from a measured ~128s stall, #1045).

The defect

A consumer-delete during watch disarm is cleanup. Its timeout means the broker did not answer within the request deadline, not that the endpoint is unusable, and an ephemeral/idle consumer the broker still holds will age out on its own. Letting that rejection escape unhandled turns a slow link into a process kill for an observer that had nothing wrong with it.

The rejection escapes whatever calls disarmMembershipWatch at endpoint.js:1811 (dist line; the disarm path around rearmMembershipWatches), so nothing above it catches the promise.

Direction

Disarm/teardown of a membership watch should treat a delete timeout as best-effort cleanup: catch it, surface it as an endpoint error event (the observer already routes those to the console), and continue. The repo's no-fallbacks rule is about degrading function silently; this is the inverse case, cleanup of a resource the broker reaps anyway, and today it fails louder than the function it was cleaning up after.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:corePrimary affected area: core.bugSomething isn't workingseverity:highConfirmed high-impact defect or security issue.triage:confirmedReported defect reproduces, or requested non-bug gap is independently verified.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions