Background
The OpenHands agent-server uses bubus (via browser_use) for browser tool event handling. When a browser launch hangs on a C-level lock (e.g. missing browser deps on a headless server), the handler timeout fires after 30s but the worker thread is never freed. Zombie threads accumulate until the thread pool is exhausted — the server reports healthy (/health → 200) but new conversation creation stalls. Only a process restart clears it.
Observed in production: 21 zombie threads across 11 conversations. Tracked in OpenHands/software-agent-sdk#4598.
Bug
execute_handler's timeout cancels the asyncio task but CancelledError can't be delivered to a handler blocked in synchronous code (run_in_executor, C-level lock). The finally block waits only 0.1s then silently abandons the task, leaving a zombie thread:
if handler_task and not handler_task.done():
handler_task.cancel()
try:
await asyncio.wait_for(handler_task, timeout=0.1)
except (asyncio.CancelledError, TimeoutError):
pass
Reproduce
git clone https://github.com/browser-use/bubus.git && cd bubus && uv sync
Save as tests/test_handler_timeout_zombie_thread.py:
import asyncio, threading, time, pytest
from bubus import BaseEvent, EventBus
class BlockEvent(BaseEvent[str]):
pass
@pytest.mark.asyncio
async def test_timeout_does_not_leave_zombie_thread():
thread_started = threading.Event()
thread_should_stop = threading.Event()
def _block_sync():
thread_started.set()
thread_should_stop.wait(timeout=30)
return "done"
async def blocking_handler(event: BlockEvent) -> str:
loop = asyncio.get_event_loop()
return await loop.run_in_executor(None, _block_sync)
bus = EventBus(name="test_zombie")
bus.on(BlockEvent, blocking_handler)
bus._start()
bus.dispatch(BlockEvent(event_timeout=1.0))
try:
await asyncio.wait_for(bus.step(timeout=1.0), timeout=5.0)
except (asyncio.TimeoutError, TimeoutError, Exception):
pass
thread_should_stop.set()
time.sleep(1)
await bus.stop(timeout=1, clear=True)
assert thread_started.is_set()
zombies = [t for t in threading.enumerate()
if t is not threading.main_thread() and t.is_alive() and not t.daemon]
assert len(zombies) == 0, f"{len(zombies)} zombie thread(s) alive: {[t.name for t in zombies]}"
uv run pytest tests/test_handler_timeout_zombie_thread.py -xvs
FAILED AssertionError: 1 zombie thread(s) alive: ['asyncio_0']
Root Cause
Two code paths, same problem:
- Sync handlers run directly in the event loop thread (
handler(event)), blocking the loop — timeout never fires.
- Async handlers using
run_in_executor: CancelledError is queued but never delivered (no await point in the blocking thread).
Proposed Fixes
Background
The OpenHands agent-server uses
bubus(viabrowser_use) for browser tool event handling. When a browser launch hangs on a C-level lock (e.g. missing browser deps on a headless server), the handler timeout fires after 30s but the worker thread is never freed. Zombie threads accumulate until the thread pool is exhausted — the server reports healthy (/health→ 200) but new conversation creation stalls. Only a process restart clears it.Observed in production: 21 zombie threads across 11 conversations. Tracked in OpenHands/software-agent-sdk#4598.
Bug
execute_handler's timeout cancels the asyncio task butCancelledErrorcan't be delivered to a handler blocked in synchronous code (run_in_executor, C-level lock). Thefinallyblock waits only 0.1s then silently abandons the task, leaving a zombie thread:Reproduce
Save as
tests/test_handler_timeout_zombie_thread.py:Root Cause
Two code paths, same problem:
handler(event)), blocking the loop — timeout never fires.run_in_executor:CancelledErroris queued but never delivered (noawaitpoint in the blocking thread).Proposed Fixes
anyio.to_thread.run_sync(abandon_on_cancel=True)so the timeout can actually fire.