Skip to content

One silent unauthenticated TCP connection can stall StayRTR's SSH RTR listener for unrelated clients #171

Description

@wangxin1115

Summary

During local and authorized testing of StayRTR, I found that a single silent unauthenticated TCP connection to the configured SSH RTR listener can prevent unrelated legitimate SSH RTR clients from receiving application-level SSH service.

The blocking client does not need to authenticate, does not need valid credentials, does not send an RTR query, and in the tested case sends exactly 0 application bytes.

The observed sequence is:

C0 connects to the SSH RTR listener
    |
    | sends 0 bytes
    | remains unauthenticated
    v
StayRTR blocks in the SSH handshake path
    |
    v
C1, an unrelated legitimate client,
can establish TCP but receives no SSH server identification
    |
    | C0 remains connected
    v
SSH RTR application-level service remains stalled
    |
    | C0 is intentionally disconnected
    v
The same pending C1 immediately begins receiving SSH service
and completes an RTR Reset Query through End Of Data

I first reproduced the behavior in a bounded 5-second test, and then performed an additional 10-minute corroboration using the same zero-byte unauthenticated connection.

The same C0 continuously maintained the cross-client SSH RTR listener stall for 603.053 seconds. At checkpoints through T+600, unrelated clients still received no SSH server identification. No automatic StayRTR disconnect was observed during that interval.

Plain RTR remained healthy throughout the test.

Tested revision:v0.6.4

Expected behavior

One incomplete pre-authentication SSH connection should not prevent unrelated later clients from receiving application-level service on the same SSH RTR listener.

Conceptually:

C0: incomplete pre-auth connection
        |
        +---- should remain isolated to C0

C1: unrelated legitimate client
        |
        +---- should still receive SSH server identification
              and continue normal SSH/RTR progress

Actual behavior

The actual behavior was:

C0:
  TCP connected
  application bytes sent = 0
  unauthenticated
  no valid SSH credentials used
  no RTR query sent
  remains connected

C1:
  TCP connection succeeds
  no SSH server identification received
  no application-level SSH progress

C0 disconnects

C1:
  immediately begins receiving SSH service
  completes SSH progress
  sends RTR Reset Query
  reaches End Of Data

This demonstrates a cross-client effect: the blocking connection does not merely stall itself; it prevents unrelated legitimate clients from receiving application-level service on the configured SSH RTR listener.

Source-level cause

The issue appears to come from the placement of the SSH handshake in the shared accept path.

StayRTR synchronously calls:

ssh.NewServerConn(...)

inside the shared SSH listener accept-loop callback.

The per-connection goroutine is only created after the SSH handshake completes successfully.

Conceptually, the current flow is equivalent to:

Accept C0
    |
    v
ssh.NewServerConn(C0)
    |
    | waits on peer-controlled SSH handshake progress
    |
    | C0 sends nothing
    v
blocked
    |
    | next Accept/service progress does not occur
    v
C1 remains without application-level SSH service

The important ordering is therefore:

Accept()
    ↓
synchronous ssh.NewServerConn()
    ↓
successful SSH handshake
    ↓
per-connection goroutine

rather than isolating the potentially blocking handshake before allowing the shared listener path to continue.

The reviewed local golang.org/x/crypto/ssh source performs peer-dependent blocking reads during SSH identification/version exchange, key exchange, and authentication.

I did not find a finite StayRTR application-level SSH handshake/authentication timeout on the tested path.

I also reviewed the official StayRTR documentation and did not find a documented server-side SSH handshake, authentication, or idle pre-authentication timeout for this path.

I am not claiming that no external timeout can ever exist. Deployment-specific firewalls, NAT devices, load balancers, operating-system behavior, TCP failure, or operator intervention may still bound connection lifetime.

Runtime evidence

Baseline

Without a blocking connection:

Legitimate SSH RTR client
    ↓
SSH service available
    ↓
SSH progress succeeds
    ↓
RTR Reset Query succeeds
    ↓
End Of Data received

Target

Exactly one C0 connection was used:

Blocking connections:     1
Application bytes sent:   0
Authenticated:            no
Valid credentials used:   no
RTR query sent:           no
Malformed RTR PDU used:   no

While C0 remained connected, an unrelated legitimate C1:

TCP connect:                succeeded
SSH server identification: not received
SSH application progress:  none

10-minute duration corroboration

The same C0 remained continuously connected for:

603.053 seconds

Cross-client SSH stall was confirmed at all of the following checkpoints:

Checkpoint Same C0 connected C0 authenticated Unrelated client receives SSH identification
T0 Yes No No
T+30 Yes No No
T+60 Yes No No
T+120 Yes No No
T+300 Yes No No
T+600 Yes No No

No automatic StayRTR disconnect was observed during the 10-minute interval.

The same StayRTR process remained alive throughout.

Recovery

After the blocking C0 connection was intentionally released:

C0 disconnects
    ↓
shared SSH listener progress resumes
    ↓
the same pending C1 receives SSH service
    ↓
C1 completes SSH progress
    ↓
C1 sends RTR Reset Query
    ↓
C1 reaches End Of Data

This recovery behavior strongly correlates the application-level stall with the single blocking C0 connection.

No StayRTR restart was required.

Plain RTR remained healthy

The issue was scoped to the configured SSH RTR listener in the tested runtime.

During the blocking condition:

SSH RTR listener:
  application-level service for later clients stalled

Plain RTR:
  remained healthy

StayRTR process:
  remained alive

Crash:
  not observed

OOM:
  not observed

Therefore, I am not claiming a whole-server DoS.

Minimal reproduction

A minimal reproduction is:

  1. Start StayRTR with a configured SSH RTR listener and a valid RTR payload.

  2. Verify that a legitimate SSH RTR client can normally:

    • connect;
    • receive SSH service;
    • complete SSH progress;
    • send an RTR Reset Query;
    • receive End Of Data.
  3. Open one raw TCP connection C0 to the SSH RTR listener.

  4. Send zero application bytes.

  5. Keep C0 connected and unauthenticated.

  6. Start an unrelated legitimate SSH RTR client C1.

  7. Observe that:

    • C1 can establish TCP;
    • C1 receives no SSH server identification;
    • C1 cannot make application-level SSH progress.
  8. Keep the same C0 connected.

  9. In my bounded test, the same cross-client stall remained present at:

    T0
    T+30s
    T+60s
    T+120s
    T+300s
    T+600s
    
  10. Disconnect C0.

  11. Observe that SSH RTR service resumes and the same pending C1 can continue normally through RTR End Of Data.

Why this matters

The directly confirmed issue is not that a slow or silent client merely delays its own session.

The cross-client effect is:

One silent unauthenticated TCP connection, sending zero application bytes, can prevent unrelated legitimate clients from receiving application-level service on StayRTR's configured SSH RTR listener.

The source-level mechanism is implementation-specific:

one pre-auth connection
        ↓
synchronous ssh.NewServerConn()
inside the shared accept path
        ↓
no next application-level SSH client progress
until the first connection is released

This differs from a traditional many-connection resource-exhaustion scenario.

The tested behavior required:

1 connection
0 application bytes
no authentication
no valid credentials
no RTR query
no malformed RTR PDU

Duration boundary

The runtime directly confirmed the stall for at least:

603.053 seconds

I am not claiming:

infinite duration
permanent outage
never times out in every deployment

The accurate statement is:

The same silent unauthenticated zero-byte TCP connection continuously maintained the cross-client SSH RTR listener stall for at least 10 minutes in the bounded local runtime, with no automatic StayRTR disconnect observed during that interval.

In addition, the reviewed StayRTR source and official documentation did not reveal a finite StayRTR server-side SSH handshake/authentication timeout on the tested path.

Possible fix direction

A possible fix direction would be to prevent one peer-controlled SSH handshake from serializing subsequent clients behind it.

Potential approaches include:

  1. Isolate ssh.NewServerConn() in a per-connection goroutine before returning to the shared accept path.

  2. Apply a finite SSH identification/handshake/authentication deadline.

  3. Combine per-connection isolation with explicit deadline and cleanup handling.

  4. Ensure one incomplete pre-authentication connection cannot prevent later connections from independently receiving SSH service.

I am not claiming that any one of these is necessarily the preferred final fix.

Suggested regression test

A minimal regression test could use:

P0 — normal baseline

No blocker
    ↓
C1 receives SSH server identification
    ↓
C1 completes normal SSH RTR service

T1 — one silent blocker

C0 connects
C0 sends 0 bytes
C0 remains unauthenticated

Then start unrelated legitimate C1.

Expected after a fix:

C1 should still receive SSH server identification
and continue independent SSH progress
without requiring C0 to disconnect.

The regression should verify:

  • one silent C0 does not block C1;
  • C1 receives SSH server identification within a bounded expected interval;
  • C1 can continue SSH/RTR progress;
  • plain RTR remains unaffected;
  • cleanup succeeds.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions