Summary
During local and authorized testing of StayRTR, I found that a single silent unauthenticated TCP connection to the configured SSH RTR listener can prevent unrelated legitimate SSH RTR clients from receiving application-level SSH service.
The blocking client does not need to authenticate, does not need valid credentials, does not send an RTR query, and in the tested case sends exactly 0 application bytes.
The observed sequence is:
C0 connects to the SSH RTR listener
|
| sends 0 bytes
| remains unauthenticated
v
StayRTR blocks in the SSH handshake path
|
v
C1, an unrelated legitimate client,
can establish TCP but receives no SSH server identification
|
| C0 remains connected
v
SSH RTR application-level service remains stalled
|
| C0 is intentionally disconnected
v
The same pending C1 immediately begins receiving SSH service
and completes an RTR Reset Query through End Of Data
I first reproduced the behavior in a bounded 5-second test, and then performed an additional 10-minute corroboration using the same zero-byte unauthenticated connection.
The same C0 continuously maintained the cross-client SSH RTR listener stall for 603.053 seconds. At checkpoints through T+600, unrelated clients still received no SSH server identification. No automatic StayRTR disconnect was observed during that interval.
Plain RTR remained healthy throughout the test.
Tested revision:v0.6.4
Expected behavior
One incomplete pre-authentication SSH connection should not prevent unrelated later clients from receiving application-level service on the same SSH RTR listener.
Conceptually:
C0: incomplete pre-auth connection
|
+---- should remain isolated to C0
C1: unrelated legitimate client
|
+---- should still receive SSH server identification
and continue normal SSH/RTR progress
Actual behavior
The actual behavior was:
C0:
TCP connected
application bytes sent = 0
unauthenticated
no valid SSH credentials used
no RTR query sent
remains connected
C1:
TCP connection succeeds
no SSH server identification received
no application-level SSH progress
C0 disconnects
C1:
immediately begins receiving SSH service
completes SSH progress
sends RTR Reset Query
reaches End Of Data
This demonstrates a cross-client effect: the blocking connection does not merely stall itself; it prevents unrelated legitimate clients from receiving application-level service on the configured SSH RTR listener.
Source-level cause
The issue appears to come from the placement of the SSH handshake in the shared accept path.
StayRTR synchronously calls:
inside the shared SSH listener accept-loop callback.
The per-connection goroutine is only created after the SSH handshake completes successfully.
Conceptually, the current flow is equivalent to:
Accept C0
|
v
ssh.NewServerConn(C0)
|
| waits on peer-controlled SSH handshake progress
|
| C0 sends nothing
v
blocked
|
| next Accept/service progress does not occur
v
C1 remains without application-level SSH service
The important ordering is therefore:
Accept()
↓
synchronous ssh.NewServerConn()
↓
successful SSH handshake
↓
per-connection goroutine
rather than isolating the potentially blocking handshake before allowing the shared listener path to continue.
The reviewed local golang.org/x/crypto/ssh source performs peer-dependent blocking reads during SSH identification/version exchange, key exchange, and authentication.
I did not find a finite StayRTR application-level SSH handshake/authentication timeout on the tested path.
I also reviewed the official StayRTR documentation and did not find a documented server-side SSH handshake, authentication, or idle pre-authentication timeout for this path.
I am not claiming that no external timeout can ever exist. Deployment-specific firewalls, NAT devices, load balancers, operating-system behavior, TCP failure, or operator intervention may still bound connection lifetime.
Runtime evidence
Baseline
Without a blocking connection:
Legitimate SSH RTR client
↓
SSH service available
↓
SSH progress succeeds
↓
RTR Reset Query succeeds
↓
End Of Data received
Target
Exactly one C0 connection was used:
Blocking connections: 1
Application bytes sent: 0
Authenticated: no
Valid credentials used: no
RTR query sent: no
Malformed RTR PDU used: no
While C0 remained connected, an unrelated legitimate C1:
TCP connect: succeeded
SSH server identification: not received
SSH application progress: none
10-minute duration corroboration
The same C0 remained continuously connected for:
Cross-client SSH stall was confirmed at all of the following checkpoints:
| Checkpoint |
Same C0 connected |
C0 authenticated |
Unrelated client receives SSH identification |
| T0 |
Yes |
No |
No |
| T+30 |
Yes |
No |
No |
| T+60 |
Yes |
No |
No |
| T+120 |
Yes |
No |
No |
| T+300 |
Yes |
No |
No |
| T+600 |
Yes |
No |
No |
No automatic StayRTR disconnect was observed during the 10-minute interval.
The same StayRTR process remained alive throughout.
Recovery
After the blocking C0 connection was intentionally released:
C0 disconnects
↓
shared SSH listener progress resumes
↓
the same pending C1 receives SSH service
↓
C1 completes SSH progress
↓
C1 sends RTR Reset Query
↓
C1 reaches End Of Data
This recovery behavior strongly correlates the application-level stall with the single blocking C0 connection.
No StayRTR restart was required.
Plain RTR remained healthy
The issue was scoped to the configured SSH RTR listener in the tested runtime.
During the blocking condition:
SSH RTR listener:
application-level service for later clients stalled
Plain RTR:
remained healthy
StayRTR process:
remained alive
Crash:
not observed
OOM:
not observed
Therefore, I am not claiming a whole-server DoS.
Minimal reproduction
A minimal reproduction is:
-
Start StayRTR with a configured SSH RTR listener and a valid RTR payload.
-
Verify that a legitimate SSH RTR client can normally:
- connect;
- receive SSH service;
- complete SSH progress;
- send an RTR Reset Query;
- receive End Of Data.
-
Open one raw TCP connection C0 to the SSH RTR listener.
-
Send zero application bytes.
-
Keep C0 connected and unauthenticated.
-
Start an unrelated legitimate SSH RTR client C1.
-
Observe that:
C1 can establish TCP;
C1 receives no SSH server identification;
C1 cannot make application-level SSH progress.
-
Keep the same C0 connected.
-
In my bounded test, the same cross-client stall remained present at:
T0
T+30s
T+60s
T+120s
T+300s
T+600s
-
Disconnect C0.
-
Observe that SSH RTR service resumes and the same pending C1 can continue normally through RTR End Of Data.
Why this matters
The directly confirmed issue is not that a slow or silent client merely delays its own session.
The cross-client effect is:
One silent unauthenticated TCP connection, sending zero application bytes, can prevent unrelated legitimate clients from receiving application-level service on StayRTR's configured SSH RTR listener.
The source-level mechanism is implementation-specific:
one pre-auth connection
↓
synchronous ssh.NewServerConn()
inside the shared accept path
↓
no next application-level SSH client progress
until the first connection is released
This differs from a traditional many-connection resource-exhaustion scenario.
The tested behavior required:
1 connection
0 application bytes
no authentication
no valid credentials
no RTR query
no malformed RTR PDU
Duration boundary
The runtime directly confirmed the stall for at least:
I am not claiming:
infinite duration
permanent outage
never times out in every deployment
The accurate statement is:
The same silent unauthenticated zero-byte TCP connection continuously maintained the cross-client SSH RTR listener stall for at least 10 minutes in the bounded local runtime, with no automatic StayRTR disconnect observed during that interval.
In addition, the reviewed StayRTR source and official documentation did not reveal a finite StayRTR server-side SSH handshake/authentication timeout on the tested path.
Possible fix direction
A possible fix direction would be to prevent one peer-controlled SSH handshake from serializing subsequent clients behind it.
Potential approaches include:
-
Isolate ssh.NewServerConn() in a per-connection goroutine before returning to the shared accept path.
-
Apply a finite SSH identification/handshake/authentication deadline.
-
Combine per-connection isolation with explicit deadline and cleanup handling.
-
Ensure one incomplete pre-authentication connection cannot prevent later connections from independently receiving SSH service.
I am not claiming that any one of these is necessarily the preferred final fix.
Suggested regression test
A minimal regression test could use:
P0 — normal baseline
No blocker
↓
C1 receives SSH server identification
↓
C1 completes normal SSH RTR service
T1 — one silent blocker
C0 connects
C0 sends 0 bytes
C0 remains unauthenticated
Then start unrelated legitimate C1.
Expected after a fix:
C1 should still receive SSH server identification
and continue independent SSH progress
without requiring C0 to disconnect.
The regression should verify:
- one silent C0 does not block C1;
- C1 receives SSH server identification within a bounded expected interval;
- C1 can continue SSH/RTR progress;
- plain RTR remains unaffected;
- cleanup succeeds.
Summary
During local and authorized testing of StayRTR, I found that a single silent unauthenticated TCP connection to the configured SSH RTR listener can prevent unrelated legitimate SSH RTR clients from receiving application-level SSH service.
The blocking client does not need to authenticate, does not need valid credentials, does not send an RTR query, and in the tested case sends exactly 0 application bytes.
The observed sequence is:
I first reproduced the behavior in a bounded 5-second test, and then performed an additional 10-minute corroboration using the same zero-byte unauthenticated connection.
The same C0 continuously maintained the cross-client SSH RTR listener stall for 603.053 seconds. At checkpoints through
T+600, unrelated clients still received no SSH server identification. No automatic StayRTR disconnect was observed during that interval.Plain RTR remained healthy throughout the test.
Tested revision:v0.6.4
Expected behavior
One incomplete pre-authentication SSH connection should not prevent unrelated later clients from receiving application-level service on the same SSH RTR listener.
Conceptually:
Actual behavior
The actual behavior was:
This demonstrates a cross-client effect: the blocking connection does not merely stall itself; it prevents unrelated legitimate clients from receiving application-level service on the configured SSH RTR listener.
Source-level cause
The issue appears to come from the placement of the SSH handshake in the shared accept path.
StayRTR synchronously calls:
inside the shared SSH listener accept-loop callback.
The per-connection goroutine is only created after the SSH handshake completes successfully.
Conceptually, the current flow is equivalent to:
The important ordering is therefore:
rather than isolating the potentially blocking handshake before allowing the shared listener path to continue.
The reviewed local
golang.org/x/crypto/sshsource performs peer-dependent blocking reads during SSH identification/version exchange, key exchange, and authentication.I did not find a finite StayRTR application-level SSH handshake/authentication timeout on the tested path.
I also reviewed the official StayRTR documentation and did not find a documented server-side SSH handshake, authentication, or idle pre-authentication timeout for this path.
I am not claiming that no external timeout can ever exist. Deployment-specific firewalls, NAT devices, load balancers, operating-system behavior, TCP failure, or operator intervention may still bound connection lifetime.
Runtime evidence
Baseline
Without a blocking connection:
Target
Exactly one C0 connection was used:
While C0 remained connected, an unrelated legitimate C1:
10-minute duration corroboration
The same C0 remained continuously connected for:
Cross-client SSH stall was confirmed at all of the following checkpoints:
No automatic StayRTR disconnect was observed during the 10-minute interval.
The same StayRTR process remained alive throughout.
Recovery
After the blocking C0 connection was intentionally released:
This recovery behavior strongly correlates the application-level stall with the single blocking C0 connection.
No StayRTR restart was required.
Plain RTR remained healthy
The issue was scoped to the configured SSH RTR listener in the tested runtime.
During the blocking condition:
Therefore, I am not claiming a whole-server DoS.
Minimal reproduction
A minimal reproduction is:
Start StayRTR with a configured SSH RTR listener and a valid RTR payload.
Verify that a legitimate SSH RTR client can normally:
Open one raw TCP connection
C0to the SSH RTR listener.Send zero application bytes.
Keep
C0connected and unauthenticated.Start an unrelated legitimate SSH RTR client
C1.Observe that:
C1can establish TCP;C1receives no SSH server identification;C1cannot make application-level SSH progress.Keep the same
C0connected.In my bounded test, the same cross-client stall remained present at:
Disconnect
C0.Observe that SSH RTR service resumes and the same pending
C1can continue normally through RTR End Of Data.Why this matters
The directly confirmed issue is not that a slow or silent client merely delays its own session.
The cross-client effect is:
The source-level mechanism is implementation-specific:
This differs from a traditional many-connection resource-exhaustion scenario.
The tested behavior required:
Duration boundary
The runtime directly confirmed the stall for at least:
I am not claiming:
The accurate statement is:
In addition, the reviewed StayRTR source and official documentation did not reveal a finite StayRTR server-side SSH handshake/authentication timeout on the tested path.
Possible fix direction
A possible fix direction would be to prevent one peer-controlled SSH handshake from serializing subsequent clients behind it.
Potential approaches include:
Isolate
ssh.NewServerConn()in a per-connection goroutine before returning to the shared accept path.Apply a finite SSH identification/handshake/authentication deadline.
Combine per-connection isolation with explicit deadline and cleanup handling.
Ensure one incomplete pre-authentication connection cannot prevent later connections from independently receiving SSH service.
I am not claiming that any one of these is necessarily the preferred final fix.
Suggested regression test
A minimal regression test could use:
P0 — normal baseline
T1 — one silent blocker
Then start unrelated legitimate
C1.Expected after a fix:
The regression should verify: