-
Notifications
You must be signed in to change notification settings - Fork 451
RATIS-2629. Handle Netty based request asynchronously and improve exception handling #1538
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
spacemonkd
wants to merge
5
commits into
apache:master
Choose a base branch
from
spacemonkd:RATIS-2629
base: master
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
+259
−69
Open
Changes from 4 commits
Commits
Show all changes
5 commits
Select commit
Hold shift + click to select a range
4776b4c
RATIS-2629. Handle Netty based request asynchronously and improve exc…
spacemonkd 269a570
Mirror gRPC async handling
spacemonkd 9441369
Address CI issues
spacemonkd 1282dce
Switch to a fixed pool
spacemonkd 9a6720b
Address timeout duration, fix minor issue which can cause leak
spacemonkd File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
So, it's still a single worker pool (corePoolSize is 0). But instead of using TCP backpressure, we are pushing everything to the unlimited queue and creating pressure on the memory. Moreover, previously, Netty’s worker EventLoops handled connection shards independently, so a blocked request delayed only that shard’s RPCs. The new default single request worker queues all inbound RPCs, including heartbeats, so one slow request can delay heartbeats and trigger leader election.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Yes this was one issue which I missed and faced before (hence the test failure in flaky test suite).
corePoolSize=0+ an unbounded queue causesrequestExecutorto be single-threaded, and since handle() blocks until commit, this causes the timeouts.I have addressed this by switching to a fixed pool for now as the related change would increase LoC.
Filed https://issues.apache.org/jira/browse/RATIS-2637 for the improvement as a follow up.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Backpressure control should be a part of these changes. You have removed the one that was going through TCP flow control and don't provide any replacement. Another thing: we have ThreadPoolExecutor with corePoolSize=0 and unbounded queue. Let me quote "Java Concurrency in Practice":
[3] Developers are sometimes tempted to set the core size to zero so that the worker threads will eventually be torn down and therefore won't prevent the JVM from exiting, but this can cause some strange‐seeming behavior in thread pools that don't use a SynchronousQueue for their work queue (as newCachedThreadPool does). If the pool is already at the core size, ThreadPoolExecutor creates a new thread only if the work queue is full. So tasks submitted to a thread pool with a work queue that has any capacity and a core size of zero will not execute until the queue fills up, which is usually not what is desired.
So, it's still a single thread.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
So I have effectively switched back to the earlier behaviour for the time being.
Let's avoid making this patch bigger and address separately as right now this is effective the previous behaviour.