Perf coverage baseline: X7 + X8 macros, cold-start (issue #557), latency curve, bandwidth - #7
Merged
Merged
Conversation
Tagar
force-pushed
the
perf-coverage-bandwidth-coldstart
branch
from
May 21, 2026 14:19
5854164 to
af235ef
Compare
This PR establishes a measurement baseline that the open optimization PRs (decode_bytearray #5, Nagle #6) can be validated against, plus adds regression-guard coverage for several blind spots. Designed to be merged FIRST so subsequent perf PRs see the new scenarios as already-existing baselines in CodSpeed. ## CodSpeed scenarios (perf framework) * **X7 — bytes round-trip RECV** (1k / 16k / 256k). Each call round-trips `ByteBuffer.array()` over the wire and routes through OUTPUT_CONVERTER[BYTES_TYPE] -> decode_bytearray. Without X7 the decode_bytearray path is a measurement blind spot (every prior macro returns int / list / void / callback). X7-16k joins the CodSpeed CI scenario list. * **X8 — bytes SEND** (1k / 16k / 256k). Complements X7 on the encode_bytearray + Python -> Java direction via ByteArrayOutputStream.write(byte[], int, int). X8-16k joins the CodSpeed list. Together X7-16k and X8-16k cover both halves of the byte-codec bandwidth surface that's Nagle-sensitive on Linux. ## perf_coverage_test.py (run with `pytest -v -s`) * **ColdStartLatencyTest (issue py4j#557)** — three measurements with progressively more of the cold path excluded so a regression can be localized: - `test_full_cold_start` — fresh `java` subprocess + listening- socket wait + py4j handshake + first call. Includes JVM class loading; issue py4j#557's reported window. Caps <30 s. - `test_py4j_handshake_under_500ms` — JVM is already listening; only `JavaGateway()` + first call is timed. Isolates py4j's own contribution (<500 ms). - `test_first_vs_warm_call_ratio` — first call vs warm median, catches regressions in protocol-cache fill. * **StringArgLatencyCurveTest** — per-call latency at 16 / 256 / 4K / 16K / 64K byte string args. Fills the gap between M2b (~5 B arg) and X7-class large payloads. Surfaces the Nagle knee around the BufferedWriter buffer boundary (~8 K) on Linux. * **BandwidthSummaryTest** — MB/s at 1K / 16K / 256K for both directions (recv via ByteBuffer.array, send via BAOS.write). Reports the numbers; asserts >0.1 MB/s floor to catch protocol-layer regressions without flaking on hardware variance. ## Validation handle for the open PRs * PR #5 (decode_bytearray) — `BandwidthSummaryTest` recv MB/s should jump substantially after the list-comprehension fix. X7-16k drops on the CodSpeed dashboard. * PR #6 (Nagle) — `StringArgLatencyCurveTest` flattens across the 8 K boundary; X7-16k and X8-16k drop dramatically on Linux CodSpeed runners (issue py4j#516's 4.4s / 100 calls reproducer is exactly X7-16k's pattern). ## Local sample output (macOS loopback) ``` [full-cold-start] subprocess.Popen -> first call: 0.31 s [py4j-handshake-only] gateway+first-call: 1.4 ms [first-vs-warm] first: 196us warm-median: 108.4us ratio: 1.8x [string-arg latency curve] 16 B arg: 361.2 us/call 256 B arg: 270.6 us/call 4096 B arg: 339.6 us/call 16384 B arg: 387.6 us/call 65536 B arg: 462.9 us/call [bandwidth summary] size recv MB/s send MB/s 1024B 7.1 6.8 16384B 60.3 79.3 262144B 94.1 244.7 ``` macOS loopback shows no Nagle penalty (Linux CodSpeed runners will); send is ~2.5x faster than recv at 256K because recv pays the base64-decode + list-comprehension cost in decode_bytearray that PR #5 fixes. Co-authored-by: Isaac
Tagar
force-pushed
the
perf-coverage-bandwidth-coldstart
branch
from
May 21, 2026 14:23
af235ef to
3438d78
Compare
Tagar
added a commit
that referenced
this pull request
May 21, 2026
The Gradle buildscript declared `jcenter()` as the sole repository
for buildscript classpath resolution (spotless 1.3.2 + bnd-platform
1.3.0). JFrog sunset JCenter / Bintray in May 2021; the endpoint is
now flaky — CI runners whose gradle cache happens to have the
artifacts cached pass; fresh runners (notably windows-latest)
intermittently fail dependency resolution:
Could not resolve com.diffplug.gradle.spotless:spotless:1.3.2.
> Could not get resource
'https://jcenter.bintray.com/com/diffplug/gradle/spotless/spotless/1.3.2/spotless-1.3.2.pom'.
Observed on the master push CI for #7 (1 of 56 cells failed:
Python 3.9 / Java 17 / windows-latest), and intermittently on prior
runs.
Replace with mavenCentral() + gradlePluginPortal(). Both classpath
artifacts are mirrored on Maven Central; gradlePluginPortal() is
added as a defensive fallback for any future plugin lookups.
Verified locally: clean gradle cache (`rm -rf
~/.gradle/caches/modules-2/files-2.1/com.diffplug.gradle.spotless`
and bnd-platform) + `./gradlew --refresh-dependencies classes
testClasses` → BUILD SUCCESSFUL with the new repositories.
Co-authored-by: Isaac
Tagar
added a commit
that referenced
this pull request
May 21, 2026
…03)" This reverts commit 5026653 — removing FindBugs was overreach on weak evidence: * The 403 was on a SINGLE cell in PR #9's matrix (Python 3.9 / Java 8 / ubuntu-latest). Other cells in the SAME matrix run resolved `findbugs:3.0.+` successfully — proving the 403 is transient (likely Maven Central IP-throttling fresh runners), not a permanent policy change. * Master CI on PRs #4 / #5 / #6 / #7 / #8 has been passing the same FindBugs resolution step reliably for months. * Removing static analysis to "fix" a single flake degrades code quality on every future build. The right defensive measures are already in this PR: * shell-level retry around `./gradlew check && assemble` * `shell: bash` for cross-platform consistency If FindBugs ever does become permanently unavailable, that's the moment to switch to SpotBugs — a real plugin migration, not a panic delete. Co-authored-by: Isaac
This was referenced May 21, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merge this PR first. It establishes the baseline measurements (X7 recv + X8 send macros in CodSpeed, plus a comprehensive local test suite) that the open optimization PRs validate against.
Merge order
CodSpeed scenarios (perf framework)
ByteBuffer.array()over the wire, routes throughOUTPUT_CONVERTER[BYTES_TYPE]→decode_bytearray. Without X7 the decode_bytearray path is a measurement blind spot.Python→Java byte[]viaBAOS.write(byte[], int, int). Exercisesencode_bytearrayand the send-side BufferedWriter overflow boundary.Both X7-16k and X8-16k are registered for CodSpeed CI.
perf_coverage_test.py — run with
pytest -v -sColdStartLatencyTest (issue py4j#557)
Three measurements, progressively excluding more of the cold path so a regression can be localized:
test_full_cold_start— freshjavasubprocess + listening-socket wait + py4j handshake + first call. JVM init dominates (matches issue Slow first call to Java from Python py4j/py4j#557's reported ~3 s window). Caps <30 s.test_py4j_handshake_under_500ms— JVM already listening; onlyJavaGateway()+ first call timed. Isolates py4j's contribution (~1 ms locally).test_first_vs_warm_call_ratio— first call vs warm median. Catches regressions in protocol-cache fill.StringArgLatencyCurveTest
Per-call latency at string-argument sizes 16 / 256 / 4K / 16K / 64K bytes. Fills the gap between M2b (~5 B arg) and X7-class large payloads. Surfaces the Nagle knee around the BufferedWriter buffer boundary (~8 K) on Linux.
BandwidthSummaryTest
MB/s at 1K / 16K / 256K for both directions (recv via
ByteBuffer.array, send viaBAOS.write). Asserts >0.1 MB/s floor.Local sample output (macOS loopback)
macOS loopback shows no Nagle penalty (Linux CodSpeed runners will); send is ~2.5× faster than recv at 256K because recv pays the base64-decode + list-comprehension cost in
decode_bytearraythat PR #5 fixes.Co-authored-by: Isaac