Skip to content

Perf coverage baseline: X7 + X8 macros, cold-start (issue #557), latency curve, bandwidth - #7

Merged
Tagar merged 1 commit into
masterfrom
perf-coverage-bandwidth-coldstart
May 21, 2026
Merged

Tagar merged 1 commit into
masterfrom
perf-coverage-bandwidth-coldstart

Conversation

@Tagar

@Tagar Tagar commented May 21, 2026 •

Copy link
Copy Markdown
Member

Merge this PR first. It establishes the baseline measurements (X7 recv + X8 send macros in CodSpeed, plus a comprehensive local test suite) that the open optimization PRs validate against.

Merge order

  1. Perf coverage baseline: X7 + X8 macros, cold-start (issue #557), latency curve, bandwidth #7 (this) — first. Adds X7/X8 macros + coverage tests. CodSpeed dashboard now tracks bytes recv + bytes send at 16K.
  2. perf: decode_bytearray (closes #570) + hot-path smart_decode shortcuts (credit @markjm #575) #5 (decode_bytearray fix). Recv side gets ~7.5× faster in microbench; X7-16k drops on the CodSpeed dashboard.
  3. perf: disable Nagle's algorithm on py4j socket connections (closes #516) #6 (Nagle disable, closes Performance Issue due to TCP Nagle's Algorithm write,write,read Situation py4j/py4j#516). On Linux runners, X7-16k and X8-16k drop dramatically (issue Performance Issue due to TCP Nagle's Algorithm write,write,read Situation py4j/py4j#516's 4.4 s/100 calls reproducer is exactly the X7-16k pattern).

CodSpeed scenarios (perf framework)

  • X7 — bytes round-trip RECV at 1k / 16k / 256k. Each call round-trips ByteBuffer.array() over the wire, routes through OUTPUT_CONVERTER[BYTES_TYPE] → decode_bytearray. Without X7 the decode_bytearray path is a measurement blind spot.
  • X8 — bytes SEND at 1k / 16k / 256k. Encode + Python→Java byte[] via BAOS.write(byte[], int, int). Exercises encode_bytearray and the send-side BufferedWriter overflow boundary.

Both X7-16k and X8-16k are registered for CodSpeed CI.

perf_coverage_test.py — run with pytest -v -s

ColdStartLatencyTest (issue py4j#557)

Three measurements, progressively excluding more of the cold path so a regression can be localized:

  • test_full_cold_start — fresh java subprocess + listening-socket wait + py4j handshake + first call. JVM init dominates (matches issue Slow first call to Java from Python py4j/py4j#557's reported ~3 s window). Caps <30 s.
  • test_py4j_handshake_under_500ms — JVM already listening; only JavaGateway() + first call timed. Isolates py4j's contribution (~1 ms locally).
  • test_first_vs_warm_call_ratio — first call vs warm median. Catches regressions in protocol-cache fill.

StringArgLatencyCurveTest

Per-call latency at string-argument sizes 16 / 256 / 4K / 16K / 64K bytes. Fills the gap between M2b (~5 B arg) and X7-class large payloads. Surfaces the Nagle knee around the BufferedWriter buffer boundary (~8 K) on Linux.

BandwidthSummaryTest

MB/s at 1K / 16K / 256K for both directions (recv via ByteBuffer.array, send via BAOS.write). Asserts >0.1 MB/s floor.

Local sample output (macOS loopback)

[full-cold-start] subprocess.Popen -> first call: 0.31 s
[py4j-handshake-only] gateway+first-call: 1.4 ms
[first-vs-warm] first: 196us  warm-median: 108.4us  ratio: 1.8x
[string-arg latency curve]
       16 B arg:    361.2 us/call
      256 B arg:    270.6 us/call
     4096 B arg:    339.6 us/call
    16384 B arg:    387.6 us/call
    65536 B arg:    462.9 us/call
[bandwidth summary]
       size    recv MB/s    send MB/s
      1024B          7.1          6.8
     16384B         60.3         79.3
    262144B         94.1        244.7

macOS loopback shows no Nagle penalty (Linux CodSpeed runners will); send is ~2.5× faster than recv at 256K because recv pays the base64-decode + list-comprehension cost in decode_bytearray that PR #5 fixes.

Co-authored-by: Isaac

@Tagar
Tagar force-pushed the perf-coverage-bandwidth-coldstart branch from 5854164 to af235ef Compare May 21, 2026 14:19
This PR establishes a measurement baseline that the open optimization
PRs (decode_bytearray #5, Nagle #6) can be validated against, plus
adds regression-guard coverage for several blind spots.

Designed to be merged FIRST so subsequent perf PRs see the new
scenarios as already-existing baselines in CodSpeed.

## CodSpeed scenarios (perf framework)

* **X7 — bytes round-trip RECV** (1k / 16k / 256k). Each call
  round-trips `ByteBuffer.array()` over the wire and routes through
  OUTPUT_CONVERTER[BYTES_TYPE] -> decode_bytearray. Without X7 the
  decode_bytearray path is a measurement blind spot (every prior
  macro returns int / list / void / callback). X7-16k joins the
  CodSpeed CI scenario list.

* **X8 — bytes SEND** (1k / 16k / 256k). Complements X7 on the
  encode_bytearray + Python -> Java direction via
  ByteArrayOutputStream.write(byte[], int, int). X8-16k joins
  the CodSpeed list.

Together X7-16k and X8-16k cover both halves of the byte-codec
bandwidth surface that's Nagle-sensitive on Linux.

## perf_coverage_test.py (run with `pytest -v -s`)

* **ColdStartLatencyTest (issue py4j#557)** — three measurements with
  progressively more of the cold path excluded so a regression can
  be localized:
  - `test_full_cold_start` — fresh `java` subprocess + listening-
    socket wait + py4j handshake + first call. Includes JVM class
    loading; issue py4j#557's reported window. Caps <30 s.
  - `test_py4j_handshake_under_500ms` — JVM is already listening;
    only `JavaGateway()` + first call is timed. Isolates py4j's
    own contribution (<500 ms).
  - `test_first_vs_warm_call_ratio` — first call vs warm median,
    catches regressions in protocol-cache fill.

* **StringArgLatencyCurveTest** — per-call latency at 16 / 256 /
  4K / 16K / 64K byte string args. Fills the gap between M2b
  (~5 B arg) and X7-class large payloads. Surfaces the Nagle
  knee around the BufferedWriter buffer boundary (~8 K) on Linux.

* **BandwidthSummaryTest** — MB/s at 1K / 16K / 256K for both
  directions (recv via ByteBuffer.array, send via BAOS.write).
  Reports the numbers; asserts >0.1 MB/s floor to catch
  protocol-layer regressions without flaking on hardware variance.

## Validation handle for the open PRs

* PR #5 (decode_bytearray) — `BandwidthSummaryTest` recv MB/s
  should jump substantially after the list-comprehension fix.
  X7-16k drops on the CodSpeed dashboard.
* PR #6 (Nagle) — `StringArgLatencyCurveTest` flattens across
  the 8 K boundary; X7-16k and X8-16k drop dramatically on
  Linux CodSpeed runners (issue py4j#516's 4.4s / 100 calls
  reproducer is exactly X7-16k's pattern).

## Local sample output (macOS loopback)

```
[full-cold-start] subprocess.Popen -> first call: 0.31 s
[py4j-handshake-only] gateway+first-call: 1.4 ms
[first-vs-warm] first: 196us  warm-median: 108.4us  ratio: 1.8x
[string-arg latency curve]
       16 B arg:    361.2 us/call
      256 B arg:    270.6 us/call
     4096 B arg:    339.6 us/call
    16384 B arg:    387.6 us/call
    65536 B arg:    462.9 us/call
[bandwidth summary]
       size    recv MB/s    send MB/s
      1024B          7.1          6.8
     16384B         60.3         79.3
    262144B         94.1        244.7
```

macOS loopback shows no Nagle penalty (Linux CodSpeed runners will);
send is ~2.5x faster than recv at 256K because recv pays the
base64-decode + list-comprehension cost in decode_bytearray that
PR #5 fixes.

Co-authored-by: Isaac
@Tagar
Tagar force-pushed the perf-coverage-bandwidth-coldstart branch from af235ef to 3438d78 Compare May 21, 2026 14:23
@Tagar Tagar changed the title Perf coverage: bandwidth, cold-start (issue #557), latency curve Perf coverage baseline: X7 + X8 macros, cold-start (issue #557), latency curve, bandwidth May 21, 2026
@Tagar
Tagar merged commit b7e694f into master May 21, 2026
58 checks passed
Tagar added a commit that referenced this pull request May 21, 2026
The Gradle buildscript declared `jcenter()` as the sole repository
for buildscript classpath resolution (spotless 1.3.2 + bnd-platform
1.3.0). JFrog sunset JCenter / Bintray in May 2021; the endpoint is
now flaky — CI runners whose gradle cache happens to have the
artifacts cached pass; fresh runners (notably windows-latest)
intermittently fail dependency resolution:

    Could not resolve com.diffplug.gradle.spotless:spotless:1.3.2.
    > Could not get resource
      'https://jcenter.bintray.com/com/diffplug/gradle/spotless/spotless/1.3.2/spotless-1.3.2.pom'.

Observed on the master push CI for #7 (1 of 56 cells failed:
Python 3.9 / Java 17 / windows-latest), and intermittently on prior
runs.

Replace with mavenCentral() + gradlePluginPortal(). Both classpath
artifacts are mirrored on Maven Central; gradlePluginPortal() is
added as a defensive fallback for any future plugin lookups.

Verified locally: clean gradle cache (`rm -rf
~/.gradle/caches/modules-2/files-2.1/com.diffplug.gradle.spotless`
and bnd-platform) + `./gradlew --refresh-dependencies classes
testClasses` → BUILD SUCCESSFUL with the new repositories.

Co-authored-by: Isaac
Tagar added a commit that referenced this pull request May 21, 2026
…03)"

This reverts commit 5026653 — removing FindBugs was overreach on
weak evidence:

* The 403 was on a SINGLE cell in PR #9's matrix (Python 3.9 /
  Java 8 / ubuntu-latest). Other cells in the SAME matrix run
  resolved `findbugs:3.0.+` successfully — proving the 403 is
  transient (likely Maven Central IP-throttling fresh runners),
  not a permanent policy change.
* Master CI on PRs #4 / #5 / #6 / #7 / #8 has been passing the
  same FindBugs resolution step reliably for months.
* Removing static analysis to "fix" a single flake degrades code
  quality on every future build.

The right defensive measures are already in this PR:
* shell-level retry around `./gradlew check && assemble`
* `shell: bash` for cross-platform consistency

If FindBugs ever does become permanently unavailable, that's the
moment to switch to SpotBugs — a real plugin migration, not a
panic delete.

Co-authored-by: Isaac
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Performance Issue due to TCP Nagle's Algorithm write,write,read Situation

1 participant