Skip to content

perf: add Arrow-native encode and coordinate decode - #38

Merged
mailletf merged 1 commit into
perf/arrow-wkb-outputfrom
perf/arrow-codec-coords
Sep 16, 2026
Merged

mailletf merged 1 commit into
perf/arrow-wkb-outputfrom
perf/arrow-codec-coords

Conversation

@mailletf

@mailletf mailletf commented Aug 25, 2026 •

Copy link
Copy Markdown
Member

Problem

Same boundary cost as #37, for the rest of the bulk API. encode_many, decode_many and decode_many_exactly all spend most of their time building Python objects — 500k strings, or 500k tuples — rather than doing geohash work.

Fix

Arrow twins for all three. encode_many_arrow reuses the fixed-width column trick: every hash is exactly precision characters, so the buffer is allocated once and filled across threads. The decoders write one row-major cell buffer, so each row is touched by exactly one thread, then split it into columns.

Benchmarks

N = 500,000 at precision 7, medians of 15 runs:

list API Arrow
encode_many 24.8 ms 3.4 ms 7.3×
decode_many 52.8 ms 1.9 ms 27.8×
decode_many_exactly 68.8 ms 2.8 ms 24.6×

decode_many gains most because a tuple per row is the most expensive thing any of these functions does.

Notes

  • The decoders return a RecordBatch rather than an array of tuples — (lng, lat) and (lng, lat, lng_err, lat_err) — which is the shape a dataframe or a DuckDB table actually wants, and keeps each column contiguous.
  • Nulls propagate: a null geohash gives null coordinates, and a row is null if either input coordinate is null.
  • Precision is validated up front, so encode_many_arrow reports a bad precision as such rather than failing per row.
  • The list-returning functions are untouched.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds Arrow-native bulk encode/decode APIs to reduce Python object materialization overhead in the hot path, extending the existing Arrow-focused approach used for WKB/EWKB output.

Changes:

  • Added encode_many_arrow producing a LargeUtf8 geohash column from Float64 coordinate arrays.
  • Added decode_many_arrow / decode_many_exactly_arrow returning Arrow RecordBatches with decoded coordinate columns.
  • Extended Python typing stubs and added PyArrow-based tests to validate parity with list-based APIs, null handling, and input validation.

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

File Description
src/lib.rs Implements Arrow-native encode/decode functions and shared helpers; registers new PyO3 functions; adds Rust unit tests for the new Arrow column logic.
tests/test_arrow_ops.py Adds PyArrow integration tests for the new Arrow encode/decode APIs (parity with list APIs, null propagation, and validation).
geohash_polygon/__init__.pyi Updates the Python type stub to include the new Arrow API entry points and their expected shapes.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/lib.rs Outdated
Comment on lines +828 to +833
let nulls = hashes.nulls();
Ok((0..columns)
.map(|c| {
let values: Vec<f64> = (0..rows).map(|i| cells[i * columns + c]).collect();
Float64Array::new(values.into(), nulls.clone())
})

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good call — fixed in f53ff32. The single buffer is now laid out column-major instead of row-major, so each output column is already contiguous and goes to Arrow as a ScalarBuffer window onto it. That removes both the gather pass and the per-column Vec, so peak memory is one copy rather than two.

Threads still split by row: each column is cut into 4096-row windows up front and the windows regrouped by index, which hands every task one disjoint slice of every column.

@mailletf
mailletf force-pushed the perf/arrow-codec-coords branch from 07381ea to f53ff32 Compare August 25, 2026 14:45
@mailletf
mailletf force-pushed the perf/arrow-wkb-output branch from 24025e5 to 0dcf5a8 Compare August 25, 2026 14:45
@mailletf
mailletf force-pushed the perf/arrow-codec-coords branch from f53ff32 to a33d93a Compare September 1, 2026 18:19
@mailletf
mailletf force-pushed the perf/arrow-wkb-output branch from 0dcf5a8 to 1554d79 Compare September 1, 2026 18:19
@mailletf
mailletf force-pushed the perf/arrow-codec-coords branch from a33d93a to d27b1ce Compare September 1, 2026 18:33
@mailletf
mailletf force-pushed the perf/arrow-wkb-output branch from 1554d79 to 98b8997 Compare September 1, 2026 18:33
Completes the Arrow bulk codec started for WKB output. encode_many,
decode_many and decode_many_exactly all spend most of their time building
Python objects — 500k strings, or 500k tuples — rather than doing geohash
work.

  N = 500,000 at precision 7, medians of 15 runs

  encode_many          (list -> list[str])       24.8 ms
  encode_many_arrow    (arrow -> arrow)           3.4 ms    7.3x
  decode_many          (list -> list[tuple])     52.8 ms
  decode_many_arrow    (arrow -> arrow)           1.9 ms   27.8x
  decode_many_exactly  (list -> list[tuple])     68.8 ms
  decode_many_exactly_arrow                       2.8 ms   24.6x

decode_many gains the most because a tuple per row is the most expensive
thing any of these functions does.

encode_many_arrow reuses the fixed-width column trick: every hash is exactly
`precision` characters, so the buffer is allocated once and filled across
threads. The decoders write one row-major cell buffer, so each row is
touched by exactly one thread, then split it into columns.

The decoders return a RecordBatch rather than an array of tuples — (lng,
lat) and (lng, lat, lng_err, lat_err) — which is the shape a dataframe or a
DuckDB table actually wants, and keeps each column contiguous.

Nulls propagate: a null geohash gives null coordinates, and a row is null if
either input coordinate is null. Precision is validated up front, so
encode_many_arrow reports a bad precision as such rather than failing per
row.

The list-returning functions are untouched.

Squashed with:
- perf: decode coordinates straight into the output columns

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@mailletf
mailletf force-pushed the perf/arrow-codec-coords branch from d27b1ce to f43b43f Compare September 1, 2026 18:54
@mailletf
mailletf merged commit 69cc202 into perf/arrow-wkb-output Sep 16, 2026
16 checks passed
@mailletf
mailletf deleted the perf/arrow-codec-coords branch September 16, 2026 20:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants