Repository navigation
perf: add Arrow-native encode and coordinate decode - #38
Conversation
|
This change is part of the following stack:
Change managed by git-spice. |
There was a problem hiding this comment.
Pull request overview
Adds Arrow-native bulk encode/decode APIs to reduce Python object materialization overhead in the hot path, extending the existing Arrow-focused approach used for WKB/EWKB output.
Changes:
- Added
encode_many_arrowproducing aLargeUtf8geohash column fromFloat64coordinate arrays. - Added
decode_many_arrow/decode_many_exactly_arrowreturning ArrowRecordBatches with decoded coordinate columns. - Extended Python typing stubs and added PyArrow-based tests to validate parity with list-based APIs, null handling, and input validation.
Reviewed changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.
| File | Description |
|---|---|
src/lib.rs |
Implements Arrow-native encode/decode functions and shared helpers; registers new PyO3 functions; adds Rust unit tests for the new Arrow column logic. |
tests/test_arrow_ops.py |
Adds PyArrow integration tests for the new Arrow encode/decode APIs (parity with list APIs, null propagation, and validation). |
geohash_polygon/__init__.pyi |
Updates the Python type stub to include the new Arrow API entry points and their expected shapes. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| let nulls = hashes.nulls(); | ||
| Ok((0..columns) | ||
| .map(|c| { | ||
| let values: Vec<f64> = (0..rows).map(|i| cells[i * columns + c]).collect(); | ||
| Float64Array::new(values.into(), nulls.clone()) | ||
| }) |
There was a problem hiding this comment.
Good call — fixed in f53ff32. The single buffer is now laid out column-major instead of row-major, so each output column is already contiguous and goes to Arrow as a ScalarBuffer window onto it. That removes both the gather pass and the per-column Vec, so peak memory is one copy rather than two.
Threads still split by row: each column is cut into 4096-row windows up front and the windows regrouped by index, which hands every task one disjoint slice of every column.
07381ea to
f53ff32
Compare
24025e5 to
0dcf5a8
Compare
f53ff32 to
a33d93a
Compare
0dcf5a8 to
1554d79
Compare
a33d93a to
d27b1ce
Compare
1554d79 to
98b8997
Compare
Completes the Arrow bulk codec started for WKB output. encode_many, decode_many and decode_many_exactly all spend most of their time building Python objects — 500k strings, or 500k tuples — rather than doing geohash work. N = 500,000 at precision 7, medians of 15 runs encode_many (list -> list[str]) 24.8 ms encode_many_arrow (arrow -> arrow) 3.4 ms 7.3x decode_many (list -> list[tuple]) 52.8 ms decode_many_arrow (arrow -> arrow) 1.9 ms 27.8x decode_many_exactly (list -> list[tuple]) 68.8 ms decode_many_exactly_arrow 2.8 ms 24.6x decode_many gains the most because a tuple per row is the most expensive thing any of these functions does. encode_many_arrow reuses the fixed-width column trick: every hash is exactly `precision` characters, so the buffer is allocated once and filled across threads. The decoders write one row-major cell buffer, so each row is touched by exactly one thread, then split it into columns. The decoders return a RecordBatch rather than an array of tuples — (lng, lat) and (lng, lat, lng_err, lat_err) — which is the shape a dataframe or a DuckDB table actually wants, and keeps each column contiguous. Nulls propagate: a null geohash gives null coordinates, and a row is null if either input coordinate is null. Precision is validated up front, so encode_many_arrow reports a bad precision as such rather than failing per row. The list-returning functions are untouched. Squashed with: - perf: decode coordinates straight into the output columns Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
d27b1ce to
f43b43f
Compare
Problem
Same boundary cost as #37, for the rest of the bulk API.
encode_many,decode_manyanddecode_many_exactlyall spend most of their time building Python objects — 500k strings, or 500k tuples — rather than doing geohash work.Fix
Arrow twins for all three.
encode_many_arrowreuses the fixed-width column trick: every hash is exactlyprecisioncharacters, so the buffer is allocated once and filled across threads. The decoders write one row-major cell buffer, so each row is touched by exactly one thread, then split it into columns.Benchmarks
N = 500,000 at precision 7, medians of 15 runs:
encode_manydecode_manydecode_many_exactlydecode_manygains most because a tuple per row is the most expensive thing any of these functions does.Notes
RecordBatchrather than an array of tuples —(lng, lat)and(lng, lat, lng_err, lat_err)— which is the shape a dataframe or a DuckDB table actually wants, and keeps each column contiguous.encode_many_arrowreports a bad precision as such rather than failing per row.