Skip to content

feat(serialize): opt-in bounded-memory zstd encoder profile (#79) - #80

Merged
oss-amikos merged 8 commits into
mainfrom
claude/gallant-allen-72pzmi
Sep 27, 2026
Merged

oss-amikos merged 8 commits into
mainfrom
claude/gallant-allen-72pzmi

Conversation

@oss-amikos

@oss-amikos oss-amikos commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Fixes #79.

Encode reuses one shared zstd encoder per compression mode, and each encoder keeps one worker per GOMAXPROCS. A level-15 worker holds about 36 MB of match tables, so a 16-CPU process retained about 560 MB after its first encode. Until now the only way to shrink that was a lower compression level.

This PR adds an opt-in bounded-memory encoder profile. It keeps a single worker at the same compression level, and the output bytes are identical.

Status: Verified ✓ (quick-260925-b7m, hold-out expectations E1–E11 all pass)

Changes

Encoder profile

  • New EncoderProfile type: EncoderProfileDefault (today's behaviour) and EncoderProfileBoundedMemory.
  • The bounded profile is WithEncoderConcurrency(1) + WithLowerEncoderMem(true). The window size is unchanged, so output is byte-identical to the default profile at every level.
  • The shared encoder cache is keyed by (zstd mode, profile), so the two profiles never share an instance. It holds at most 4 × 2 entries.

Selecting it

  • WithEncoderProfile(p) is a ConfigOption. Every encode path that reads the index config uses it: Encode, WriteSidecar, EncodeToMetadata, S3 sidecars.
  • WithEncodeProfile(p) is an EncodeOption that overrides the config per call, including an explicit EncoderProfileDefault.
  • The profile is runtime-only and never serialized. A decoded index reads EncoderProfileDefault. Wire format stays v11.
  • Unknown profile values are rejected by the option, by config validation and by the encode call.

CLI

  • gin-index build, extract and experiment gain -low-memory. extract passes the profile per call, because a decoded index always carries the default.

Benchmarks and docs

  • BenchmarkEncoderProfile covers both profiles × levels 15 and 3 × small and high-cardinality fixtures × cold and repeated runs. It reports retained MB, allocations, time and compressed size.
  • make bench-encoder-profile runs it at GOMAXPROCS 1, 4 and 16. Results are in docs/encoder-profile-benchmarks.md.
  • README has a new "Encoder memory profile" subsection with the trade-off. CHANGELOG has an Unreleased entry, and CLAUDE.md is updated.

Retained heap after the first level-15 encode:

GOMAXPROCS Default Bounded
1 51 MB 42 MB
4 153 MB 42 MB
16 561 MB 42 MB

Review round

A high-effort self-review of the first commit found five issues; all are fixed in the second commit:

  • The byte-identity test now covers a payload of more than 128 KB, so the multi-block path where low-memory buffers apply is tested.
  • Forced GCs are out of the cold benchmark's timed region, and the cold timings were re-measured.
  • -low-memory is added to extract and experiment -o.
  • The cache-key and eviction helpers are deduplicated.
  • The benchmark serializes each payload once.

Key files: serialize.go, gin.go, cmd/gin-index/main.go, cmd/gin-index/experiment.go, serialize_profile_test.go, serialize_concurrency_test.go, benchmark_test.go, Makefile, README.md, CHANGELOG.md, docs/encoder-profile-benchmarks.md

Requirements Addressed

  • ISSUE-79: the bounded profile's level-15 output decodes to the same index, with decoder limits unchanged.
  • ISSUE-79: mixed-profile and mixed-level concurrent encodes are race-free.
  • ISSUE-79: benchmarks run at GOMAXPROCS 1, 4 and 16, cold and repeated, at levels 15 and 3, on small and high-cardinality fixtures.
  • ISSUE-79: the memory, time and size trade-off is documented.

Verification

  • go build ./... and go vet ./... pass
  • make lint: 0 issues
  • go test -race -short ./... passes in all packages
  • Full suite via gotestsum (./... ./testdata/phase20): 1199 tests pass, 1 pre-existing skip (missing testdata/test.parquet)
  • make bench-encoder-profile ran at GOMAXPROCS 1, 4 and 16
  • Note: make test exits non-zero in the authoring container only because that Go toolchain lacks go tool covdata for -coverprofile. No tests failed. CI's runner has the tool.

Key Decisions

  • The per-call option overrides the config option, mirroring WithEncodeSignals / WithSignals.
  • The profile is a named enum, not a worker count. That keeps the cache bounded and keeps klauspost's pool model out of the public API.
  • Bounded encoders are cached like default ones. About 42 MB is the level-15 floor, so it is paid once rather than per call.
  • The default profile is unchanged. Capping default workers would silently reduce parallel-encode throughput on large hosts. It is logged as a deferred follow-up with these numbers.
  • The window size is left alone, so byte-identical output holds as a tested invariant and the profile is purely a memory knob.

Review round 2 (local multi-agent review, 15 findings)

Summary

 WriteSidecar(path, idx)
 EncodeToMetadata(idx)
 RebuildWithIndex(path, idx)
 S3Client.WriteSidecar(ctx?, bucket, key, idx)
+  ...opts ...EncodeOption   // e.g. WithEncodeProfile; a decoded index always carries Default
  • New TestEncoderProfileBoundedUsesSingleWorker reflects into the zstd encoder and asserts one worker and low-memory buffers. Before this, dropping both options passed every test.
  • experiment -low-memory without -o now warns instead of doing nothing.
  • Memory prose aligned with the benchmark tables: about 34 MB per additional worker, about 560 MB at 16 CPUs, about 42 MB bounded.
  • Tautological wire-format check replaced with a byte-equality check across profiles. Mixed-profile concurrency test regained its decode-and-re-encode check.
  • Doc comments, README cache-key and level-3 wording, CHANGELOG, and the klauspost pin in CLAUDE.md corrected.

Evidence

  • Before: removing WithEncoderConcurrency(1) / WithLowerEncoderMem(true) from newZstdEncoder left the suite green.
    After: TestEncoderProfileBoundedUsesSingleWorker fails on channel capacity or lowMem.
  • go test -race -short . ./cmd/... passes. make lint reports 0 issues (Go 1.25.5 toolchain).

Merge Danger

Door: two-way. Trailing variadic parameters are additive; wire format stays v11.

Blast Radius: small. Only callers that store these functions in a typed variable with the old signature would need a change.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Dh45hVm1GJoj9sTGDah84X


Generated by Claude Code

claude and others added 8 commits September 25, 2026 07:16
The shared zstd encoder keeps one worker per GOMAXPROCS and each level-15
worker retains about 36 MB, so a 16-CPU host held roughly 560 MB after its
first encode with no way to shrink it short of lowering the level.

Add EncoderProfile with EncoderProfileDefault and
EncoderProfileBoundedMemory. The bounded profile keeps a single worker with
zstd's lower-memory buffers (about 42 MB at level 15 regardless of
GOMAXPROCS) and produces byte-identical output at every level. Select it per
call with WithEncodeProfile (EncodeOption) or on the config with
WithEncoderProfile (ConfigOption) so WriteSidecar, EncodeToMetadata and S3
sidecars use it; the per-call option wins. The encoder cache is keyed by
zstd mode and profile so the two profiles never share an instance. The
profile is runtime-only and not serialized; the default profile and wire
format v11 are unchanged.

gin-index build gains -low-memory. BenchmarkEncoderProfile and
make bench-encoder-profile report retained memory, allocations, time and
compressed size at GOMAXPROCS 1, 4 and 16; results and the trade-off are
documented in docs/encoder-profile-benchmarks.md and the README.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dh45hVm1GJoj9sTGDah84X
- Test byte-identical output on a multi-block (>128 KB) payload too, so
  the path where WithLowerEncoderMem changes buffer sizing is covered.
- Keep forced GCs out of the cold benchmark's timed region and refresh the
  cold timings in docs/encoder-profile-benchmarks.md.
- Add -low-memory to gin-index extract and experiment; extract passes the
  profile per call because a decoded index always carries the default.
- Share one cache-key helper and one eviction helper across tests and
  benchmarks; serialize each benchmark payload once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dh45hVm1GJoj9sTGDah84X
Close the encoder-profile review gaps on the bounded-memory zstd encoder
feature (#79):

- I1: WriteSidecar, EncodeToMetadata, RebuildWithIndex, and
  S3Client.WriteSidecar/WriteSidecarContext accept a trailing
  opts ...EncodeOption so a caller can force a profile even when idx.Config
  carries the default (post-decode) profile.
- I2: add TestEncoderProfileBoundedUsesSingleWorker, asserting via reflection
  that the bounded profile configures a single zstd worker with lowMem=true.
- I3/S10: correct EncoderProfileDefault/BoundedMemory doc comments to the
  measured numbers (~34 MB per additional worker, ~560 MB at GOMAXPROCS=16,
  ~42 MB bounded) and note the worker pool size is fixed at first encode.
- S2: replace the tautological Header.Version check in
  TestEncoderProfileConfigReachesAllEncodePaths with a byte-equality check
  proving the serialized config payload does not carry the profile.
- S5: wrap zstd encoder construction errors with level and profile context.
- S6: use key.profile (not the shadowed profile param) as the single cache
  key identity source in sharedZstdEncoder.
- S8: rewrite TestEncodeDecodeConcurrentMixedProfiles's doc comment and
  restore the decode-then-re-encode correctness check.
- S9: clarify the zstdEncoderKey doc comment on who Closes cached encoders.
- S11: improve the unknown-profile validation error message.

Also updates the S3Client structural-compatibility test in
boundary_observability_test.go for the new WriteSidecar/WriteSidecarContext
variadic signature.

Ref #79
…ts (#79)

I4: gin-index experiment --low-memory without -o was a silent no-op; print
"Warning: -low-memory has no effect without -o" to stderr instead.

Extracts experimentGINConfig as a standalone, independently testable helper
(mirroring buildGINConfig in main.go) instead of building the config inline
in runExperiment.

S1: add TestExperimentGINConfigLowMemorySelectsBoundedProfile (config-level
profile assertion for the experiment command) and
TestRunBuildLowMemoryProducesDecodableIndex (end-to-end runBuild -low-memory
coverage for both sidecar and -embed modes). Amend the doc comments on
TestRunExtractLowMemoryWritesSameBytes and
TestRunExperimentLowMemoryWritesSameSidecarBytes to state they guard output
bytes and exit code only.

Ref #79
I3: align every prose memory figure with the measured benchmark tables --
"about 34 MB per additional worker", "about 560 MB at GOMAXPROCS=16", and
"about 42 MB bounded" -- across README.md, CHANGELOG.md,
docs/encoder-profile-benchmarks.md, and the -low-memory CLI flag help in
cmd/gin-index/main.go (serialize.go's doc comments were already corrected).

S4: clarify that Encode's zstd encoder cache is keyed by compression mode
and profile (not the raw numeric level), and scope the "level 3 is larger"
claim to the high-cardinality fixture, since the small fixture is smaller at
level 3.

S7: bump the klauspost/compress version noted in CLAUDE.md from v1.18.3 to
v1.19.2 to match go.mod.

Adds a footnote to docs/encoder-profile-benchmarks.md's retained-memory table
naming the highcard fixture as the L3 column source and noting level-3
retention is sensitive to first-payload size.

Ref #79
Vary the fixture status literal in TestRunBuildLowMemoryProducesDecodableIndex
so it does not push an existing string past golangci-lint's goconst
threshold, and document+suppress the unparam finding on
sharedZstdEncoderCached: every current call site checks the bounded-memory
profile, but the parameter mirrors evictSharedZstdEncoder's signature for a
future default-profile assertion.

Ref #79
- CHANGELOG: note trailing EncodeOption on WriteSidecar, EncodeToMetadata,
  RebuildWithIndex and S3Client sidecar helpers
- README: per-call helper example; label level-3 retention by fixture
- benchmarks doc: 'per additional worker'; reconcile level-3 bullet with tables
- serialize_profile_test: strengthen per-call override test with a positive
  default-profile assertion and drop the unparam nolint; fix stale comments
- main_test: correct which tests cover extract -low-memory wiring
- serialize_concurrency_test: reference identifiers instead of line numbers
@oss-amikos
oss-amikos merged commit 29f24ee into main Sep 27, 2026
11 checks passed
@tazarov
tazarov deleted the claude/gallant-allen-72pzmi branch September 27, 2026 17:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[PERF] Add opt-in bounded-memory zstd profile for level-15 encoding

2 participants