Skip to content

perf(lexical/index): stored field values are written twice per segment (docs + dv) with no opt-out #1047

Description

@mosuka

Background

Every stored field value is persisted twice per lexical segment: once in the docs part and once in the dv (DocValues) part. There is no way to opt a field out. Found while re-triaging the lexical/index backlog; verified on a6ad63c7.

Both sinks are fed from the same analyzed_doc.stored_fields:

  • DocValues, at add time: laurus/src/lexical/index/inverted/writer.rs:494-498 clones every (field_name, value) into doc_values_writer. The rebuild path repeats it verbatim (writer.rs:1854-1866). DocValuesWriter::add_value (laurus/src/lexical/index/structures/doc_values.rs:83-88) has no gate of any kind.
  • Stored documents, at flush time: write_stored_documents (writer.rs:1312-1401) serializes the same doc.stored_fields for every buffered document.

Consequences:

  • Any stored: true field is duplicated on disk regardless of whether it is ever sorted, faceted, or read as a doc value — including DataValue::Bytes and DataValue::Vector payloads.
  • The schema has no doc_values / fast flag to control it: FieldOption variants carry only indexed / stored / term_vectors (laurus/src/lexical/core/field.rs).
  • The read side amplifies it further: DocValuesReader::load (doc_values.rs:174-249) materializes every field of the segment into memory on the first doc-value probe, even when the query sorts on a single field.

Note the two encodings are not byte-identical (the docs part is type-tagged, dv is rkyv), so this is duplicated content, not duplicated bytes; the point is the write amplification and the absence of any opt-out.

Proposal

  1. Add a per-field doc-values flag to the schema, defaulting to on so existing indexes keep working, and gate writer.rs:494-498 on it.
  2. Independently of the flag, make DocValuesReader::load lazy per field — the existing layout already length-prefixes each field record (doc_values.rs:126-162) and StorageInput is Read + Seek, so this needs no format change. (This half overlaps the narrowed scope of perf(lexical/index): columnar packed DocValues #547.)

Acceptance criteria

  • A field can be excluded from DocValues without losing stored-field retrieval
  • Sorting on one field does not deserialize the other fields' doc values
  • Segment size measurably drops for a schema with large non-sortable stored fields
  • Existing indexes (no flag recorded) behave exactly as today

Refs #547, #555, #548

Task List

Scope: full implementation, as decided in the investigation comment above. The prior comments' PR-1 content (type-based exclusion, sort-side fallback) is already merged via #1053/#1108. This adds: a latent mixed-segment fallback bug fix (prerequisite for the flag), lazy per-field DocValues loading with bounds checking and deterministic ordering, and the doc_values: bool field option itself with full merge/rebuild/server/binding wiring. DocumentParser's unrelated stored/indexed gap is filed as a separate issue.

Phase 1: Mixed-segment fallback bug fix (read-side correctness)

  • Add stored-document fallback to TopFieldCollector::get_field_value on a DocValues miss (Ok(None)/Err), via document_fields
  • Add the missing document_fields override on PerSegmentReaderView
  • Make FacetCollector::collect_doc's document fetch lazy so a DV-miss document still gets a fallback contribution
  • Update has_doc_values/get_doc_value doc comments (semantics + fallback contract)
  • Unit tests (mock reader: has_doc_values=true, get_doc_value=Ok(None))
  • Integration test with an actually column-absent segment (deleted .dv file)
  • cargo test -p laurus / cargo fmt / cargo clippy clean

Phase 2: Lazy loading + robustness (no behavior change)

  • Move alloc_bounds from vector::index to util, generalize error wording
  • DocValuesWriter.fields -> BTreeMap, sort values_vec by doc_id
  • DocValuesReader lazy per-field loading (storage + directory + cache)
  • Bounds-check header-declared sizes in DocValuesReader::load
  • Remove unused get_field; keep field_names (directory-backed)
  • Rewrite doc_values_load_once_per_segment as a stability check
  • Lazy-load tests (byte-count based) + bounds-check tests + determinism test
  • cargo test -p laurus / cargo fmt / cargo clippy clean

Phase 3: doc_values flag — core wiring

  • Add doc_values: bool to 7 field-option structs (not BytesOption)
  • Add store_doc_values to InvertedIndexConfig/LexicalIndexConfig
  • Add field_doc_values/stores_doc_values to InvertedIndexWriterConfig, gate both feed sites
  • Add SegmentReader::doc_values_field_names() descoped: the merge path's detection only ever needed a per-field has_doc_values(field) check, so the enumeration API ended up with zero real callers and was removed rather than kept dead
  • New integration tests (flag effect, byte-size comparison)

Phase 4: Merge / field-rebuild path

  • MergeConfig::{field_doc_values, default_doc_values}, schema-first resolution
  • Wire perform_merge/reconstruct_segment*/rebuild_field_across_segments
  • Wire InvertedIndex::writer()/rebuild_field/merge_segment_set
  • Merge tests (both field-name orderings, schema-priority pin, no-values-present pin)
  • [~] Field-rebuild tests: covered indirectly through the merge-path tests above (rebuild_field_across_segments shares the same reconstruction code). No Engine::update_field end-to-end test was added: Phase 1's stored-document fallback means a doc_values-only change never changes an observable result (unlike term_vectors/analyzer changes, which do), so there is nothing for such a test to assert beyond what classify_change_table and the merge tests already pin

Phase 5: Schema-change classification

  • Add classify_doc_values, fold into the 7 lexical arms
  • Add cases to classify_change_table

Phase 6: Server surface

  • optional bool doc_values on 7 proto messages
  • convert/schema.rs both directions
  • gateway/convert.rs both directions
  • Server conversion tests

Phase 7: Bindings + CLI

  • Add doc_values parameter to 7 methods across all 5 bindings
  • Add a doc_values prompt to the CLI wizard's two option-builder functions
  • Python/Ruby integration tests

Phase 8: Docs

  • Update EN + JA docs (schema_and_fields.md, faceting.md, schema_format.md, grpc_api.md, 5x api_reference.md) -- lexical_indexing.md / http_gateway.md / laurus/api_reference.md left unchanged (no term_vectors precedent to mirror there either; reasons noted in the Phase 8 PR comment)
  • mdbook build clean for both docs trees

Phase 9: Verification, implementation report, PR

  • cargo test --workspace (full suite) -- 106 binaries all green (also caught and fixed one missed doctest)
  • cargo fmt --all -- --check / cargo clippy --workspace --all-targets --features embeddings-all -- -D warnings -- both clean
  • cargo check on every binding crate (python/ruby/nodejs/php/mcp + wasm --target wasm32-unknown-unknown --tests)
  • Direct Python/Ruby scripts confirming the doc_values effect -- covered by Phase 7's integration tests
  • Manual backward-compat check against an existing sample schema -- loaded resources/schema.toml (no doc_values key) via Python, committed, searched, and round-tripped through to_toml(); an omitted doc_values correctly resolves to true and existing behavior is unchanged
  • File a separate issue for DocumentParser's stored/indexed gap -- fix(lexical): DocumentParser ignores indexed/stored field options #1114
  • PR (Closes #1047, Refs #547 #555 #548) -- perf(lexical): add a doc_values opt-out for stored-field DocValues duplication #1115

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions