You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Every stored field value is persisted twice per lexical segment: once in the docs part and once in the dv (DocValues) part. There is no way to opt a field out. Found while re-triaging the lexical/index backlog; verified on a6ad63c7.
Both sinks are fed from the same analyzed_doc.stored_fields:
DocValues, at add time: laurus/src/lexical/index/inverted/writer.rs:494-498 clones every (field_name, value) into doc_values_writer. The rebuild path repeats it verbatim (writer.rs:1854-1866). DocValuesWriter::add_value (laurus/src/lexical/index/structures/doc_values.rs:83-88) has no gate of any kind.
Stored documents, at flush time: write_stored_documents (writer.rs:1312-1401) serializes the same doc.stored_fields for every buffered document.
Consequences:
Any stored: true field is duplicated on disk regardless of whether it is ever sorted, faceted, or read as a doc value — including DataValue::Bytes and DataValue::Vector payloads.
The schema has no doc_values / fast flag to control it: FieldOption variants carry only indexed / stored / term_vectors (laurus/src/lexical/core/field.rs).
The read side amplifies it further: DocValuesReader::load (doc_values.rs:174-249) materializes every field of the segment into memory on the first doc-value probe, even when the query sorts on a single field.
Note the two encodings are not byte-identical (the docs part is type-tagged, dv is rkyv), so this is duplicated content, not duplicated bytes; the point is the write amplification and the absence of any opt-out.
Proposal
Add a per-field doc-values flag to the schema, defaulting to on so existing indexes keep working, and gate writer.rs:494-498 on it.
Independently of the flag, make DocValuesReader::load lazy per field — the existing layout already length-prefixes each field record (doc_values.rs:126-162) and StorageInput is Read + Seek, so this needs no format change. (This half overlaps the narrowed scope of perf(lexical/index): columnar packed DocValues #547.)
Acceptance criteria
A field can be excluded from DocValues without losing stored-field retrieval
Sorting on one field does not deserialize the other fields' doc values
Segment size measurably drops for a schema with large non-sortable stored fields
Existing indexes (no flag recorded) behave exactly as today
Scope: full implementation, as decided in the investigation comment above. The prior comments' PR-1 content (type-based exclusion, sort-side fallback) is already merged via #1053/#1108. This adds: a latent mixed-segment fallback bug fix (prerequisite for the flag), lazy per-field DocValues loading with bounds checking and deterministic ordering, and the doc_values: bool field option itself with full merge/rebuild/server/binding wiring. DocumentParser's unrelated stored/indexed gap is filed as a separate issue.
Add doc_values: bool to 7 field-option structs (not BytesOption)
Add store_doc_values to InvertedIndexConfig/LexicalIndexConfig
Add field_doc_values/stores_doc_values to InvertedIndexWriterConfig, gate both feed sites
Add SegmentReader::doc_values_field_names() descoped: the merge path's detection only ever needed a per-field has_doc_values(field) check, so the enumeration API ended up with zero real callers and was removed rather than kept dead
New integration tests (flag effect, byte-size comparison)
[~] Field-rebuild tests: covered indirectly through the merge-path tests above (rebuild_field_across_segments shares the same reconstruction code). No Engine::update_field end-to-end test was added: Phase 1's stored-document fallback means a doc_values-only change never changes an observable result (unlike term_vectors/analyzer changes, which do), so there is nothing for such a test to assert beyond what classify_change_table and the merge tests already pin
Phase 5: Schema-change classification
Add classify_doc_values, fold into the 7 lexical arms
Add cases to classify_change_table
Phase 6: Server surface
optional bool doc_values on 7 proto messages
convert/schema.rs both directions
gateway/convert.rs both directions
Server conversion tests
Phase 7: Bindings + CLI
Add doc_values parameter to 7 methods across all 5 bindings
Add a doc_values prompt to the CLI wizard's two option-builder functions
Python/Ruby integration tests
Phase 8: Docs
Update EN + JA docs (schema_and_fields.md, faceting.md, schema_format.md, grpc_api.md, 5x api_reference.md) -- lexical_indexing.md / http_gateway.md / laurus/api_reference.md left unchanged (no term_vectors precedent to mirror there either; reasons noted in the Phase 8 PR comment)
mdbook build clean for both docs trees
Phase 9: Verification, implementation report, PR
cargo test --workspace (full suite) -- 106 binaries all green (also caught and fixed one missed doctest)
cargo check on every binding crate (python/ruby/nodejs/php/mcp + wasm --target wasm32-unknown-unknown --tests)
Direct Python/Ruby scripts confirming the doc_values effect -- covered by Phase 7's integration tests
Manual backward-compat check against an existing sample schema -- loaded resources/schema.toml (no doc_values key) via Python, committed, searched, and round-tripped through to_toml(); an omitted doc_values correctly resolves to true and existing behavior is unchanged
Background
Every stored field value is persisted twice per lexical segment: once in the
docspart and once in thedv(DocValues) part. There is no way to opt a field out. Found while re-triaging the lexical/index backlog; verified ona6ad63c7.Both sinks are fed from the same
analyzed_doc.stored_fields:laurus/src/lexical/index/inverted/writer.rs:494-498clones every(field_name, value)intodoc_values_writer. The rebuild path repeats it verbatim (writer.rs:1854-1866).DocValuesWriter::add_value(laurus/src/lexical/index/structures/doc_values.rs:83-88) has no gate of any kind.write_stored_documents(writer.rs:1312-1401) serializes the samedoc.stored_fieldsfor every buffered document.Consequences:
stored: truefield is duplicated on disk regardless of whether it is ever sorted, faceted, or read as a doc value — includingDataValue::BytesandDataValue::Vectorpayloads.doc_values/fastflag to control it:FieldOptionvariants carry onlyindexed/stored/term_vectors(laurus/src/lexical/core/field.rs).DocValuesReader::load(doc_values.rs:174-249) materializes every field of the segment into memory on the first doc-value probe, even when the query sorts on a single field.Note the two encodings are not byte-identical (the
docspart is type-tagged,dvis rkyv), so this is duplicated content, not duplicated bytes; the point is the write amplification and the absence of any opt-out.Proposal
writer.rs:494-498on it.DocValuesReader::loadlazy per field — the existing layout already length-prefixes each field record (doc_values.rs:126-162) andStorageInputisRead + Seek, so this needs no format change. (This half overlaps the narrowed scope of perf(lexical/index): columnar packed DocValues #547.)Acceptance criteria
Refs #547, #555, #548
Task List
Scope: full implementation, as decided in the investigation comment above. The prior comments' PR-1 content (type-based exclusion, sort-side fallback) is already merged via #1053/#1108. This adds: a latent mixed-segment fallback bug fix (prerequisite for the flag), lazy per-field DocValues loading with bounds checking and deterministic ordering, and the
doc_values: boolfield option itself with full merge/rebuild/server/binding wiring.DocumentParser's unrelatedstored/indexedgap is filed as a separate issue.Phase 1: Mixed-segment fallback bug fix (read-side correctness)
TopFieldCollector::get_field_valueon a DocValues miss (Ok(None)/Err), viadocument_fieldsdocument_fieldsoverride onPerSegmentReaderViewFacetCollector::collect_doc's document fetch lazy so a DV-miss document still gets a fallback contributionhas_doc_values/get_doc_valuedoc comments (semantics + fallback contract)has_doc_values=true,get_doc_value=Ok(None)).dvfile)cargo test -p laurus/cargo fmt/cargo clippycleanPhase 2: Lazy loading + robustness (no behavior change)
alloc_boundsfromvector::indextoutil, generalize error wordingDocValuesWriter.fields->BTreeMap, sortvalues_vecby doc_idDocValuesReaderlazy per-field loading (storage + directory + cache)DocValuesReader::loadget_field; keepfield_names(directory-backed)doc_values_load_once_per_segmentas a stability checkcargo test -p laurus/cargo fmt/cargo clippycleanPhase 3:
doc_valuesflag — core wiringdoc_values: boolto 7 field-option structs (notBytesOption)store_doc_valuestoInvertedIndexConfig/LexicalIndexConfigfield_doc_values/stores_doc_valuestoInvertedIndexWriterConfig, gate both feed sitesAdddescoped: the merge path's detection only ever needed a per-fieldSegmentReader::doc_values_field_names()has_doc_values(field)check, so the enumeration API ended up with zero real callers and was removed rather than kept deadPhase 4: Merge / field-rebuild path
MergeConfig::{field_doc_values, default_doc_values}, schema-first resolutionperform_merge/reconstruct_segment*/rebuild_field_across_segmentsInvertedIndex::writer()/rebuild_field/merge_segment_setrebuild_field_across_segmentsshares the same reconstruction code). NoEngine::update_fieldend-to-end test was added: Phase 1's stored-document fallback means adoc_values-only change never changes an observable result (unliketerm_vectors/analyzer changes, which do), so there is nothing for such a test to assert beyond whatclassify_change_tableand the merge tests already pinPhase 5: Schema-change classification
classify_doc_values, fold into the 7 lexical armsclassify_change_tablePhase 6: Server surface
optional bool doc_valueson 7 proto messagesconvert/schema.rsboth directionsgateway/convert.rsboth directionsPhase 7: Bindings + CLI
doc_valuesparameter to 7 methods across all 5 bindingsPhase 8: Docs
mdbook buildclean for both docs treesPhase 9: Verification, implementation report, PR
cargo test --workspace(full suite) -- 106 binaries all green (also caught and fixed one missed doctest)cargo fmt --all -- --check/cargo clippy --workspace --all-targets --features embeddings-all -- -D warnings-- both cleancargo checkon every binding crate (python/ruby/nodejs/php/mcp + wasm --target wasm32-unknown-unknown --tests)resources/schema.toml(no doc_values key) via Python, committed, searched, and round-tripped throughto_toml(); an omitted doc_values correctly resolves to true and existing behavior is unchangedDocumentParser'sstored/indexedgap -- fix(lexical): DocumentParser ignores indexed/stored field options #1114Closes #1047,Refs #547 #555 #548) -- perf(lexical): add a doc_values opt-out for stored-field DocValues duplication #1115