Repository navigation
How to search for long hashes and IDs with dict='keywords_32k' #4960
githubmanticore
announced in
Blog
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Originally published on the Manticore Search website on August 30, 2026
How to search for long hashes and IDs with dict='keywords_32k'
A practical guide to searching long hashes, event IDs, message IDs, and email addresses in Manticore Search: limits, exact matching, wildcard search, tokenization, migration, and limitations.
Full-text search usually works with ordinary words: product names, titles, comments, and descriptions. Such tokens are rarely longer than a few dozen characters.
Logs and technical data are different. A SHA-256 hash is 64 characters long, while message IDs, correlation keys, event identifiers, and some email addresses can be even longer. The value is often meaningful only as a whole: if its tail is lost, one ID can easily be confused with another.
Manticore Search provides
dict='keywords_32k'for these cases.The problem with the regular dictionary
By default, Manticore uses
dict='keywords'. Its maximum token length is 42 bytes after normalization.The limit is measured in bytes, not characters. One ASCII character takes one byte, but a UTF-8 character can take several bytes.
If a token exceeds 42 bytes, Manticore truncates it:
As a result, querying the complete long value does not necessarily return zero results. Because the query is truncated too, the document may be found—but only by the first 42 bytes.
This creates a more serious problem: two different IDs with the same first 42 bytes become indistinguishable to full-text search. It is also impossible to find a value by a fragment located after the 42nd byte.
What
keywords_32kchangesdict='keywords_32k'increases the maximum normalized token length to 32768 bytes, or 32 KB.dict='keywords'dict='keywords_32k'The setting applies to the entire table:
Regular short words in the same table continue to use the configured morphology. Tokens longer than 42 bytes are stored in their original normalized form, without stemming or lemmatization.
For machine identifiers, this is usually exactly what you need: a hash or event ID has no useful word stem.
When to use
keywords_32kUse it when both conditions are true:
MATCH(), by prefix, or by substring.Typical examples include:
However,
keywords_32kis not required for every ID lookup.If you only need exact equality
If the application always receives the complete ID and only needs to check exact equality, a string attribute is sufficient:
Then use a regular filter:
If you need both exact matching and full-text search
Use
string attribute indexed:In this case, Manticore:
WHERE;MATCH()and wildcard search.This is usually the most convenient schema for technical identifiers.
Searching by the complete token
Create a table and add a 64-character SHA-256 hash:
Use the complete normalized token in a full-text search:
The
@event_idoperator restricts the search to the required field. Without it, Manticore searches for the value in all full-text fields of the table.Quotes are not needed around a single simple alphanumeric token here. Quotes denote a phrase search; they do not turn
MATCH()into a byte-for-byte comparison of the original string.A full-text match is not the same as exact equality
MATCH()operates on the result of tokenization and normalization. It can be affected by:charset_table;blend_chars;ignore_chars;For a strict comparison of the stored value, use a string attribute:
binaryenables byte-for-byte string comparisons in the current SQL session. This setting does not affect full-text search behavior.The practical rule is simple:
WHERE event_id = ...— a strict comparison of the stored string;MATCH('@event_id ...')— a normalized-token search;MATCH('@event_id prefix*')— a prefix search;MATCH('@event_id *fragment*')— a substring search.How to inspect tokenization
Before loading a large volume of data, check that Manticore actually sees the value as a single token:
The
normalizedcolumn should contain the complete 64-character hash.CALL KEYWORDSis especially useful for values containing:@symbol;This lets you inspect the actual token boundaries before indexing the data.
Prefix and substring search
To search inside a token, enable
min_infix_len:Prefix search:
Substring search:
A positive
min_infix_lenalso enables prefix search. This example uses4to prevent excessively short and broad patterns.Why short patterns are dangerous
With
dict='keywords_32k', as with regulardict='keywords', Manticore does not precompute every possible substring. Instead, at query time it expands a wildcard pattern into matching dictionary terms.For example, the pattern
*ab*may match a huge number of values. The more expansions it produces, and the more documents that contain each matching term, the more expensive the query becomes.For production systems:
min_infix_lenbased on real data;expansion_limitto limit the number of expansions;@field.index_exact_words='1'is not required for wildcard search itself. You need it if you want to distinguish exact matches from wildcard matches when ranking, usually together withexpand_keywords.Email addresses and other values with separators
keywords_32kchanges only the maximum token length. It does not determine where a token begins and ends.By default, a period,
@, hyphen, and other characters may split a value into multiple parts. If an email address or message ID must also be indexed as a whole, you can useblend_chars:Blended characters are indexed in two ways:
This makes it possible to search both the complete email address and individual words within it.
Searching for the complete value:
The quotes matter here because
@is also used in full-text query syntax. Inside a phrase, the parser can process it as a blended character.Searching by a fragment:
Inspect the tokenization result:
In a real application, values passed to
MATCH()must be escaped correctly. Simply appending user input to a query string can change the query's meaning because of@,-,|,!,",*, and other operators.How to convert an existing table
For an RT table, you can change the setting with
ALTER TABLE:However, this affects only documents added or replaced after the setting is changed.
Existing documents are not automatically tokenized again. Their long tokens remain in the old truncated form until the documents are reindexed.
Follow these steps:
dict.SHOW CREATE TABLE.CALL KEYWORDSandMATCH().For a plain table:
dict = keywords_32k.ALTER TABLE ... RECONFIGUREif this fits your update workflow.Until the data is reindexed, the same table may contain both:
This can produce different results for documents that appear identical.
Why
dict='crc'does not solve this problemdict='crc'stores keyword checksums instead of their original text. However, it does not increase the allowed token length.The exception to the regular 42-byte limit is implemented specifically by
dict='keywords_32k'.The
keywordsandkeywords_32kdictionaries also store term text, which allows Manticore to expand prefix and infix wildcard queries against the dictionary.If you need to search long machine identifiers,
crcis not a replacement forkeywords_32k.Current limitations
At the time of publication,
dict='keywords_32k'has several limitations:CALL SUGGESTandCALL QSUGGESTare not supported;indextool --dumpdictcannot dump this type of dictionary;REGEXoperator works withdict='keywords', but not withkeywords_32k.The last point should not be confused with the
REGEX()function for filtering string attributes. If the field is declared asstring attribute indexed, attribute filtering and full-text search remain two separate mechanisms.Do not index secrets
Being able to search a long value does not mean that you should store it in a search index.
Do not index the following unless absolutely necessary:
keywords_32ksolves the search problem, but does not protect the value from being read by users, backups, query logs, or system administrators.If a secret must be matched by its exact value, it is safer to calculate a suitable hash in advance and store only the hash.
Quick checklist
Before enabling
keywords_32k, check:WHERE value = ...?CALL KEYWORDSsee the entire value as one token?blend_charsfor periods, hyphens,@, and other separators?min_infix_lenlarge enough?SUGGEST, percolate, or full-textREGEX?Summary
dict='keywords_32k'solves a specific problem: it allows a full-text index to store normalized tokens up to 32768 bytes long instead of the regular 42 bytes.It is well suited to long hashes, event IDs, message IDs, email addresses, and other machine identifiers. Keep three points in mind:
keywords_32kincreases the token length but does not change tokenization rules;MATCH()on a complete token is not the same as a strict comparison of the original string;If you only need exact equality, use a string attribute. If you need both exact matching and partial-value search, use
string attribute indexedtogether withdict='keywords_32k'.Documentation
dictandkeywords_32klimitationsblend_charsCALL KEYWORDSAll reactions