Manticore Search has introduced a new feature that handles document chunking directly within the database engine. By adding a chunk_strategy to a vector column in a CREATE TABLE statement, users can automatically split long documents into smaller pieces for embedding and search.
Product LaunchesManticore Search
Manticore Search Adds In-Engine Chunking for Improved Vector Search on Long Documents
This feature addresses the limitation of model input windows. For example, when using a model with a 512-token limit on a 5,000-token document, standard embedding processes would only capture the beginning of the text. Manticore's approach allows the engine to split the document, embed every chunk, and search them all, ensuring that relevant information hidden deep within a document is not lost.
Manticore supports several chunking strategies:
- Truncate: The default method, which uses the model's own token limit.
- Mean: Splits the document and averages the chunk vectors into a single vector, which prevents data loss but may dilute specific topics.
- Fixed: Cuts text at specific token intervals.
- Recursive: Splits text at natural boundaries like blank lines, line breaks, or sentence ends, which scored highest in the company's benchmarks.
- Sentence: Detects sentence boundaries using Unicode UAX #29 to ensure chunks do not break mid-sentence.
The implementation requires no additional ingestion pipelines, external splitter libraries, or secondary tables for chunks. Users can manage different strategies for different columns within the same table by using multiple vector columns. This functionality aims to make semantic search on long-form content, such as documentation or legal contracts, more precise and easier to manage.
Sources
- Better Vector Search for Long Documents: Chunking Inside Manticore Search (Hacker News Frontpage, 2026-09-17)