Better Vector Search for Long Documents: Chunking Inside Manticore Searchnew
Manticore Search added automatic document chunking for vector columns, lifting long-document recall@5 from 55.1% to 83.3% in its benchmarks.
Manticore Search introduced a chunk_strategy option for model-backed vector columns in CREATE TABLE, offering five strategies (truncate, mean, fixed, recursive, sentence) with tunable max_tokens, overlap_tokens, and max_chunks, eliminating external splitters and separate chunk tables. On its 189-page, ~298k-word manual, sentence chunking improved recall@5 from 55.1% to 83.3% and MRR from 0.44 to 0.70, at roughly 2.5x RAM and 4x ingest time. Documents still return as single results; queries are never chunked.
20