Repository navigation
Better Vector Search for Long Documents: Chunking Inside Manticore Search #4939
githubmanticore
announced in
Blog
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Originally posted by Dmitrii Kuzmenkov on https://mnt.cr/go/Ku6gwD on September 15, 2026
Better Vector Search for Long Documents: Chunking Inside Manticore Search
An embedding model reads only the first few hundred tokens of a document and silently drops the rest. Manticore Search now splits long documents for you at INSERT time: add chunk_strategy to the vector column and pick one of five strategies. No ingest pipeline, no splitter library. On our own manual, recall@5 for deep content went from 55% to 83%.
Say you are building search over your team's internal documentation — guides, runbooks, postmortems. You have a table with auto embeddings: you insert text, Manticore runs the model and fills the vector column for you. (If that is new to you, start with vector search in Manticore.) You load a 4,000-word document. The insert succeeds. The search works. Everything looks fine.
Except the model you picked has a 512-token input window, and that document is about 5,000 tokens long. The model read the first 380 words and threw away the other 3,600. Nothing in the document past that point can ever be retrieved, and nothing anywhere told you. The embedding may not represent the document as a whole either.
Until now, you would usually split the document into several pieces yourself, create embeddings for each one, and then work out how to combine the results if you wanted document search rather than chunk search. Manticore now handles this in the table definition: add
chunk_strategyto the vector column inCREATE TABLE, and Manticore splits each document into chunks, embeds every chunk, and searches all of them:That is the whole feature. No ingest pipeline, no splitter library, no second table for chunks, no
GROUP BYto fold chunk hits back into documents.TL;DR
truncate(the old default),mean,fixed,recursive,sentence. Set withchunk_strategyon a model-backed vector column.truncateandmeanproduce one vector per document and work on afloat_vectorcolumn.fixed,recursiveandsentenceproduce many, so they need afloat_vector_arraycolumn.knn_dist()reporting the distance to its closest chunk.kcounts documents, not chunks.max_tokens(chunk size),overlap_tokens(shared tokens between neighbors),max_chunks(ceiling per document).The problem, shown with a small example
Suppose you have four documents:
You can create the table and add the documents using the commands below.
So, what we have is: one table, three vector columns with the same source text — one column per strategy. A single
INSERTfills all three, so the comparison conditions are identical:Insert the four documents
Now ask a question whose answer lives in the runbook's last section, once per strategy:
truncate(default)meansentence, 128 tokens, 32 overlapWith
truncate, the document that actually answers the question loses to a decoy that merely looks like it is about certificates. The runbook's single vector was built from its opening pages on backup schedules and restore drills, because that is all the model was allowed to read.With
sentencechunking, the runbook is stored as nine vectors instead of one:One of those nine is the certificate-rotation paragraph. It matches the query almost exactly, so the document wins by a wide margin: 0.254 against 0.700.
More about the chunking strategies
truncatefloat_vectormeanfloat_vectorfixedfloat_vector_arraymax_tokenstokens.recursivefloat_vector_arraymax_tokens.sentencefloat_vector_arraymax_tokens.The important distinction is not how the text is cut. It is what a match means.
With one vector per document, search asks: "is this document, as a whole, similar to the query?" A single relevant paragraph is diluted by everything around it, and a document that covers five topics ends up not really matching any of them.
With one vector per chunk, search asks: "does this document contain something similar?" Each chunk competes on its own merits, and Manticore returns the document once, scored by its best chunk.
truncate— keep it when your documents are shortThis is what you already have;
chunk_strategy='truncate'is the default and you never have to write it. It's the right choice — and the fastest, and the smallest — whenever your text genuinely fits the model's window: product titles, short descriptions, tags, chat messages, log lines, search queries, commit subjects.How much fits? More than most people assume, and less than they hope.
all-MiniLM-L6-v2takes 512 tokens, roughly 380 English words.text-embedding-3-smalltakes 8,192. If your 95th-percentile document is comfortably under the limit, stop reading and keeptruncate.When it hurts: anything long-form. Documentation pages, knowledge-base articles, contracts, transcripts, email threads, wiki pages, README files, incident postmortems.
mean— one vector, but the whole documentManticore splits the document, embeds every chunk, and averages the chunk vectors into a single normalized vector. Storage and search cost are identical to
truncate— one vector per document, one HNSW node — but nothing is thrown away.Use it when:
float_vectorand you can't change the type (for example you're adding the column to an existing table withALTER, which multi-vector strategies don't support).Do not use it when a document covers several unrelated topics. Averaging a legal contract's indemnity clause with its payment terms produces a vector that sits between them and is close to neither. In our benchmark below,
meanrecovered about a third of the gap that chunking closes — a real improvement, and clearly not the same thing.fixed— predictable, cheapest to reason aboutCut every
max_tokenstokens, no matter what the text is doing at that point. Chunk count is a straight function of document length, so index size is predictable before you load anything.Use it when the text has no reliable structure to exploit: OCR output, scraped HTML that lost its paragraphs, machine transcripts without punctuation, log dumps, minified content. Also a fine default when you simply want the cheapest thing that stops truncation.
The cost: a boundary can land mid-sentence, and a chunk that begins in the middle of a thought embeds badly. That is exactly what
overlap_tokensis for — see below.recursive— the best general default for proseSame token budget as
fixed, but each cut is pulled back to the nearest natural boundary: a blank line first, then a line break, then a sentence end, then a space. A chunk stops where the text stops, not where the counter runs out. The boundary is never dragged back past the midpoint of the chunk, so you don't get a stream of tiny fragments.If you have used LangChain's
RecursiveCharacterTextSplitter, this is the same idea, except it runs inside the database on the model's real tokens instead of characters, and there is nothing to install.Use it for: Markdown and HTML documentation, wiki pages, knowledge bases, blog posts, README files, structured reports — anything written by a human in paragraphs. This scored highest on deep content in our benchmark.
sentence— when a chunk must be a complete thoughtDetects sentence boundaries with Unicode UAX #29, then greedily packs whole sentences until the token budget is reached. A chunk never starts or ends mid-sentence. A single sentence longer than the budget is split by the token window, as a last resort.
Use it for: support tickets and email threads, chat and meeting transcripts, legal and policy text, news, customer reviews, medical and scientific abstracts — anything where a fragment of a sentence changes or destroys the meaning. It is also the strategy to pick when chunks will be fed to an LLM afterwards, because a chunk that ends mid-clause reads badly in a prompt.
sentenceis a little more conservative thanrecursive: it produced fewer, cleaner chunks in our tests and scored about the same on recall@5.The three knobs
max_tokensis capped at what the model can actually accept — ask for 4,096 on a 512-token model and you still get 512, not an error. Smaller chunks mean sharper matches and more vectors; larger chunks mean more context per vector and fewer of them. For English prose, 128–512 covers almost every use case; we used 256 throughout the benchmark.overlap_tokensrepeats the tail of each chunk at the head of the next, so a sentence that straddles a boundary still appears intact somewhere. 10–20% ofmax_tokensis the usual setting. Manticore guarantees forward progress:fixedandrecursivecap the overlap at half the chunk size, andsentencere-seeds the next chunk with at mostoverlap_tokensworth of trailing whole sentences while always advancing by at least one sentence. It requires an explicit non-zeromax_tokens— overlap against "whatever the model's limit happens to be" isn't a meaningful setting, so Manticore rejects it.max_chunkslimits the impact of unusually large documents. Without it, a 400-page PDF pasted into one row becomes thousands of HNSW nodes. With it, Manticore merges the overflow into the last kept chunk, then truncates it to the model's window when embedding:Use it as a guard rail against outliers, not as a way to save memory across the board.
What search looks like
Nothing about your query changes from before. There is no chunk table, no nested field, no join, no
GROUP BY. Here is the complete example:max_tokens='32'is small on purpose here, so that these short notes actually split and you can see the multi-vector behaviour on a toy dataset.LENGTH()on the vector column shows how each document was divided:Six vectors, three rows back. Search follows the rules described in Multiple vectors per document:
knn_dist()is the distance to its closest chunk.kcounts documents, not vectors.knn(chunks, 3, ...)means three documents.The same query over HTTP:
Everything else on the KNN page keeps working as before: filtering, prefilter and postfilter strategies, quantization, early termination, and rescoring.
Does it actually help? Numbers on our own manual
We tested the feature on the Manticore English manual — 189 pages and about 298,000 words, ranging from a two-paragraph note to a 39,000-word changelog.
The query set is generated mechanically, not hand-picked. For every page we took its section headings, kept only headings that are unique across the whole manual, and split them in two:
truncatecan partly see. The gap below is therefore an understatement, not an exaggeration.truncatecan already see.A query is a hit if KNN returns the page the heading came from, within the top k. Model:
Xenova/all-MiniLM-L6-v2(384 dims, 512-token window) running on Manticore's ONNX backend. Hardware: 32 threads.max_tokens='256',overlap_tokens='32'for the multi-vector strategies. Quality numbers are deterministic for a given index; the timings are a single run per strategy on an otherwise idle box.Deep content — what chunking is for
truncatemeanfixedrecursivesentenceChunking turns a coin flip into a working search. recall@5 goes from 55.1% to 83.3%, and the rank of the right answer improves just as much — MRR 0.44 → 0.70. Of the queries
truncatecould not answer in the top 5 at all,recursiverecovers roughly two thirds.meanlands where you would expect: it recovers about a third of the gap for free, because it costs exactly nothing extra to store or search.Head content — the control group
truncatemeanfixedrecursivesentenceFor completeness, the control group is worth reading carefully. For content that the model could already see,
truncateis still the most precise at rank 1 — 65.9% against 58.0% forrecursive. A whole-document vector carries the page's overall topic, and when the query is about the page's opening subject, that context helps.By rank 5 the difference is gone:
recursivematchestruncateexactly at 86.4%. So the trade is a few points of top-1 precision on content near the beginning, in exchange for +28 points of recall on everything else. For a documentation search, a help center, or any RAG retriever that feeds 5–10 passages to an LLM, that is not a close call.Cost
OPTIMIZE, though Manticore builds it across all your cores.If you use a paid embedding API, read that ingest number as a bill: chunking sends your whole corpus to the model instead of the head of each document, and you pay for every token of it. Local ONNX models have no per-token cost, which is a large part of why we made them fast.
Recommendations for choosing a chunking strategy
truncatefloat_vectormeanrecursive,max_tokens128–256sentence,max_tokens128–256fixed,max_tokens256, plus overlapsentence,max_tokens384–512Overlap is deliberately absent from most of those: our sweep below could not measure a benefit from it on structured prose, and it costs vectors. Add it when a thought routinely straddles a boundary — unstructured transcripts, OCR, long narrative without paragraph breaks.
How big should a chunk be?
Chunk size is the setting that actually affects your results. The trade is direct: a smaller chunk is a sharper match on one idea, a larger chunk carries more context but dilutes each idea inside it. A paragraph buried in a long document only becomes findable once the chunk size is small enough to give it a vector of its own.
We ran another test:
recursiveon the same 189-page manual, three chunk sizes × three overlap settings, and the same 419 deep queries. Two trends stand out: quality rises as chunks get smaller, while overlap adds cost without improving quality much.Note that the Y axis starts at 78%, not zero — the whole spread is about six points, so a zero-based axis would flatten it into a straight line. The numbers behind the chart:
max_tokensoverlap_tokensWhat we see:
Smaller chunks win, consistently. Going from 512 to 128 tokens buys about five points of recall@5 (80.2% → 85.2%) and a large jump in ranking quality (MRR 0.655 → 0.718). It costs 4× the vectors and 2.5× the index RAM. Below 128 the chunks stop containing a whole thought, so this is not a slope you ride forever — but on long technical prose, 128–256 beat 512 every time.
Overlap did essentially nothing for quality, and was not free. At 128 tokens, going from no overlap to 25% overlap moved recall@5 from 85.2% to 85.2% while adding 41% more vectors and 5 MB of RAM. The pattern holds at every size: the spread across overlap settings (±1.5 points) is within the noise of a 419-query set, while the cost is not. This lines up with Chroma's chunking evaluation, where plain recursive splitting at 200 tokens with no overlap scored 88.1% recall — within a few points of an LLM-driven splitter at 91.9% — and it is the opposite of the "always use 10–20% overlap" advice you will read in most RAG guides.
The honest caveat: this is one corpus, one model, and queries that look like section headings. Overlap earns its keep when a single fact routinely straddles a boundary — long unbroken narrative, transcripts without structure — and
recursivealready snaps cuts to paragraph and sentence boundaries, which does much of the same work. So treat "start at 128–256 with no overlap, add overlap only if you can measure it helping" as the default, and check it on your own data with the recipe below.Comparing settings
You do not have to guess, and you do not need two tables. A table can carry several model-backed vector columns, each with its own strategy, all filled from the same fields on the same
INSERT:Load your corpus once, then run the same query against each column and compare. For example:
On a short runbook whose certificate section sits at the end,
sentence/256 fits the whole document in a single chunk and answers at distance 0.515;recursive/128 splits it in two, isolates the certificate paragraph, and answers at 0.310. Same row, same model, same query — only the chunk size differs.Build a dataset of real queries with answers you trust — even 50 is enough — and compare recall@5 across two or three columns, exactly as we did on the manual above. Then drop the losing column with
ALTER TABLE ... DROP COLUMNand keep the winner.Recipes
Documentation and help center search. Long Markdown pages, users asking questions in their own words. Chunk on structure and search across it:
Note
from='title,body': the fields are joined before chunking, so the page title lands in the first chunk and gives it context. For a worked end-to-end example of this shape, see Vector search on GitHub.Support tickets and email threads. A thread is a sequence of complete messages; cutting one mid-sentence loses the fact you need. Keep the chunk count bounded, because threads have no natural length limit:
Filtering works exactly as it does for a single-vector column.
Contracts and policy documents. Clause-level retrieval is the entire point — nobody wants "the contract" back, they want the indemnity clause. Smaller chunks, generous overlap:
Product catalog with long descriptions. One product is one topic, and catalogs are large, so pay nothing extra:
RAG: retrieval for an LLM. Whatever you retrieve gets pasted into a prompt, so chunks should read as prose — this is the retrieval half of conversational search. Larger chunks, sentence boundaries, and ask for more of them:
Adding chunking to a table you already have. Multi-vector columns can't be added by
ALTER— existing rows have no vectors and there's no way to backfill them yet. A single-vector strategy can:For a multi-vector column, create the new table with the column in place and reindex into it.
How other engines handle this
Every vector engine now generates embeddings for you. Far fewer will split your document before doing it — and of those, most make you assemble it out of pipeline stages.
chunk_strategyon the vector columnsemantic_texthides the chunksEvery cell in the chunking column links to a source. A “yes” links to the feature's own documentation. An “app-side” links to that vendor's own guidance on chunking in your application — which is what they publish instead of an in-engine option. If we missed a feature, or one has shipped since, tell us and we will fix it.
Versions checked - 4 September 2026
The latest stable release of each product available that day:
Two things stand out.
Chunking in-engine is still rare. Milvus, Qdrant, Weaviate, Pinecone, MongoDB Atlas, Typesense, Meilisearch and — since the 9.8 LLM module — Apache Solr will all run the embedding model for you, and every one of them will happily truncate your 4,000-word document without saying so. The splitting is your problem, in your application, in a language and a library that has no idea what tokenizer the model uses.
Where chunking exists, the plumbing usually leaks. OpenSearch gets you there with a
text_chunkingprocessor feeding atext_embeddingprocessor writing into a nested field, queried with a nested query and a score mode. Azure AI Search wants a skillset with a Split skill, an embedding skill and index projections — and returns one result row per chunk, so grouping back to documents is on you. pgai Vectorizer writes chunks to a second table, so every query is a join plus aDISTINCT ON. Elasticsearch'ssemantic_textis genuinely close to Manticore's model: chunking settings on the inference endpoint, chunks hidden inside the field, one hit per document.Manticore does the same thing with less surface area: the strategy is an option on the column, the chunks are the column's value, and search returns documents. If you are weighing the whole stack rather than this one feature, we have written up the comparison with Elasticsearch and with Turbopuffer too.
What chunking does not fix
Chunking solves one problem well — a document longer than the model's window is no longer half-invisible. It does not make retrieval perfect, and two known gaps are worth naming.
A chunk does not know where it came from. Split a document and you get a paragraph that says "do one node at a time and confirm every peer reports as synced" with no indication of what is being rotated, or which product it belongs to. Anthropic's contextual retrieval work put numbers on this: prepending a short, chunk-specific description of the surrounding document before embedding cut top-20 retrieval failures by 35%, and by 49% combined with a contextual BM25 index.
Manticore does not do this for you.
FROMjoins its fields with a space before chunking, so listingtitlefirst puts the title at the head of the text that gets split — which means it lands in the first chunk and only that one. Every chunk after it is on its own:If you need every chunk to carry context, you have to build it into the stored text yourself before inserting — for example by repeating a short heading at the start of each section of
body. There is no per-chunk prefix option today.Chunk boundaries are decided before the model sees the text. Manticore splits, then embeds each piece independently — the standard approach, and what every engine with in-engine chunking in the table above does. An alternative called late chunking inverts it: run a long-context model over the whole document first, then pool the token embeddings into chunks, so each chunk vector carries context from the rest of the document. It needs a long-context model and more compute per document, and Manticore does not do it today. If your documents depend heavily on cross-paragraph context, it is worth knowing the option exists.
Neither gap changes the basic result: for long documents, chunked retrieval beats truncated retrieval by a wide margin, and reaching it means adding
chunk_strategyto one column.Limits and gotchas
max_chunksdiscards text. Manticore merges overflow into the last kept chunk, then truncates it to the model's window. Nothing warns you. It's a guard rail for outliers, not a way to save memory across the board.Remote models chunk by bytes, not tokens. OpenAI, Voyage and Jina have no local tokenizer, so Manticore falls back to a deliberately conservative estimate of 3 bytes per token — a chunk lands under the provider's cap rather than over it. In practice
max_tokens='N'becomes anN × 3-byte window. We measured it against a stub endpoint with a 3,599-byte document and thefixedstrategy:max_tokensEnglish prose runs closer to 4 bytes per token, so on a remote model you get chunks roughly a quarter smaller than the number you asked for — set
max_tokensabout 30% higher than you would for a local model to land in the same place. If exact boundaries matter, use a local model, where splitting is done on the model's real tokens.Multi-vector columns can't be added with
ALTER.ALTER TABLE ... ADD COLUMNandALTER TABLE ... REBUILD EMBEDDINGSon a model-backedfloat_vector_arrayare rejected. Recreate the table instead. Both work normally on afloat_vector, including withmean.Chunking applies to auto embeddings only. Vectors you insert yourself are stored exactly as given — Manticore never re-cuts data you supplied.
chunk_strategywithoutmodel_nameis a DDL error, on purpose.embeddingsis a reserved word.EMBEDDINGSis a DDL keyword (ALTER TABLE ... REBUILD EMBEDDINGS), so a column literally namedembeddingsis a syntax error unless escaped. Use escaping if you need that name.Queries are not chunked. A query is embedded whole, as a single vector. That is what you want: chunking exists to make a long document findable, not to split a fifteen-word question.
The DDL tells you when a combination is wrong, at
CREATE TABLEtime rather than at the first insert:Try it
The shortest path to a working chunked semantic search:
No model to download by hand, no splitter to pick, no pipeline to maintain. One column option, and the parts of your documents that used to be invisible start showing up in results.
Full reference: Chunking strategies and Multiple vectors per document in the manual. Questions and bug reports on GitHub.
All reactions