Today we're introducing topk-embed-v1, a family of multi-vector embedding models for text and visual documents. Two smaller models, xsmall and small, are open-sourced, while the production-size model is offered inside TopK for $0.05/1M tokens.
We co-designed the models and storage to enable efficient multi-vector retrieval at scale. Our production model scores 65.16 nDCG@10 at 11.27 KiB per page on native ViDoRe v3 image queries, delivering strong retrieval quality at storage size comparable to dense embeddings.
Highest quality for $0.05/1M tokens
ViDoRe v3 tests retrieval over documents such as financial reports, scientific papers and industrial manuals. We evaluate page images, which preserve layout, tables and figures, and Markdown, which tests retrieval from the text. We measure quality with nDCG@10: how well the ten highest-ranked results match the query (higher is better).
Across the eight public datasets, topk-embed-v1 scores 66.33 nDCG@10 on native images and 65.09 nDCG@10 on native Markdown. It outperforms comparable models from Voyage, Cohere, Google Gemini and Mixedbread at significantly lower cost for both text and image inputs.
Frontier performance from small models
A dense model represents a document or chunk with one vector. A multi-vector model keeps several, letting different parts of a query match different parts of a document. A relevant table row or passage can stand out even when the rest of the page is about something else.
Our xsmall (0.8B parameters) model already exceeds the reported Qwen 8B dense reference in both text and image retrieval tasks. Moving from xsmall to small to the production model improves quality across both document formats and query sets.
What about multi-vector storage?
Storing hundreds of full-width vectors per page can make multi-vector retrieval expensive. We treat the representation budget as part of model-storage co-design, ensuring that we preserve retrieval quality while making storage cost effective at scale.
The smallest representations use 5.64 KiB for images and 4.87 KiB for Markdown. Both score above the corresponding dense reference on native and multilingual queries, while fitting below the 8 KiB fp16 budget of a 4,096-dimensional dense vector.
With more room, the high-quality representations stay within one point of the uncompressed reference in both query sets, while reducing storage size by 256× for images and 64× for Markdown.
Scaling text-only retrieval
On nine public RTEB tasks across Finance, Healthcare and Legal, average nDCG@10 rises from 66.72 for xsmall to 69.01 for small to 72.25 for the production model. Retrieval quality and representation efficiency both improve with model size across all domains.
Agentic search
Retrieval quality also matters when an agent makes several searches to answer a question. On BrowseComp-Plus, with GPT-5 as the agent, topk-embed-v1 reaches 82.65% accuracy with an average of 14.37 search calls per question.
Compared with the dense Qwen3 Embed 8B baseline, TopK improves accuracy by 10.96 percentage points while using 34% fewer search calls. BM25 reaches 57.59% accuracy at 23.23 calls. Mixedbread leads this comparison at 90.48% accuracy and 11.53 calls, although their search endpoint is likely a compound retriever rather than just a single embedding model.
Where we go next
We've identified several axes for improving retrieval quality, including model size, representation size, training data construction and co-design of model representations and our retrieval engine. We'll use the current generation of models to improve the future versions, delivering higher quality retrieval at the lowest cost for our customers.
Start with topk-embed-v1-xsmall or topk-embed-v1-small on Hugging Face, or use the production model inside TopK for $0.05/1M tokens.
