Today we're introducing topk-embed-v1, a family of multi-vector embedding models for text and visual documents. Two smaller models, xsmall and small, are open-sourced, while the production-size model is offered inside TopK for $0.05/1M tokens.

We co-designed the models and storage to enable efficient multi-vector retrieval at scale. Our production model scores 65.16 nDCG@10 at 11.27 KiB per page on native ViDoRe v3 image queries, delivering strong retrieval quality at storage size comparable to dense embeddings.

Highest quality for $0.05/1M tokens

ViDoRe v3 tests retrieval over documents such as financial reports, scientific papers and industrial manuals. We evaluate page images, which preserve layout, tables and figures, and Markdown, which tests retrieval from the text. We measure quality with nDCG@10: how well the ten highest-ranked results match the query (higher is better).

Across the eight public datasets, topk-embed-v1 scores 66.33 nDCG@10 on native images and 65.09 nDCG@10 on native Markdown. It outperforms comparable models from Voyage, Cohere, Google Gemini and Mixedbread at significantly lower cost for both text and image inputs.

Retrieval quality versus cost

ViDoRe v3 · managed providers · upper left is better

TopKCohereVoyageGeminiMixedbread
Managed provider input cost per 1,000 images versus nDCG@10nDCG@10455055606570$0.06$0.1$0.2$0.5$1$2$4$/1k imagesTopKCohereVoyageGeminiMixedbread
View chart data and sources
topk-embed-v1 · production$0.075 / 1,000 images66.33 nDCG@10
Managed providerUSD / 1,000 imagesnDCG@10
topk-embed-v1 · production$0.07566.33
Cohere Embed 4≈$2.4058.23
Voyage Multimodal 3.5$0.6057.54
Gemini Embedding 2$0.1256.00
Mixedbread Wholembed V3$3.0064.24

nDCG@10 on a 0–100 scale, averaged equally across eight public ViDoRe v3 datasets. Native uses each dataset’s original query language; multilingual includes all six languages, including native.

TopK: $0.075/1,000 images. Gemini: $0.00012/image. Mixedbread: 2,000 content tokens/image at $1.50/M. Voyage: one-megapixel images at $0.60/billion pixels. Cohere is an estimate for one-megapixel images, applying AWS’s token formula (pixels ÷ 784 × 4) to Cohere’s $0.47/M image-token rate; These are billing estimates, not measured benchmark bills.

Rates: Google, Cohere, Voyage, Mixedbread (September 24, 2026). Mixedbread’s $1.50/M rate includes indexing.

Frontier performance from small models

A dense model represents a document or chunk with one vector. A multi-vector model keeps several, letting different parts of a query match different parts of a document. A relevant table row or passage can stand out even when the rest of the page is about something else.

Our xsmall (0.8B parameters) model already exceeds the reported Qwen 8B dense reference in both text and image retrieval tasks. Moving from xsmall to small to the production model improves quality across both document formats and query sets.

Quality scales with model size

Three sizes. The same evaluation recipe.

xsmallsmallproduction
Model size versus retrieval quality. Open the data table below to inspect measurements.nDCG@105055606570xsmallsmallproductionModel sizeQwen3-VL Embed 8B · 61.11Qwen3-VL Embed 8B62.8865.2266.67
View chart data and sources
production66.67 nDCG@10
MeasurementModel sizenDCG@10
xsmallxsmall62.88
smallsmall65.22
productionproduction66.67

nDCG@10 on a 0–100 scale, averaged equally across eight public ViDoRe v3 datasets. Native uses each dataset’s original query language; multilingual includes all six languages, including native.

What about multi-vector storage?

Storing hundreds of full-width vectors per page can make multi-vector retrieval expensive. We treat the representation budget as part of model-storage co-design, ensuring that we preserve retrieval quality while making storage cost effective at scale.

Multi-vector, near dense storage

Selected production representations · upper left is better

SmallestCompactBalancedHigh qualityQwen3-VL 8B
Storage (KiB/page) versus retrieval quality. Open the data table below to inspect measurements.nDCG@1050545862667001020304050Storage (KiB/page)Dense fp16 · 8 KiB · 8.0064.0165.1665.7266.01Qwen3-VL 8B · 61.11 nDCG@10
View chart data and sources
Smallest5.64 KiB/page64.01 nDCG@10
MeasurementStorage (KiB/page)nDCG@10
Smallest5.6464.01
Compact11.2765.16
Balanced23.9565.72
High quality45.0966.01
Qwen3-VL 8B8.0061.11

nDCG@10 on a 0–100 scale, averaged equally across eight public ViDoRe v3 datasets. Native uses each dataset’s original query language; multilingual includes all six languages, including native.

Average vector payload across 19,252 pages. Qwen’s reported score is paired with an assumed 4096-dim fp16 payload (8 KiB).

The smallest representations use 5.64 KiB for images and 4.87 KiB for Markdown. Both score above the corresponding dense reference on native and multilingual queries, while fitting below the 8 KiB fp16 budget of a 4,096-dimensional dense vector.

With more room, the high-quality representations stay within one point of the uncompressed reference in both query sets, while reducing storage size by 256× for images and 64× for Markdown.

Scaling text-only retrieval

On nine public RTEB tasks across Finance, Healthcare and Legal, average nDCG@10 rises from 66.72 for xsmall to 69.01 for small to 72.25 for the production model. Retrieval quality and representation efficiency both improve with model size across all domains.

Text retrieval by domain

RTEB · 9 selected tasks

xsmallsmallproduction
Model size versus retrieval quality. Open the data table below to inspect measurements.nDCG@105060708090xsmallsmallproductionModel size66.7269.0172.25
View chart data and sources
production72.25 nDCG@10
MeasurementModel sizenDCG@10
xsmallxsmall66.72
smallsmall69.01
productionproduction72.25

Full-width representations; nDCG@10 on a 0–100 scale. Overall averages these nine tasks equally, excluding code and encyclopaedic tasks. Scores average queries within each source, then sources within each task, including all query languages.

Finance: FinanceBench, HC3 Finance, FinQA. Healthcare: ChatDoctor, CURE. Legal: AILA Casedocs, AILA Statutes, LegalSummarization, LegalQuAD.

Retrieval quality also matters when an agent makes several searches to answer a question. On BrowseComp-Plus, with GPT-5 as the agent, topk-embed-v1 reaches 82.65% accuracy with an average of 14.37 search calls per question.

Agentic search

BrowseComp-Plus · GPT-5 agent · upper left is better

MixedbreadTopKqwen3-embed-8B (dense)BM25 (sparse)
Tool calls per question versus Accuracy (%). Compare accuracy and search calls per question.Accuracy (%)50607080901001014182226Tool calls per question90.48%82.65%71.69%57.59%
TopK (topk-embed-v1)14.37 search calls82.65 % accuracy
View chart data and sources
MeasurementTool calls per questionAccuracy (%)
Mixedbread11.5390.48
TopK (topk-embed-v1)14.3782.65
qwen3-embed-8B (dense)21.7471.69
BM25 (sparse)23.2357.59

BrowseComp-Plus results with GPT-5 as the agent. Accuracy is the percentage of correctly answered questions; search calls are averaged per question.

Compared with the dense Qwen3 Embed 8B baseline, TopK improves accuracy by 10.96 percentage points while using 34% fewer search calls. BM25 reaches 57.59% accuracy at 23.23 calls. Mixedbread leads this comparison at 90.48% accuracy and 11.53 calls, although their search endpoint is likely a compound retriever rather than just a single embedding model.

Where we go next

We've identified several axes for improving retrieval quality, including model size, representation size, training data construction and co-design of model representations and our retrieval engine. We'll use the current generation of models to improve the future versions, delivering higher quality retrieval at the lowest cost for our customers.

Start with topk-embed-v1-xsmall or topk-embed-v1-small on Hugging Face, or use the production model inside TopK for $0.05/1M tokens.