Cohere released Embed 5 on 30 September 2026, a pair of embedding models that share one vector space: Embed 5 Pro for indexing and Embed 5 Fast for queries, at $0.12 and $0.08 per million text tokens. Because they share that space, you can build your index once with Pro and answer live queries with Fast, and lose 1.6 points of retrieval quality in Cohere’s own tests.
What is new in Cohere Embed 5?
An embedding model turns text or an image into a list of numbers, so a search system can find the passages closest in meaning to a question. Until now you picked one model and used it for everything. Index quality and query speed pulled in opposite directions, and you paid for the compromise on every search.
Embed 5 splits that choice in two. Pro and Fast produce vectors in the same space, so they are interchangeable at query time against an index built by either one. Fast handles roughly 2.4 times the document throughput of Pro in Cohere’s tests. Both models take text, images, or text and images fused together, both carry a 128,000 token context window, and both cover more than 100 languages.
Cohere Embed 5 benchmarks and retrieval quality
The numbers worth reading are the document benchmarks and the mix-and-match quality loss, not the headline averages.
| What is measured | Embed 5 result | Named comparison |
|---|---|---|
| ViDoRe V3 (visually rich documents) | 85.8 (Pro) | Voyage 4 Large: 83.7, OpenAI text-embedding-3-large: 75.5 |
| FinanceBench (questions over company filings) | 80.1 (Pro) | First place in Cohere’s comparison set |
| FinQA (numbers inside financial documents) | 90.0 (Pro) | First place in Cohere’s comparison set |
| ViDoRe Finance (finance documents as images) | 85.0 (Pro) | First place in Cohere’s comparison set |
| Retrieval quality: index with Pro, query with Fast | 98.4 | Pro for both: 100 |
| Retrieval quality: index and query with Fast | 96.6 | Pro for both: 100 |
These are Cohere’s own evaluations across 40 datasets. Cohere scores them with RCP-nDCG@10, which reranks a fixed set of candidate passages rather than searching your full corpus from scratch. That is a friendlier test than production retrieval, so expect the gap between Pro and Fast to be wider on your own data than the 1.6 points suggest.
Embed 5 Pro and Embed 5 Fast pricing and access
- Embed 5 Pro: $0.12 per million text tokens, $0.40 per million image tokens
- Embed 5 Fast: $0.08 per million text tokens, $0.40 per million image tokens
- Context window: 128,000 tokens on both models
- Output dimensions: 2048, 1536, 1024, 768, 512 or 256, with float, int8 or binary vectors
- Languages: more than 100
- Where to run it: Cohere’s API, Microsoft Foundry, Amazon SageMaker, or self-hosted through Model Vault
The dimension and format choice is where the real money sits. Cohere’s worked example takes a corpus of 100 million chunks: stored as 2048-dimensional float32 vectors it needs around 819 GB, and as 256-dimensional binary vectors around 3.2 GB. That is a vector database bill falling by more than two orders of magnitude, and it is a decision you make once, at index time.
What the benchmarks do not tell you
Cohere publishes throughput but no latency. There is no millisecond figure for Pro against Fast, which is odd for a model whose selling point is speed at query time. Throughput is a batch measure and it is not what a user feels when they press enter.
Cohere also names no failure cases. The 98.4 is an average across 40 datasets, and an average hides the datasets where mixing Pro and Fast cost far more than 1.6 points. If your corpus is unusual, scanned forms, dense tables, one low-resource language, you are the tail of that average and Cohere has not said how long the tail is.
What this means
Worth testing now if you run retrieval at scale and your query volume dwarfs your indexing volume. A support search handling a million queries a day against a corpus you reindex weekly is the clearest case: you pay Pro prices once and Fast prices forever after. Run both setups side by side on your own queries for a week and measure the quality drop yourself, because 1.6 points on a reranking test is not the same number you will get.
Safe to ignore if your retrieval already works and your embedding bill is not a line item you notice. Swapping embedding models means reindexing everything, and that cost is real whatever the per-token price says. The genuinely interesting part here is not the benchmark lead over Voyage 4 Large, which is 2.1 points and will not survive the next release. It is the shared vector space, which turns index quality and query cost into two separate decisions instead of one.
For more information, visit the official announcement of Embed 5 on the Cohere blog.