> ## Content Index
> Fetch the complete content index at: https://www.tao.media/llms.txt
> Use this file to discover other available public pages before exploring further.

# Cohere Launches Embed 5 Pro and Fast Frontier Embedding Models
- URL: https://www.tao.media/cohere-launches-embed-5-pro-and-fast-frontier-embedding-models/
- Published: 2026-09-30T22:20:00.000Z
- Updated: 2026-09-30T22:23:19.000Z
- Description: The new embeddings family pairs a max-quality Pro model with a lower-cost Fast model in one shared embedding space for enterprise RAG, search, and agentic retrieval.
- Author: Tristan Hillerich
- Tags: News, AI, Cohere

[Cohere](https://cohere.com/?ref=tao.media) has launched [Embed 5](https://cohere.com/blog/embed-5?ref=tao.media), a new family of enterprise embedding models for search, retrieval-augmented generation, multimodal document understanding, and agentic retrieval workflows.

The release introduces two tiers: Embed 5 Pro, Cohere's highest-quality option, and Embed 5 Fast, a lighter model built for lower-latency and lower-cost workloads. Both models support 128K-token context windows, multimodal inputs, more than 100 languages, multiple output dimensions, Matryoshka embeddings, and float, int8, and binary embedding formats.

Embed 5 Pro is priced at $0.12 per million text tokens, while Embed 5 Fast costs $0.08 per million text tokens. Image embeddings cost $0.40 per million tokens for both tiers.

Pro and Fast share the same embedding space, so companies can index documents with Embed 5 Pro, then query that index with either Pro or Fast without rebuilding the vector database. That gives enterprise teams managing large retrieval systems a more flexible tradeoff between indexing quality, live-query latency, and operating cost.

Embed 5 is generally available through the Cohere API, Cohere Model Vault, Microsoft Foundry, Amazon SageMaker, and Cohere's North platform. Cohere said both models can also be self-hosted with vLLM, while batch embedding is available for large ingest jobs.

## Why Cohere Split Embed 5 Into Pro and Fast

Embedding models convert text, images, documents, or other inputs into vectors that retrieval systems can compare. They are a core part of modern enterprise AI because they determine which documents, passages, charts, tables, or records get sent into a language model before an answer is generated.

That retrieval layer matters for ordinary search, but it becomes even more important in RAG and agentic systems. If the embedding model retrieves the wrong context, a downstream language model may produce incomplete or inaccurate answers even if the generator itself is strong. If retrieval is too slow or expensive, companies may limit how often agents search, how much data they index, or how many workflows they automate.

Cohere's Pro and Fast split is aimed at that tension. Embed 5 Pro is positioned for quality-critical retrieval across complex enterprise corpora, including financial documents, parsed PDFs, multilingual datasets, code, and multimodal content. Embed 5 Fast is built for interactive search, high-volume RAG, and agent loops where every live request may require one or more embedding calls.

Cohere's recommended pattern for many customers is to index with Pro and query with Fast. In the company's testing across 40 development datasets, cross-model retrieval stayed close to same-model baselines, with average losses of 1.6% for Fast queries and 2.7% for Pro queries against mixed corpus/query pairings. Teams can pay for higher-quality document indexing once, then run cheaper and faster queries against the same vectors.

Fast also improves throughput. Cohere said Embed 5 Fast delivered an average of 2.4 times higher document throughput than Pro across tested context sizes, making it more suitable for large indexing jobs and repeated query calls inside automated workflows.

## Enterprise Retrieval Is the Benchmark Focus

Cohere is framing Embed 5 around enterprise retrieval rather than general-purpose embedding leaderboards alone. The company's launch materials emphasize financial documents, parsed PDFs, visually rich files, multilingual retrieval, and multimodal search, where document structure can matter as much as raw text.

According to Cohere's benchmarks, Embed 5 Pro achieved the highest average score among the models it tested on ViDoRe V3, a benchmark covering visually rich enterprise documents such as financial filings, technical manuals, regulatory material, government reports, textbooks, and lectures. Cohere reported an 85.8 average for Embed 5 Pro, compared with 83.7 for Voyage 4 Large, 83.2 for Gemini Embedding 2, 77.0 for Embed 4, and 75.5 for OpenAI text-embedding-3-large. Embed 5 Fast averaged 84.7 in the same category, ahead of Gemini Embedding 2 and Voyage 4 Large in Cohere's reported results.

![](https://storage.ghost.io/c/78/0b/780ba906-b1a7-4bf0-873c-bdd5c32e5331/content/images/2026/09/vidore-v3.png)

ViDoRe V3 retrieval chart | [Cohere](https://x.com/cohere/status/2105285145267687637?ref=tao.media)

The company also said Embed 5 Pro ranked first across three public financial benchmarks it highlighted: FinanceBench, FinQA, and ViDoRe V3 Finance. Fast ranked second on each, according to Cohere, despite being the lower-cost tier.

![](https://storage.ghost.io/c/78/0b/780ba906-b1a7-4bf0-873c-bdd5c32e5331/content/images/2026/09/finance-retrieval.png)

Finance retrieval chart | [Cohere](https://x.com/cohere/status/2105285149000954276?ref=tao.media)

Enterprise retrieval often fails in places that simple text search does not capture. PDFs can lose structure when parsed, tables can lose row and column relationships, charts may disappear during text extraction, and multi-column layouts can scramble context. Cohere is positioning Embed 5 as a model family built for those practical failure modes.

The company said Embed 5 is also its first model family evaluated with RCP-nDCG@10, a retrieval method Cohere introduced to score retrieved documents against query-specific relevance criteria rather than only a fixed set of labels. Cohere argues that this approach better reflects how retrieval systems behave on real enterprise corpora, where relevance can depend on the specific user question and document context.