Voyage AI introduced the Rerank 3 series, the next generation of Voyage rerankers.
Rerankers are the second stage of a retrieval pipeline. The first stage, usually embeddings and vector search, selects a list of candidates from a large corpus. The reranker then scores each candidate against the query and reorders the list so that the most relevant documents land at the top. Because the reranker only sees a short list, it can afford to be more accurate per document than an embedding model, and it can be added to an existing retrieval system with a single API call and no re-indexing. For an introduction to rerankers, see our earlier post. For a comparison against LLMs used as rerankers, see this post.
Since its release in August 2025,
Today, we are excited to announce
Improvements on long documents and code
The Rerank 3 series is designed as an upgrade to Rerank 2.5 that requires no changes to your code. The API, the 32K token context length, instruction-following, and pricing are the same. Relevance scores are calibrated to match the score distributions of the corresponding Rerank 2.5 models, so score thresholds tuned on Rerank 2.5 continue to work.
Long documents.
Code. As we noted in the voyage-code-4 release, coding agents now issue many of the code retrieval queries we serve. On our code retrieval datasets,
Evaluation Details
Datasets. We evaluate across 9 domains: technical documentation, code, law, finance, web reviews, multilingual, long documents, medical, and conversations, 95 datasets in total. The multilingual domain is composed of 51 datasets from 31 languages. Detailed information about each of the domains and languages can be found in the rerank-2 release blog. To evaluate instruction-following, we use the MAIR benchmark, which consists of 123 tasks with task-specific instructions. Examples of instruction-following can be found in the Rerank 2.5 release blog.
Method and Metrics. We evaluate the retrieval quality of various rerankers on top of four first-stage search methods: (1) lexical search with BM25, (2) OpenAI v3 large (text-embedding-3-large), (3)
Baselines: We compare our models against
Results
Results across domains. The first bar chart below shows the average accuracy of each reranker across the 9 domains. Specifically:
- Averaged across the four first-stage retrieval methods, rerank-3 outperforms Cohere Rerank v4.0 Pro, Cohere Rerank v4.0 Fast, Qwen3-Reranker-8B, and rerank-2.5 by 2.72%, 5.74%, 3.02%, and 0.96%, respectively.
- rerank-3-lite, while optimized for latency, outperforms Cohere Rerank v4.0 Pro, Cohere Rerank v4.0 Fast, Qwen3-Reranker-8B, and rerank-2.5-lite by 2.23%, 5.25%, 2.53%, and 1.10%, respectively.
- rerank-3-lite outperforms Qwen3-Reranker-8B, the leading open-weights reranker, despite being over an order of magnitude smaller.
- Both models provide a significant quality improvement on top of all first-stage retrieval results. rerank-3 improves NDCG@10 by 7.65% atop voyage-3-large, 6.38% atop voyage-4-large, 16.67% atop OpenAI v3 large, and 20.33% atop BM25.
Long documents. On LongEmbed,
Code. On our code datasets (DS-1000, APPS, CodeChef C++, RepoBench Java, and WikiSQL),
Multilingual. Averaged across the four first-stage retrieval methods,
Instruction-following. Both models retain the instruction-following capability introduced in Rerank 2.5. On MAIR,
Comparison with Jev. We also evaluate the very recent Jev model in domains requiring strong expertise. Across 26 datasets,
Detailed results. Numeric results for all evaluations are available in this spreadsheet.
Try rerank-3 and rerank-3-lite today!
As our results show, combining Voyage embedding models with Voyage rerankers delivers the highest possible retrieval accuracy.
Both
For new users, head over to our docs to get started and learn more; the first 200M tokens are free.
Follow us on Twitter and LinkedIn to stay up-to-date with our latest releases.
Contributors
Tengyu Ma, Project Advisor
Minghan Li, Project Lead
In collaboration with the Voyage team at MongoDB.