Opens in a new tab

Perplexity releases pplx-embed-v2-context-9b-preview

01-10-2026

Perplexity released pplx-embed-v2-context-9b-preview on 30 September 2026, an MIT-licensed embedding model with open weights on Hugging Face.

Written by:

Jorick van Weelie

Marketing Lead at DataNorth | AI Enthusiast & Tech Storyteller

perplexity releases pplx embed v2 context 9b preview, an mit licensed embedding model
Sign up for our Newsletter

Perplexity released pplx-embed-v2-context-9b-preview on 30 September 2026, an open-weights embedding model under the MIT licence that scores 45.5% answer recall on context-bench, 14.4 points ahead of Voyage Context 4. It reads a document of up to 32,768 tokens in one pass and returns not just the passage holding the answer but the passages that support it.

What does pplx-embed-v2-context-9b-preview do differently?

Most retrieval systems cut a document into chunks and embed each chunk on its own. A chunk that says “it rose to 14 percent in the same period” loses its meaning the moment it leaves the page it came from, because the model never sees what “it” refers to.

This model uses late chunking: it encodes the whole document in one pass, then pools a vector per chunk from that shared reading. Every chunk keeps the context around it. Perplexity also changed the training target. Instead of marking one gold passage per question, a teacher model scores individual tokens, so the model learns to retrieve the chain of evidence behind an answer rather than the single best quote.

pplx-embed-v2-context-9b-preview benchmarks

The headline is the gap on answer recall. The evidence numbers are the ones that matter for anything an agent will cite.

What is measurespplx-embed-v2-context-9b-previewComparison
Answer recall at K=10 (a retrieved chunk holds the answer)45.5%14.4 points ahead of Voyage Context 4
Evidence recall at K=10 (retrieved chunks support that answer)40.6%No rival figure published
All-evidence recall at K=10 (every supporting chunk found)31.1%No rival figure published
ConTEB average nDCG@10 (contextual retrieval suite)Highest of the models testedPer-model scores not published
Document length read in one pass32,768 tokensChunk-by-chunk models see each chunk alone
Storage at 1024 dimensions, int81 KB per vectorMatches Voyage Context 4 at lower storage

These are Perplexity’s own evaluations, published in its research post. Perplexity does not say whether it reran Voyage Context 4 itself or took a published figure, so the 14.4 point lead is a claim with a method behind it that has not been shown. The 31.1% all-evidence number is the honest one: on two thirds of questions the model still misses at least one supporting passage.

Licence, weights and how to run it

  • Licence: MIT, which permits commercial use
  • Weights: open on Hugging Face, under perplexity-ai/pplx-embed-v2-context-9b-preview
  • Size: Perplexity describes it as 9B parameters, the model card lists 8B
  • Output: 2048 dimensions, reducible to 1024 through Matryoshka training, with native int8
  • Requirements: transformers 5.4.0 or newer, loaded with trust_remote_code set to true
  • Access: self-hosted only for now, with no Perplexity API endpoint and no published price

One usage detail will break your pipeline if you miss it. Queries and documents were trained with different prefixes, so you call encode_queries() for a question and encode() for a chunk. Mix them up and retrieval quality drops without an error to tell you why.

What Perplexity is not saying

This is a preview, and Perplexity says plainly that weights, embeddings and the interface may change without backward compatibility. In practice that means a future version can invalidate every vector you have stored. There is also no API, no price and no date for either, so the only way to use it today is to host it yourself on hardware you pay for.

The parameter count disagrees with itself across Perplexity’s own two sources, 9B in the name and the post, 8B on the model card. It is a small thing, but it is the kind of small thing that tells you how finished this release is.

What this means

Worth testing now if you run retrieval-augmented search on your own servers and you need to show users where an answer came from. A compliance or legal search team that has to produce the supporting passages, not just the answer, is the clearest fit, and the MIT licence means no procurement conversation. The thing to test first is all-evidence recall on your own documents, because 31.1% is the number your users will experience as a missing citation.

Safe to ignore if you buy retrieval as a managed service. There is no endpoint, no price and an explicit warning that the weights will change. Cohere shipped Embed 5 the same day with an API, pricing and three hosting options, and for a team without GPU capacity that is simply the usable product. Revisit this one when Perplexity drops the preview label.

For more information, visit the official announcement of pplx-embed-v2-context-9b-preview on the Perplexity blog.

Add DataNorth AI to your Google favorites