Vector Search

Vector Search and Semantic Retrieval

Text search matches tokens, so "laptop overheats" never finds "notebook runs hot". Semantic retrieval stores an embedding — a fixed-length float array from a model — beside each document and answers a query with the nearest embeddings, indexed with HNSW, a navigable small-world graph, and queried through $vectorSearch at up to 8,192 dimensions.

A vector index and a $vectorSearch query (Atlas required)
db.articles.createSearchIndex('article_vec', 'vectorSearch', {
  fields: [{ type: 'vector', path: 'embedding', numDimensions: 1536,
             similarity: 'cosine', quantization: 'scalar' },
           { type: 'filter', path: 'views' }]
});
const queryVector = await embed('why does my laptop get so hot');   // your model call
db.articles.aggregate([
  { $vectorSearch: { index: 'article_vec', path: 'embedding', queryVector,
      numCandidates: 150,   // ANN: HNSW queue size, >= 20x limit
      limit: 5,             // exact: true would run ENN over every vector instead
      filter: { views: { $gte: 100 } } } },
  { $project: { title: 1, score: { $meta: 'vectorSearchScore' } } }
]);

Both statements return SearchNotEnabled on Community, so they are shown without output. numCandidates is the knob that matters: HNSW walks that many candidates and keeps the best limit, so the documented starting point is 20 times limit, and raising it buys recall with latency. exact: true scans every vector instead: slow, but the ground truth for measuring recall. Declared filter paths apply before the graph walk, and vectorSearchScore runs 0 to 1. The embedding is your responsibility: the same model, preprocessing and numDimensions for documents and queries, or the distances mean nothing — and store the model version, because re-embedding is a migration (Zero-Downtime Data Migrations). A 1536-dimension float vector costs 6 KB, which is why quantization exists: 'scalar' stores int8 bytes, four times smaller and 3.75 times less RAM; 'binary' stores single bits, 32 times smaller and 24 times less RAM.