All concepts

Filtered Vector Search

Combining a metadata filter with an ANN graph is where recall quietly collapses — and why pre-, post-, and in-filter search are different products.

Advanced Vector Search · Advanced · ~6 min

In plain English

You know a shortcut through the city, but tonight most streets are closed. The shortcut isn't a shortcut any more — and you may not arrive at all.

Why it's worth your time

Every real query is filtered by tenant, permission, language or date, and this is where benchmark recall and production recall part company.

If you remember three things

  • Post-filter starves, pre-filter scales badly, in-filter breaks the graph
  • Selectivity decides which strategy is right — index the fields so the engine can tell
  • A stable high-cardinality filter should be a partition, not a predicate

Overview

Almost every real query is filtered: this tenant, this language, documents after this date, permissions this user holds. That interacts badly with ANN. Post-filtering searches first and drops non-matching results, which returns too few rows when the filter is selective. Pre-filtering computes the matching set and brute-forces it, which is exact but degrades as the set grows. In-filter (filtered) search evaluates the predicate during graph traversal, which is right in principle but breaks HNSW's connectivity: once most neighbours are filtered out, the greedy walk gets stranded in a disconnected region and recall falls off a cliff. Modern engines handle this with selectivity-based strategy switching and techniques like ACORN's predicate-agnostic neighbour expansion.

In an interview

Filtered vector search is where ANN indexes break. Post-filtering returns too few results when the filter is selective; pre-filtering is exact but scales with the matching set; filtering during graph traversal disconnects the HNSW graph and drops recall sharply. Production engines pick a strategy from estimated selectivity, and for a stable high-cardinality filter like tenant id, physically partitioning the index beats filtering it.

Production defaults

Rule of thumb
matching set < ~1% of the corpus → brute-force the filtered set. Above that → filtered traversal with a raised ef
ef
raise hnsw_ef to 2-4× the unfiltered value when filters are selective; it buys back much of the lost recall
Tenancy
bounded tenant count → one collection or namespace per tenant. It deletes the hardest case rather than tuning it

What breaks

  • Returns fewer than k results — Post-filtering on a selective predicate. Switch to a filter-aware search rather than raising k and hoping.
  • Recall fine in tests, poor in prod — You benchmarked unfiltered. Re-measure with the real filters attached — that is the only number that matters.

Watch it explained

21. Vector Search Optimization: Pre-filter, Re-ranking, & Metadata Filtering Explained — SH AI Academy, 7:14

Related