Combining a metadata filter with an ANN graph is where recall quietly collapses — and why pre-, post-, and in-filter search are different products.
You know a shortcut through the city, but tonight most streets are closed. The shortcut isn't a shortcut any more — and you may not arrive at all.
Every real query is filtered by tenant, permission, language or date, and this is where benchmark recall and production recall part company.
Almost every real query is filtered: this tenant, this language, documents after this date, permissions this user holds. That interacts badly with ANN. Post-filtering searches first and drops non-matching results, which returns too few rows when the filter is selective. Pre-filtering computes the matching set and brute-forces it, which is exact but degrades as the set grows. In-filter (filtered) search evaluates the predicate during graph traversal, which is right in principle but breaks HNSW's connectivity: once most neighbours are filtered out, the greedy walk gets stranded in a disconnected region and recall falls off a cliff. Modern engines handle this with selectivity-based strategy switching and techniques like ACORN's predicate-agnostic neighbour expansion.
Filtered vector search is where ANN indexes break. Post-filtering returns too few results when the filter is selective; pre-filtering is exact but scales with the matching set; filtering during graph traversal disconnects the HNSW graph and drops recall sharply. Production engines pick a strategy from estimated selectivity, and for a stable high-cardinality filter like tenant id, physically partitioning the index beats filtering it.
21. Vector Search Optimization: Pre-filter, Re-ranking, & Metadata Filtering Explained — SH AI Academy, 7:14