Measuring recall v0

ivfplus is an approximate index: instead of scanning every row, it only checks the lists closest to your query, so a search can miss some of the true nearest neighbors. Recall measures how often it finds them anyway — the fraction of the exact top-k results that the approximate index actually returns. Before relying on ivfplus in production, measure its recall against an exact scan.

Run the following in psql, replacing [...] with a real query vector matching your table's dimensionality. First, establish the exact result to compare against: disabling index scans forces a full sequential scan and sort, which is slow but guaranteed exact.

\set q '[...]'

BEGIN;
SET LOCAL enable_indexscan = off;
CREATE TEMP TABLE truth AS
SELECT id FROM items ORDER BY embedding <-> :'q' LIMIT 10;
COMMIT;

Then run the same query through the ivfplus index at the probes setting you want to test, and compare the two result sets:

SET ivfplus.probes = 32;
CREATE TEMP TABLE got AS
SELECT id FROM items ORDER BY embedding <-> :'q' LIMIT 10;

SELECT count(*)::float / 10 AS recall FROM truth JOIN got USING (id);

The recall column in the query above is this value already: the count of rows that appear in both truth and got (the join keeps only the ids present in both, so count(*) is the number of hits), divided by LIMIT (10 in this example). If you change LIMIT, update the divisor in the query to match.

Average the result over at least 50 sampled query vectors, and re-measure after any lists change. There's no single target recall: exploratory search can often tolerate 0.8 or lower, while applications that depend on finding near-exact matches usually want 0.95 or higher. Decide what your use case needs, then see Tuning and sizing for how to get there by raising probes.

Alternatively, use a benchmarking tool such as vsbt, EDB's vector-search benchmark toolkit, to measure recall on its built-in datasets or your own.