Playbooks

How to choose a reranker without a benchmark harness

You do not need a public leaderboard to choose a reranker. You need queries with known correct passages, a way to measure position, and the discipline to include unanswerable queries.

What a reranker changes

A first-stage retriever is optimised to be fast over a large candidate set, which means it approximates relevance. A reranker reads the query and the passage together and scores them properly, over a much smaller set. It improves the order of what was already found. It cannot find what the first stage missed.

That single property sets the evaluation: a reranker can only be judged on queries where the correct passage is already in the candidate set. If it is not, no reranker helps and the work belongs upstream.

The minimum viable harness

You need a query set, the correct passage identifier for each query, and a way to record the rank at which that passage appears with and without the reranker. For a few hundred queries this can be a script that writes two ordered lists per query and prints the differences. It does not need a framework.

  • Record whether the correct passage appears in the top results at all, and at what rank. Track both, because a reranker that improves the mean rank while pushing the passage out of the window has made things worse.
  • Include queries with no correct passage. A reranker will still return its best guess, and the useful measurement is whether the top score of a reranked miss stays visible as low confidence or comes out indistinguishable from a hit.
  • Keep query difficulty in the set. Queries that work without reranking contribute nothing to the comparison beyond latency.
  • Store the results as a committed baseline so a future change to the reranker or to the first stage is measured against something.

The trade nobody writes down

A cross-encoder reranker adds latency proportional to the number of candidates it scores, and that cost sits on the critical path of every request. Two questions should be answered before adoption, in writing: how many candidates will be reranked, and how much median latency is acceptable. A reranker tuned on quality alone tends to end up reranking a candidate set large enough to blow the latency budget, at which point the quality gain is spent on timeouts and retries.

Public benchmarks and why they mislead here

Public retrieval benchmarks describe a corpus and a query distribution you do not have. A model that leads on general text can trail on your documents if your corpus is structured — clauses, tables, specifications — because the training distribution rarely looks like a policy document. Treat a leaderboard as a shortlist, then decide on your own corpus with your own queries, including the refusal cases.

When not to add one

If the first stage already returns the correct passage at the top for most queries, a reranker is optimising a solved case at the cost of latency for every request. If the failures are concentrated in the tail of hard queries, the higher-leverage change is usually chunking or parsing. A reranker is the right tool when the first stage finds the passage but ranks it below other, superficially similar text — and that is a condition you can measure before committing.

What to do about it

  • Measure recall at the cutoff and mean rank of the correct passage
  • Always measure the latency the reranker adds, not only the quality
  • Test on the corpus you actually serve, not on a public benchmark