Six weeks of silent relevance regression
A retrieval change degrades answer quality without breaking a single test, because nothing compares today ranking against yesterday ranking. Users absorb the loss quietly.
The shape of the incident
A routine change ships: the index is rebuilt with a slightly different analyser, or a synonym list is added, or the corpus is re-imported after a parser upgrade. Every test passes, because every test asserts that the endpoint responds and that the answer contains something plausible. Six weeks later somebody reads the support queue in order and notices that the same complaint appears, phrased three different ways, every week.
There is no error to find. There was never an error. Ranking moved, and ranking is not a thing that fails a test.
Why end-to-end tests cannot see it
An answer-level assertion is a coarse instrument. If the assertion is that the response mentions a topic, it will pass on a paragraph that mentions the topic and answers the wrong question. If the assertion is a human-graded quality score, the grader usually grades the answer in isolation, not against the previous answer to the same question — so a change that makes every answer slightly less relevant produces slightly lower scores on a scale where slightly lower is indistinguishable from noise.
The measurement that would have caught it is comparative and mechanical: for a fixed set of queries, store the top results, and diff them whenever the retrieval configuration changes. A chunk that used to be first and is now absent from the top few is a signal, and it does not require a judge to interpret.
Making ranking changes diffable
- Keep a golden query set — a few hundred real queries is plenty — and store the ranked result identifiers for each one as a committed baseline.
- On every change to the index, the analyser, the embedding model or the chunker, regenerate the ranking snapshot and diff it. Report how many queries changed their top result, not only whether the average improved.
- Separate the two directions. Churn in one direction is often intended; churn in both directions, concentrated in a document family, is usually a parsing or analyser defect.
- Attach a reason to every intentional baseline update. A baseline that is refreshed without explanation is a regression that has been reclassified as normal.
- Watch the queries that changed from a correct answer to a wrong one, not the ones that changed from wrong to wrong. Aggregate scores hide directional damage.
The habit that prevents the second occurrence
Treat the retrieval configuration as a versioned artefact with its own release note: what changed, which queries moved, and which document families were affected. That single paragraph is what turns a six-week silent regression into a same-day conversation, because the person who made the change has the context and is still available to make it better.
It also settles the more common dispute before it starts. When someone claims relevance has degraded, the diff either shows it or it does not, and the conversation moves from impressions to a specific list of queries.
What to do about it
- End-to-end answer tests cannot see a ranking change that still returns a plausible paragraph
- Snapshot the top results per query so ranking changes become diffable
- A quality regression that does not throw an error will be found by users, not by CI