Playbooks

How to build an evaluation set from real traffic

Most evaluation sets are written from documentation and measure the documentation. Building one from logs, tickets and abandonments takes a week and is the highest-leverage task available.

Why not to write questions from the documentation

Documentation describes the system the way its authors understand it. Questions derived from it share the authors' assumptions, appear in the retrieval corpus as paraphrases, and cluster around the parts of the product that are well documented rather than the parts that are used. The resulting suite is easy to pass and slow to change.

Real traffic has the opposite distribution, which is why it is worth the effort of collecting.

Four sources, in order of usefulness

Search logs. The queries users actually typed, including the ones that returned nothing. Zero-result queries are the cheapest source of high-value evaluation items that exist, because they are pre-filtered for topics the system does not yet handle.

Support tickets and internal questions. Every time a human was asked and answered a question the system should have handled, that is a question with a known-correct answer and an identified cost. These are also the items most likely to be phrased the way a real user phrases things.

Abandoned sessions. Users who started, typed something, and left without completing. Harder to interpret, but they mark the boundary of where the system currently fails in a way satisfaction data cannot.

Expert recall, deliberately limited. Sit with two or three people who do the work and ask what they are asked most often and what the answer depends on. Use this to fill gaps the logs cannot show, such as questions asked in meetings and never typed. Cap the time spent here; its purpose is coverage, not volume.

Writing items that survive regrading

  • Store the question exactly as the user phrased it, including the typos and shorthand. Rewriting it into clean language removes the part that is hard.
  • State the expected answer and the reason it is correct. The reason is what makes the item regradable by somebody else and keeps the criteria from drifting silently.
  • Record what the correct behaviour is when the answer is not in the corpus. Roughly a fifth of items should be unanswerable, and refusing cleanly should score as a pass.
  • Tag each item with the capability it exercises — retrieval, arithmetic over retrieved values, multi-document synthesis, refusal — so failures point at a component instead of at the system.
  • Note the source and the date. Items drawn from a transient incident stop being valid when the incident's data leaves the corpus.

Running it as a habit

Sample monthly, not once. Review the suite quarterly and retire what no longer reflects the product. Keep one human-graded slice permanently, because it is the only part that notices when the definition of a good answer shifts.

The first run will be uncomfortable. A suite built from traffic typically reports substantially lower quality than the documentation-derived suite it replaces, and that gap is the amount of quality the old suite was hiding. Report both numbers side by side, attribute the difference to the change in criteria, and then work from the honest baseline.

What to do about it

  • Mine four sources: search logs, support tickets, abandoned sessions, expert recall
  • Every item needs a stated reason, not just an expected answer
  • Include unanswerable questions and score refusal as a pass condition