Playbooks

An on-call runbook for an LLM service

Classic latency and error alerting does not see the failures that matter here: quality decay, provider degradation, silent context growth. A runbook has to name the signals first.

Signals that classical monitoring misses

Latency, error rate and saturation catch the failures that look like infrastructure. An LLM service has a second family of failures that produce no errors and no slow requests.

  • Quality decay. A monthly sample of evaluation items, scored and trended. Without it, degradation is reported by users weeks later, in the form of distrust.
  • Refusal rate. A rising refusal rate is often the first visible sign of a retrieval or parsing problem, and it is cheaper to investigate than a complaint.
  • Context size per request. It grows silently as corpora and conversation history accumulate, and it moves cost and latency at the same time.
  • Output token count. A distribution shift here usually means a prompt or model change, and it shows up on the invoice before it shows up anywhere else.
  • Provider health. Error rate is only one signal from a provider; unexplained latency shifts and response shape changes matter too, and they are the leading edge of a provider-side incident.

Writing it so it can be followed at three in the morning

A runbook entry has four parts: the signal, the first action, the escalation condition, and the fallback. The first action must be mechanical and specific — not investigate, but: check this page, compare this number against its value from a week ago, and if the delta exceeds the recorded threshold, do this.

Fallbacks belong in the runbook because they are decisions, not improvisations. Possible fallbacks include routing to a different provider, dropping to a smaller model, shortening the context, disabling reranking, and turning off the feature entirely. Each needs a pre-agreed trigger and a named person who can authorise it.

The degraded modes worth having

  • Shorter context. Fewer retrieved passages, lower quality, much lower cost and latency. Cheap to implement and usually the first thing to reach for.
  • Smaller model. Quality drops unevenly across task types; know in advance which tasks degrade acceptably and which must fail rather than answer badly.
  • Cache-only. Serve previously computed answers and decline anything new. Rarely applicable, but decisive when the corpus is mostly stable questions.
  • Feature off. The final fallback, and the only one that must always be available. If turning the feature off requires a deployment, the runbook is not finished.

Practice that costs one hour a quarter

Exercise the degraded paths and the feature switch during business hours, and record how long each took. Untested switches do not work when they are needed, and the failure is usually trivial — a permission that only the deployment system has, or a configuration that was never exercised in the environment where it matters. One hour a quarter is the whole cost of knowing the answer.

What to do about it

  • Add quality, refusal rate and context size to the on-call signals
  • Every alert needs a first action that does not require reasoning under pressure
  • Keep a degraded mode documented and reachable without a deployment