Case Studies

The on-prem box that ran out of VRAM at month two

Capacity planning for self-hosted inference usually counts the weights and forgets everything else that has to live on the same card. The gap between those two numbers is what runs out first.

How it presents

A self-hosted model is deployed on hardware sized from a straightforward calculation: parameter count times bytes per parameter, with room left over. It works in testing. Two months later it begins returning errors under load, then falls back to a much smaller model, then serves one request at a time.

The card has not changed. What changed is everything else competing for the same memory.

What actually occupies the card

Weights are the term everyone plans for and, on a busy server, often not the largest consumer.

The key-value cache. This grows with concurrent requests and with context length, and it is the term that turns a fixed model into a variable footprint. Long conversations and large retrieved contexts multiply it. It is the usual culprit and the usual omission.

Activation memory and runtime overhead. Framework allocations, graph optimisations and memory pools are not free, and they vary with execution mode, batch size and the number of concurrent sequences in flight.

Fragmentation. Allocating and freeing variable-sized blocks for weeks leaves gaps that cannot be reused. A server that ran comfortably on day one can fail to allocate an identical request on day sixty with no change in the load profile.

Everything else on the host. An embedding model, a reranker, a vision encoder, a monitoring agent, an experimental deployment. Each is small in isolation and the set of them is not.

Sizing it without a surprise

  • Size for the peak you can actually justify, not the average. The failure mode is a peak, and peaks are what users remember.
  • Bound context length at the server rather than by convention. An unbounded context is an unbounded cache, and one client that resends history every turn will find the limit for you.
  • Quantise deliberately rather than by default. It is a real lever on hardware you own, but measure quality on your own evaluation set before and after — the trade is not uniform across tasks.
  • Reserve headroom for the non-model workloads and write the reservation down, so the reranker planned for next quarter has somewhere to live.
  • Track memory over time, not only at startup. A slow upward drift is fragmentation or a leak, and it will be discovered by an outage if nobody is looking.
  • Define a degraded mode in advance — fewer concurrent requests, a shorter context, a smaller model — and attach it to an explicit threshold instead of improvising during an incident.

The calculation nobody writes down

The useful artefact is not the initial sizing sheet, it is the record of what the headroom was spent on. Hardware that was provisioned with room to spare stops having room to spare one experiment at a time, and each individual decision looks reasonable when it is made. Keeping the reservation list alongside the deployment turns a series of small approvals into a visible budget.

What to do about it

  • Size the key-value cache, not just the weights
  • An unbounded context is an unbounded memory footprint
  • Watch memory over weeks — fragmentation shows up as an outage, not an alert