Playbooks

Moving a prototype onto managed inference

A prototype built against one vendor API carries assumptions that surface on a managed platform. Settle them before the first request so the migration is configuration, not a rewrite.

Assumptions that travel badly

A prototype is usually a thin wrapper around one provider. That is the right choice for a prototype and the wrong foundation for a service, because several properties that feel like facts about LLMs are actually properties of the first vendor's implementation.

  • Tokenisation. The same text is a different number of tokens on a different tokeniser. Cost estimates, context limits and truncation logic that assume otherwise break quietly, producing truncation in the middle of a retrieved passage.
  • Tool calling. The request and response shape for tools is vendor-specific, as is how strictly the model is required to produce valid arguments. A schema that one provider enforces may be advisory on another.
  • Structured output. Some platforms support enforced JSON schemas and some only encourage them in the prompt. Code that parses the response without a fallback works until the first time it does not.
  • Streaming. Chunk boundaries, usage reporting and completion signals differ, and the differences show up as intermittent rendering bugs rather than exceptions.
  • Safety filtering. Which inputs and outputs are blocked, and whether the block is an error or a substitution, varies. A prototype that never saw a refusal may see them often elsewhere.

Two source-level changes that make migration configuration

Put an adapter in front of the model call. One module that takes an internal request shape and returns an internal response shape, with one implementation per provider. The rest of the application is then written against your shape, and adding a provider is a new implementation rather than an edit in forty places.

Make the request shape explicit and versioned. Fields for the model, the context, the tools, the maximum output length, the temperature, and the correlation identifier — with the provider-specific translation happening inside the adapter. A version identifier on the shape means a change to it is a reviewable event rather than an invisible refactor.

Settle these before the first production request

  • Data handling. Whether prompts and outputs are retained, for how long, and whether they are used to improve the provider's models. This is a contractual question with a written answer, and it belongs in the security packet.
  • Rate limits and quotas. The provider's limits in the regions you serve from, and the behaviour of your client when they are hit. A client that retries on rate limit without a budget converts a soft limit into a cost incident.
  • Fallback. Which provider you use when the primary is degraded, decided in advance, with the adapter already implemented. A fallback designed during an outage is a second outage.
  • Model version pinning. Whether you pin a model version or follow the provider's latest, and who reviews the behaviour change when the provider retires the pinned one.
  • Cost attribution. Which team pays, and how usage is tagged so the answer is not a monthly reconstruction from logs.

What to migrate first

Migrate the least critical path first — an internal tool, a batch job, a low-traffic endpoint — and use it to validate the adapter, the token accounting and the fallback. The purpose is to discover the vendor-specific behaviour on a workload whose failure is cheap, before the same discoveries are made by customers.

What to do about it

  • Put a provider adapter in front of the model call before migrating
  • Treat tokenisation, tools and streaming as vendor-specific, not universal
  • Define the fallback provider before you need it