Prompt versioning without a platform
You do not need a prompt management product to stop shipping prompt changes blind. A file, a version identifier and one rule about out-of-band edits cover most of what a platform sells.
The problem with editing prompts in a console
An editable prompt in a vendor console is the highest-velocity change path in most AI systems and the least recorded one. A wording tweak is saved during an incident, behaviour changes for every user immediately, and nothing in the code repository, the review history or the release notes reflects it. When quality drops a month later, the most likely explanation is invisible.
A platform is one solution. The version of the solution that fits a small team is a set of conventions, and it takes less time to adopt than procuring something.
Four conventions that cover the essentials
Prompts live in the repository. One file per prompt, or one directory with a file per prompt, in the same review process as code. If the runtime loads prompts from a service, that service is populated by a deployment from the repository — not by a person typing into a form.
Every prompt carries a version identifier. A short hash of the content is enough, and it has the useful property of being impossible to forget to update. The identifier is logged with every response, which is what makes a regression attributable to a prompt rather than to a model update, a corpus change or a code path.
A change is accepted on an evaluation delta. The commit that changes a prompt states which evaluation items moved and in which direction. This is the only mechanism that prevents the accumulation of tweaks that each fixed one case and broke another.
Out-of-band edits are treated as incidents. If somebody edits the live prompt to stop an outage, the same edit must land in the repository before the incident is closed. Without this rule the repository diverges from production within one incident, and every subsequent version identifier is a lie.
What to record per response
- The prompt version identifier, and the model version actually served rather than the one requested.
- The parameters that affect output: temperature, output limit, any tool definitions.
- A reference to the retrieved context, so a quality complaint can be reproduced rather than discussed.
- The decision path, if the request was routed or degraded — a smaller model answering because the large one was unavailable looks identical in the response.
The comparison that justifies the effort
When quality changes, the first question is always whether the model, the prompt or the data changed. Model updates are announced, data changes are usually loud. Prompt changes are the quiet one, and with the identifier in the logs the answer takes minutes instead of a day of bisecting deployments.
The second benefit arrives at review time. A prompt diff with an attached evaluation delta is reviewable; a prompt diff alone invites an argument about wording, which nobody wins and which delays the change.
What to do about it
- Prompts belong in the repository, not in a console
- Log the prompt version with every response
- Every prompt change needs the evaluation delta it was accepted on