Playbooks

Writing the security packet for an AI feature

Security review does not fail because the system is risky. It fails because nobody could describe the risk. Five artefacts turn a three-month review into a two-week one.

What the reviewer is actually deciding

A security review is not a judgement about whether a model is accurate. It is a judgement about whether the organisation can take on the risk, and to make that judgement a reviewer needs four things: what data the system touches, where that data goes, who can see what it produces, and what happens when it is wrong. If any of those answers is a paragraph of reassurance rather than an artefact, the review extends.

None of this is specific to AI. The reason AI features are reviewed slowly is that teams rarely produce these artefacts, because the system was built to demonstrate value rather than to be described.

The five artefacts

A one-page data flow. Every source system, the boundary it crosses, what is transmitted, and whether it is retained. Name the model provider as a data processor and state explicitly whether prompts and outputs are retained or used for training, quoting the contractual term rather than the marketing page. This single page resolves most legal and privacy questions in one meeting.

A risk register. What the system can do wrong, the harm, the mitigation, and the residual. Two columns of a table. A reviewer who sees "hallucination, mitigated by citations plus a low-confidence refusal path, residual: user may still act on an uncited claim" can make a decision. A reviewer who sees "the model is reliable" cannot.

A permission demonstration. The access control model, where enforcement happens relative to retrieval, and a way to show it working. If filtering happens after retrieval, say so and explain why the ranking leak is acceptable — or fix it. Reviewers ask this question more often than any other because it is the one with a clean, checkable answer.

A retention and deletion statement. How long prompts, outputs and retrieved context are stored, where, and how a deletion request is satisfied. Include logs and evaluation datasets — evaluation sets built from real data are frequently the longest-lived copy of user content in the building.

An incident path. Who is paged, what can be switched off, and how fast. A capability switch that turns the feature off without a deployment is the strongest single answer available in a review, because it converts worst-case exposure into a bounded window.

Two habits that shorten the review

Write these before the code, not before the review. Data flow and the risk register are design documents; written afterwards they are archaeology, and they will contradict the implementation in exactly the places a reviewer probes.

Answer with the artefact, not with confidence. If the honest answer to a question is that the system can produce an unverifiable claim and the mitigation is a citation requirement that is not yet enforced, say that and give the date it will be. Reviewers approve mitigations with dates far more readily than they approve adjectives.

What to do when the answer is no

If the review blocks the feature, the useful question is which specific control would change the answer. That reframes the outcome as a scoped piece of work rather than a rejection, and it usually turns out to be one of the five artefacts missing rather than a fundamental objection. Teams that ask this question get a second review quickly; teams that appeal get a slower one.

What to do about it

  • Write down what leaves the network boundary, before writing the code
  • A risk register with named mitigations beats a reassurance
  • Demonstrate permission enforcement on request, not in a diagram