The data owner nobody assigned
AI features fail on the organisational boundary between the team that builds the system and the team that owns the data. That boundary is rarely named in the project plan, and it is where the schedule goes.
The failure that looks like an engineering problem
The system is built, the pipeline runs, and the evaluation is acceptable. Then the corpus drifts: a source system changes a field, a folder is reorganised, a document set stops being updated because the person who maintained it moved teams. Answer quality decays over weeks, and the engineering team cannot fix it, because the fix is not code.
Nothing in this story is hard. It is simply unowned.
Why ownership is the actual dependency
An AI feature built on enterprise data has a dependency on people outside the build team, and those dependencies behave differently from software dependencies. They are not versioned, they do not fail loudly, and they change when an organisation changes. Three of them account for most of the schedule risk.
Refresh. Who is responsible for the corpus reflecting the current state of the business, on what cadence, and how would anybody notice if it stopped?
Schema. Which fields exist, which are deprecated, and who is told when one changes meaning. A renamed field is a silent retrieval failure, not a crash.
Entitlement. The mapping between the source system's permission model and the index's access control list has to be maintained by someone with authority over the source. That is rarely an engineer.
Making it explicit
- Name a person, not a team, for each data source, and name a backup. Unstaffed ownership is unowned.
- Write the refresh cadence into the project plan as a dependency with a date, so a stalled refresh shows up in the same place as a stalled API integration.
- Agree an alerting channel for schema changes before ingestion begins, not after the first silent failure.
- Put corpus freshness on the same dashboard as latency and cost, so quality decay is visible before users report it as a quality problem.
- Re-validate ownership at every reorganisation. Ownership agreements survive only as long as the people who signed them.
Why it is worth the paperwork
Two documents change the outcome: a one-page data flow naming every source, its owner and its cadence, and a risk register listing what the system can do wrong with the mitigation for each. Both are boring, both take an afternoon, and both convert a class of problem that would have been discovered at month four into a set of decisions that can be made at week one.
What to do about it
- Every data source needs a named owner and an agreed refresh cadence
- Permission mapping is an organisational task with a technical surface
- Written agreements survive reorganisation; verbal ones do not