AI Governance Best Practices
AI governance is the governance practice that data teams already run, including ownership, access control, lineage, quality, and audit, now extended to cover a new class of consumer: agents, copilots, and MCP-connected tools that query your data on behalf of a person or a workflow.
The mistake most organizations make is treating AI governance as a policy problem: a new document, a new committee, a new set of guidelines for "responsible AI use." A policy document does not check permissions at query time or verify which table an agent selected, so it will not stop an agent from picking the wrong one. The controls that do that are the same operational controls you already use for human users, access control, ownership records, lineage, and quality checks, applied consistently to machine users as well. If your governance program can't answer "which agents have access to what, and what did they do with it," that is the gap to close first, before adding new policy language.
Key Takeaways
- AI governance extends existing governance practice, ownership, access, lineage, quality, audit, to agents as a new consumer class. It isn't a separate discipline.
- Permission-aware retrieval, where an agent only sees what its account is authorized to see, is the single most important control, but only when the agent is actually configured to inherit user-scoped credentials rather than a shared service account.
- An unregistered agent or MCP server is the most common governance gap teams find, not a missing policy document.
- Structure (context, semantics, memory), not model size, separates reliable agent answers from confident guesses.
- Roll out through two or three high-stakes domains first, and measure discovery time, disputes, and incident reopens, not policy-doc completion.
How AI governance can impact the business
Consider what happened at OpenAI, one of the most technically capable AI organizations in the world, while building an internal data agent. The team's own writeup describes it plainly: an AI agent incorrectly answered 5,062,338 when the real number was 800 million. The agent had the right tables and a working query engine. What it didn't have was any way to know that the metric it grabbed was defined differently than the one the business actually meant.
This kind of failure shows up wherever metadata is incomplete or inconsistent, and it isn't limited to one company or one query. Any data platform leader who has watched two teams pull different tables for the same KPI and still present confident slides has seen an early version of the same problem. A human analyst who hits a table name that looks slightly off, or an ambiguous metric like "active_users," will usually pause and ask a coworker what it means. An agent has no equivalent pause built in by default: without explicit signals about definitions, freshness, and ownership, it will return an answer with the same confidence whether that answer is right or off by orders of magnitude.
Six practices that make AI governance operational
1. Ownership resolves at query time, not in a spreadsheet
Ownership metadata that lives in a wiki page or a quarterly RACI spreadsheet is stale by the time anyone reads it. For an agent, ownership needs to resolve the moment a query runs, so the agent, or the person reviewing its answer, can see who's accountable for a given table, dashboard, or metric right now. This is what the shared metadata infrastructure that connects data assets to business meaning, quality, lineage, and ownership is for: a live system of record that updates as ownership changes, rather than a document that goes out of date the week after it ships.
2. Access control is inherited, not bypassed
This is the single most important control in the entire practice. An agent acting on a user's behalf should see exactly what that user is authorized to see, nothing more. In practice this means permission-aware retrieval, where an agent only sees what its account is authorized to see, enforced by the same role-based access controls and authorization engine as the rest of the platform. That authorization engine only closes the gap if the agent or MCP client is actually configured to pass through the requesting user's own credentials, rather than authenticating as a shared service account, which is an integration decision your team makes, not something the platform enforces on its own. Check this configuration directly: if an agent queries through a service account with broader access than any individual user has, every query it runs on someone's behalf executes under permissions that person never actually had, which is a compliance exposure to fix before the agent goes into wider use.
3. Lineage explains provenance before a citation is trusted
When an agent cites a number, the reader needs a way to check where that number came from before they act on it, not after. Column-level lineage that explains where a metric came from and what depends on it turns an agent's answer from a black box into something a data engineer can verify in seconds: which source table, which transformation, which downstream reports would break if this changed. Without that lineage record, verifying an agent's answer means manually retracing the query logic each time, which does not scale past a handful of checks a week.
4. Quality signals are exposed where agents look
An agent has no instinct for "this table looks stale" the way a human analyst does. It needs the signal made explicit, with freshness, test results, and severity levels exposed where agents look, in the same metadata layer the agent queries against, not in a separate quality dashboard nobody wired the agent to read. If a table failed its freshness check an hour ago, the agent needs that fact sitting next to the table it's about to query, not two clicks away in a tool it was never given access to.
5. A registered inventory of agents, tools, and LLMs
The most common gap in AI governance today isn't a broken control. It's an unregistered agent or an MCP server nobody logged. Someone connects a new copilot to the warehouse on a Friday afternoon to unblock a project, and six months later nobody on the data team can produce a list of every AI system with a live connection to production data. Build and maintain that list before adding new integrations: every agent, every MCP server, every LLM integration gets registered with an owner, a scope, and a review date, the same way you'd register a new service account or a new BI tool. The same inventory is also where you'd notice if write access to descriptions, tags, or glossary terms, the same fields agents are asked to trust, isn't itself access-controlled and audited; an editable metadata field an agent reads as ground truth is a channel worth locking down.
6. An auditable memory of corrections
When a human analyst gets corrected, "that metric excludes churned accounts," they remember it, and the correction quietly becomes part of how they work. When an agent gets corrected, the correction needs to be written down somewhere durable and governed, not just retained in one session's context window, so the next agent, or the same agent next week, doesn't make the identical mistake. Without a governed memory layer, each correction has to be re-taught the next time the same question comes up, since nothing was recorded for the next session to reference. This is one reason an AI agent will make an inference instead of catching the error a human analyst would: it has no memory of the exception someone explained last month unless that exception was captured somewhere the agent actually reads.
A rollout sequence that works
Don't try to govern every agent, every dataset, and every workflow in one initiative. Pick two or three high-stakes domains where a wrong answer is expensive, finance reporting, customer-facing metrics, regulatory submissions, and apply all six practices there first. Set measurable coverage goals instead of a full policy rewrite: what percentage of tables in this domain have registered owners, what percentage of agent queries pass through permission-aware retrieval, what percentage of key metrics have lineage documented back to source.
Then measure three things as you expand: discovery time, how long it takes someone to find the right table or metric, disputes, how often two people or an agent and a person disagree about a number, and incident reopens, how often a resolved data issue resurfaces because the underlying cause was never fixed. If those three numbers improve as you add domains, the six practices are being applied correctly. If they don't move, check access control and the registered inventory first, since gaps there are the most common cause of stalled progress.
What this looks like in production
The OpenAI case study is instructive because the fix wasn't a bigger model. It was structure: data a model or agent can use without guessing what a column means, built from a context layer that carries semantics, not just column names. Anthropic's own research backs this up from a different angle. Anthropic saw accuracy jump from 21% to 95%+ after adding structure, a semantic layer and skills, not from giving the agent broader raw access. When the team tested a version where the agent had more grep-style access but no added structure, accuracy moved less than one point when the shortcut was more raw access instead of structure. Structure, meaning context, semantics, and memory together, was the variable that separated a reliable agent from a confidently wrong one in both cases; model size and raw access were not.
Frequently asked questions
Is AI governance a new team or a new function?
No. It's an extension of the data governance function you already run, applied to a new consumer class. The people who own access control, lineage, and quality for human users are the right people to own it for agents.
Do we need a full AI policy before we start?
No. Start with two or three domains, apply the six operational practices, and measure discovery time, disputes, and incident reopens. A policy document that isn't backed by permission-aware retrieval and a registered agent inventory will not catch the specific failure that produced the 5,062,338 answer, since that required checking permissions and metric definitions at query time, not a written policy.
What's the single highest-priority control if we can only do one thing first?
Permission-aware retrieval. An agent that inherits a user's actual permissions, rather than querying through an overprivileged service account, closes the largest and most common security gap in agent deployments. This depends on the agent or MCP client being configured to pass through the requesting user's credentials, so confirm that integration detail rather than assuming the platform handles it automatically.
How do we know if an agent or MCP server is unregistered?
If you cannot produce a current list of every AI system with a live connection to production data, including owner, scope, and last review date, you have unregistered systems. This is worth auditing before adding new agents.
Does bigger model size fix bad answers?
No. Anthropic's testing and OpenAI's production experience point the same direction: structure, meaning clear semantics, defined context, and durable memory of corrections, is what separated a reliable answer from a confident guess in both cases, not model scale.
