[Learning Center](/learning-center)[Data Catalog](/learning-center/topic/data-catalog)

# What does AI need from a data catalog?

Your catalog probably works fine for the humans who use it. An analyst searches a term, finds a table, checks who owns it, maybe reads a description someone wrote two years ago, and moves on.

Slow, occasionally frustrating, but survivable, because the human fills every gap the catalog leaves. They know which "customer\_id" is the real one. They know the finance dashboard broke last Tuesday and to distrust anything downstream of it. They know who to ping when something looks wrong.

Now put an agent in that seat. Unlike a human analyst, an agent cannot rely on informal institutional knowledge unless that knowledge is made available through governed tools, documentation, or memory. A catalog built to help a person eventually figure things out was not built to give a machine what it needs to get things right the first time.

That gap is why so many AI-on-data projects stall after the demo. [Gartner's Rita Sallam](https://www.gartner.com/en/newsroom/press-releases/2026-05-11-gartner-says-lack-of-semantics-causes-inaccurate-artificial-intelligence-agents-and-wasted-spending) has said that without context, AI agents cannot operate accurately and are far more likely to hallucinate, introduce bias, and produce unreliable results.

Gartner has separately projected accuracy gains of up to 80% and cost reductions of up to 60% by 2027 for organizations that invest in AI-ready data foundations, a gap that runs the other direction for organizations that don't. Not weak models, but data infrastructure that was never built for a machine reader. The fix is giving the agent the same six things a competent human colleague picks up automatically: structured metadata, shared semantics, quality signals, lineage, permission-aware access, and memory.

**Key Takeaways**

*   An AI agent has no colleague to ask when a metric looks ambiguous. It answers anyway, so the catalog has to resolve ambiguity before the question reaches the model.
*   Structured, connected metadata beats more retrieval. Giving an agent everything without structure barely moves accuracy, while governed context and semantics move accuracy by multiples.
*   Quality and lineage signals need to sit beside the catalog entry, not in a Slack channel only engineers monitor, or agents end up quoting numbers that already failed a freshness test.
*   Permission-aware retrieval has to extend into what the agent retrieves and reasons over, not just what it's allowed to display at the end.
*   Memory stops every new agent session from rediscovering the same edge case a human analyst solved two years ago.
*   Auditing against these six primitives isn't a quick checklist item: some, like permission-aware retrieval, typically require re-architecting access control rather than adding a catalog field, so scope the audit before you scope the pilot.

## Article Contents

*   [Metadata an agent can parse](#metadata-an-agent-can-parse)
*   [Semantics: what the fields mean](#semantics-what-the-fields-mean)
*   [Quality signals the agent can act on before it answers](#quality-signals-the-agent-can-act-on-before-it-answers)
*   [Lineage: Does the agent know where the number came from?](#lineage-does-the-agent-know-where-the-number-came-from)
*   [Access: Agents inherit permissions, they don't get their own rules](#access-agents-inherit-permissions-they-don-t-get-their-own-rules)
*   [Memory: So the agent doesn't relearn the same lesson every session](#memory-so-the-agent-doesn-t-relearn-the-same-lesson-every-session)
*   [Why an agent's mistake costs more than a bad dashboard](#why-an-agent-s-mistake-costs-more-than-a-bad-dashboard)
*   [When the failure looks like a model problem but isn't](#when-the-failure-looks-like-a-model-problem-but-isn-t)
*   [What this means for platform leaders](#what-this-means-for-platform-leaders)
*   [Frequently asked questions](#frequently-asked-questions)

## Metadata an agent can parse

A data catalog tells you what data exists: table names, column types, who owns what, when it was last updated. That's necessary, and most catalogs already do it. But metadata built for a search bar and metadata built for an agent's reasoning loop are not the same artifact. A human sees "cust\_typ\_flag" and squints, then asks a colleague. An agent sees it and either guesses or hallucinates a join. Structured, machine-readable metadata, the kind covered in [what a data catalog is and how it works](https://www.getcollate.io/learning-center/data-catalog), is the floor. It is not the ceiling, and treating it as the whole solution is where most AI-readiness projects go wrong.

## Semantics: what the fields mean

Metadata tells you a column exists. It doesn't tell you that "active\_user" means something different in the billing system than it does in the product analytics warehouse. That gap is exactly where agents fail silently, producing a confident, wrong answer instead of an error message. Text-to-SQL benchmarks make the size of the gap concrete: on Spider 2.0, frontier models like GPT-4o and Claude Sonnet 4.5 solve roughly 10% of realistic enterprise queries when working from raw schema metadata alone, while adding governed semantic context (business definitions, approved joins, canonical metrics) lifts accuracy into the high 70s, as detailed in [how a context platform/layer differs from a semantic layer](https://www.getcollate.io/learning-center/how-is-a-context-layer-different-from-a-semantic-layer). The distinction matters operationally. A semantic layer defines the metrics and business logic. A context platform carries that meaning, plus quality and lineage signals and memory, to wherever an agent is reasoning, so the agent doesn't have to rediscover the definition of "revenue" every time it runs a query.

## Quality signals the agent can act on before it answers

A dashboard that's silently wrong sits there until someone notices. An agent that's silently wrong acts on the wrong number immediately, and it might write that number into a report, a customer email, or a downstream pipeline before anyone checks. That's the sharpest difference between AI stakes and BI stakes: a human in the loop used to be an implicit quality gate, and agents remove that gate unless you build one back in. [Data quality signals](https://www.getcollate.io/learning-center/data-quality), freshness checks, anomaly flags, and test pass and fail status need to travel with the data, visible to the agent at query time, not buried in a separate monitoring tool the agent never queries.

## Lineage: Does the agent know where the number came from?

If an agent cites a number, someone downstream will eventually ask where it came from and whether the source is trustworthy. Without [data lineage](https://www.getcollate.io/learning-center/data-lineage), the agent has no way to trace a suspicious result back to its origin, verify it hasn't been quietly deprecated, or know that the table it just queried was flagged broken an hour ago. Lineage also makes an agent's answer auditable after the fact, which matters more for a system that acts autonomously than for one where a human reviewed the query before running it.

## Access: Agents inherit permissions, they don't get their own rules

An agent querying on a user's behalf has to respect that user's actual permissions. Achieving this requires row-level security, column masking, classification-aware retrieval, and runtime enforcement, not merely a catalog API. This is where a lot of "catalog is AI-ready" claims fall apart under inspection. The real test isn't whether the catalog has an API. It's whether that API enforces the same governance rules a human would hit in the UI. Of the six primitives, this is often the most expensive to retrofit: it usually means extending access control into the retrieval path itself, not just adding a field to a catalog record. [Whether a catalog is enterprise-ready for AI](https://www.getcollate.io/learning-center/is-openmetadata-enterprise-ready) comes down largely to whether permission-aware retrieval exists and is enforced consistently across the catalog, retrieval system, and underlying data sources.

## Memory: So the agent doesn't relearn the same lesson every session

[Anthropic's own reporting on internal agent evaluations](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) says that accuracy started at 21% before structured context was introduced, then consistently exceeded 95% after it was in place. Separately, giving that same context-layer-equipped agent broader raw file and grep access on top of the structured context moved accuracy by less than a single point. The bottleneck wasn't retrieval breadth. It was structure. Part of that structure is memory: when a human corrects an agent's mistake once, does that correction persist for every future query, or does the next session start from zero? An [enterprise context platform](https://www.getcollate.io/learning-center/enterprise-context-layer) is built to carry corrections, definitions, and prior decisions forward, so the fix one analyst makes today is context the next agent run inherits, instead of a lesson that evaporates when the session ends.

## Why an agent's mistake costs more than a bad dashboard

A wrong number on a dashboard sits there until a human notices it, questions it, or ignores it. An agent's wrong number gets acted on: sent in an email, written into a report, used to make a decision, propagated to the next agent in the chain. There is no human checkpoint unless you build one in deliberately. That's the real reason "just add AI to the catalog" fails as a plan. It treats agent risk as the same shape as dashboard risk, when the blast radius and the response time are both different.

## When the failure looks like a model problem but isn't

OpenAI's own case study on building [OpenAI's Kepler agent](https://open-metadata.org/case-study/openai), a self-service data agent for internal analytics, is instructive here because the team didn't set out to fix a metadata problem. They set out to fix an AI accuracy problem, and the root cause turned out to be context. In one documented instance, Kepler answered a user-count question with 5,062,338 when the real figure was closer to 800 million, roughly 160 times smaller than the true value, not because the model reasoned badly but because it lacked the structured context (which table was canonical, what the metric actually meant, what had changed upstream) to know its answer was implausible. OpenAI's own framing describes six layers of context an agent draws on to get this right: schema, semantics, quality, lineage, usage patterns, and organizational knowledge. Fix the model and this error recurs. Fix the context layer and it doesn't.

## What this means for platform leaders

If you're the one who has to decide whether your catalog can support production agents, the useful question isn't "do we have a catalog." It's "does our catalog give an agent metadata, semantics, quality, lineage, permission-aware access, and memory, or does it give a person a search box."

[AI-ready data](https://www.getcollate.io/learning-center/ai-ready-data) rests on the pillars of metadata, meaning/semantics, and memory, carried through a governed context platform, and a catalog that only handles discovery for humans is missing most of these by design, not by oversight. Human analysts already rely on [data discovery](https://www.getcollate.io/learning-center/data-discovery) workflows built for browsing and search; agents need the same information, but retrievable programmatically and structured enough to act on without a person double-checking the result.

Before you greenlight the next agent pilot, audit your stack against all six primitives rather than the feature list a vendor hands you, and size the audit honestly. Metadata and semantics gaps are often closeable in weeks. Permission-aware retrieval and lineage coverage across legacy pipelines are frequently a much longer project, since they can require re-architecting how access control and provenance are tracked rather than adding fields to an existing catalog. Know which category each gap falls into before you commit a timeline to stakeholders.

The [metadata foundational for agentic analytics](https://www.getcollate.io/learning-center/metadata-for-agentic-analytics) argument holds regardless of which agent framework or model you standardize on: passive catalogs answer "what exists." Active context platforms answer "what can I trust, and what happens if I act on it." Getting this right also [supports better decisions](https://www.getcollate.io/learning-center/how-does-metadata-management-support-better-decision-making) further downstream, since every dashboard and report an agent feeds inherits whatever context the agent had, or didn't have, when it ran the query.

## Frequently asked questions

### Is a data catalog the same as an AI context layer?

No. A catalog inventories what data exists and who owns it. A context platform for AI combines that information with semantics, quality, lineage, memory, and machine-readable access controls. It's built to be queried by machines in real time, not just browsed by people.

### Can we just add a semantic layer and call it done?

A semantic layer solves the meaning problem, business definitions and canonical metrics, but not quality, lineage, access enforcement, or memory. It's one of six primitives, not a substitute for the rest.

### Why did our agent pilot work in the demo but fail in production?

Demos usually run against a curated, small dataset where semantics and quality are effectively hand-checked in advance. Production exposes every gap the demo hid: stale tables, undocumented joins, permission edge cases, and no memory of prior corrections.

### Does this mean we need to rebuild our whole catalog?

Not necessarily. It means auditing what you have against the six primitives and closing the gaps that matter most for your first production use cases, usually quality signals and access enforcement, before scaling to more agents. Budget more time for access enforcement specifically; it tends to be the primitive most likely to require rework beyond the catalog itself.

### How does permission-aware retrieval work for AI agents specifically?

In a properly designed system, the agent's queries and retrieved context are filtered through the same row-level security, column-masking, and classification rules that apply to the user it is acting for.

[

## Fashion Retailer Mango’s Data Journey with Collate

Read the case study

![](data:image/svg+xml,%3csvg%20xmlns=%27http://www.w3.org/2000/svg%27%20version=%271.1%27%20width=%27708%27%20height=%27470%27/%3e)![Mango](/_next/image?url=%2Fimages%2Flearning-center%2Fmango-lc.webp&w=1920&q=75)



](/resources/ebook/mango-case-study)

Sign up to receive updates for Collate services, events, and products.

## Share this article

[![](data:image/svg+xml,%3csvg%20xmlns=%27http://www.w3.org/2000/svg%27%20version=%271.1%27%20width=%2724%27%20height=%2724%27/%3e)![Share on Twitter](data:image/gif;base64,R0lGODlhAQABAIAAAAAAAP///yH5BAEAAAAALAAAAAABAAEAAAIBRAA7)

![Share on Twitter](/_next/image?url=%2Fimages%2Ffooter%2Ftwitter-x-v1.svg&w=48&q=75)

](https://twitter.com/intent/post?url=https%3A%2F%2Fgetcollate.io%2Flearning-center%2Fwhat-does-ai-need-from-a-data-catalog%3Fref%3Dtwitter-share)[![](data:image/svg+xml,%3csvg%20xmlns=%27http://www.w3.org/2000/svg%27%20version=%271.1%27%20width=%2724%27%20height=%2724%27/%3e)![Share on LinkedIn](data:image/gif;base64,R0lGODlhAQABAIAAAAAAAP///yH5BAEAAAAALAAAAAABAAEAAAIBRAA7)

![Share on LinkedIn](/_next/image?url=%2Fimages%2Ffooter%2Flinkedin-v1.svg&w=48&q=75)

](https://www.linkedin.com/feed/?linkOrigin=LI_BADGE&shareActive=true&shareUrl=https%3A%2F%2Fgetcollate.io%2Flearning-center%2Fwhat-does-ai-need-from-a-data-catalog)

Ready for trusted intelligence?

See how Collate helps teams work smarter with trusted data

[Get Started](/welcome)[Contact Us](/contact-sales)