Skip to content
Start a conversation
Governance

The data catalogue nobody opens

Almost every organisation we work with has bought a data catalogue. Rather fewer have a data catalogue that anybody opens. The gap between those two states is one of the more expensive quiet failures in enterprise data, because the licence renews whether or not the tool is used, and because the failure is easy to mistake for a tooling problem.

It is usually not a tooling problem. The catalogues on the market are broadly competent. The problem is that catalogues are deployed to answer a question users are not asking, in a place users do not go, with content users cannot trust.

The question nobody is asking

Catalogue projects are typically justified with a discovery narrative. Analysts waste time hunting for data. A searchable inventory will let them find it. Productivity improves.

Watch how analysts actually work and the narrative wobbles. Experienced analysts rarely have a discovery problem within their own domain. They know where the revenue tables are. They have known for two years. What they have is a trust problem and a change problem.

The real questions sound like this. Is this table still being loaded. Which of these three revenue tables is the one finance uses. Did the definition of this field change last month, because my number moved and nothing in my code did. If I build on this, will someone deprecate it without telling me.

A catalogue populated by crawling schemas answers none of those. It tells you a table exists, which the analyst already knew.

The trust collapse

Catalogues fail fast and then stay failed, and the mechanism is worth understanding because it is nearly always the same.

Launch happens with partial coverage, which is sensible. Descriptions are backfilled in a push, often by people who are not the domain experts, often under time pressure. Some are good. Many are restatements of the column name. A field called cust_status_cd acquires the description “customer status code”.

An analyst visits, searches three terms, and gets one useful answer and two restatements. They form a judgement: this thing does not know more than I do. That judgement is durable. They will not re-test it in six months when coverage has improved, because nobody re-litigates a settled conclusion about an internal tool.

This is why partial launches are more dangerous for catalogues than for most software. The cost of a weak first impression is not a delayed adoption curve. It is a permanently capped one.

Metadata that is generated beats metadata that is written

The most-used fields in every catalogue we have measured are the ones no human maintains.

  • Freshness. When was this last loaded successfully. Answers the single most common question directly.
  • Usage. How many distinct people queried this in the last thirty days, and who. Resolves the three-revenue-tables problem instantly, because the one finance uses is the one finance queries.
  • Downstream dependency count. A proxy for consequence. Tables with forty dependents are load-bearing whether or not anyone documented them as such.
  • Change history. Schema and definition changes with dates. Turns “my number moved” from an investigation into a lookup.
  • Test status. What is asserted about this asset and whether those assertions currently pass.

None of these degrade. None require a campaign. All of them answer questions people genuinely have. If we were sequencing a catalogue programme from scratch, we would ship these five and no prose descriptions at all, then add prose only where the generated signals demonstrably fail to explain something.

Go to where the work happens

The second structural problem is location. A catalogue is a destination. Analysts work in a query editor and a BI tool. Asking them to leave, authenticate somewhere else, search, and come back is a real cost, paid every time, against a benefit that is uncertain until after they have paid it.

Every meaningful adoption jump we have seen came from surfacing catalogue content inside an existing surface rather than from improving the catalogue itself. A hover card in the BI tool showing freshness and owner. A comment block in the query editor. A bot that answers a definition question in the channel where it was asked.

The catalogue becomes infrastructure rather than a product. That is a demotion in visibility and a promotion in usefulness, and it is usually the right trade.

Coverage is the wrong metric

Catalogue programmes are almost universally measured on percentage of assets documented. It is an easy number to produce and it drives exactly the wrong behaviour, because it rewards volume of description over usefulness of description, and it treats a table queried twice a year identically to one that feeds the board pack.

We would measure three things instead. Weekly active users as a share of the analyst population, which tells you whether the thing is alive. Question resolution, meaning the proportion of definition questions in support channels that get answered with a catalogue link rather than a human explanation. And coverage weighted by usage, so documenting the top fifty assets counts for more than documenting five hundred dormant ones.

Those numbers are harder to move and they mean something. Unweighted coverage can reach ninety percent while the tool sits unopened, and we have seen precisely that on more than one engagement.

If it has already failed

Reviving a catalogue that has lost its audience is harder than launching one, because you are arguing against a formed opinion. Relaunching with a fanfare rarely works. What has worked for us is quieter.

Pick the single most-queried domain. Get the generated signals right for it. Embed those signals in the tool that domain already uses. Say nothing. Let people encounter it in the flow of work and reach their own conclusion. Then expand.

Adoption recovers on the strength of the answer, not the strength of the announcement.

The one description worth writing by hand

We have argued that generated metadata beats written metadata, which invites the question of whether prose has any place at all. It does, in exactly one form: the caveat.

The genuinely valuable human contribution to a catalogue is not describing what a field contains. It is recording the thing you only learn by getting burned. That this table excludes cancelled orders before 2022. That this timestamp is in local time despite the column name. That this field was repurposed in March and older rows mean something different.

These cannot be generated and they are worth a great deal, because they are precisely the knowledge that currently lives in three people’s heads and leaves when they do. If you run one campaign, do not run a description drive. Ask experienced analysts what they wish they had known before their first mistake with each major asset, and record those answers.

Ready to turn complexity into your next advantage?

Book a discovery call