Understanding Metadata Systems
This article is part of a series exploring metadata systems from first principles—from why metadata exists to how modern metadata platforms are designed, implemented and operated.
Previous: Beyond Infrastructure: Seeing the Data & AI Platform as a System
Start With the Consumer, Not the Data
Modern applications in large enterprises generate enormous amounts of data. The current generation of modern data platforms can store enormous volumes of it, process it at scale, stream it to target systems in near real time, secure it, and make it available to downstream applications, analytics platforms and AI systems.
However, the availability of data does not automatically guarantee its usability or usefulness.
Suppose somebody asks for data from a customer_orders table that contains the following example record:
| customer_id | product_id | quantity | price |
|---|---|---|---|
| C10234 | P779 | 2 | 499.00 |
The above database record looks perfectly valid, and at first glance there does not seem to be much ambiguity. We have a customer, a product, a quantity and a price.
Before going deeper into the columns, or fields, to analyse the data further, as an architect I would first ask some more fundamental questions:
- Why do we need this data?
- What information are we actually looking for?
- How will this information be used?
- Who or what will consume it?
- In what form is it required?
- How current does it need to be?
For example, an analyst preparing a sales report, a finance team calculating revenue, an ML model using purchase behaviour as a feature, and an AI application answering questions about customer orders may all look at the same dataset with very different requirements.
The question then becomes:
Can these users or systems understand the data well enough to use it correctly for their respective requirements?
The Data Is There, But What Does It Mean?
The existence of a value does not always convey its meaning in a readily usable manner.
Let’s look at the price column:
price = 499.00
What do we understand by simply looking at the value 499.00?
What exactly does it represent?
- What is the currency?
- Is it the unit price or the total price for the two items?
- Does it include VAT?
- Is it before or after discount?
- Was this the price displayed to the customer or an internal accounting value?
The number itself cannot answer these questions. We need additional context.
For example:
Price
Meaning: Unit selling price
Currency: EUR
Discount: After discount
Tax: Before VAT
We haven’t changed the original value. 499.00 remains 499.00. What has changed is our ability to interpret it.

This is where metadata begins to become important.
So What Exactly Is Metadata?
Metadata Is More Than ‘Data About Data’
If you ask this question to people, you will most probably get the answer, “Metadata is data about data.” This is the common definition of metadata used widely in the industry.
There is nothing fundamentally wrong with that definition. It is useful as a starting point, but from an architecture perspective, it does not tell us enough about why metadata exists.
A more useful working definition would be:
Metadata is structured knowledge and context that enables humans and systems to identify, understand, discover, govern, operate, trust and relate data assets.
Going back to our customer_orders example:
- Knowing that
priceisDECIMAL(10,2)is metadata. - Knowing that it represents the unit selling price in EUR after discount and before VAT is metadata.
- Knowing that
customer_idcontains personally identifiable information (PII) is metadata. - Knowing that the table is refreshed every 15 minutes is metadata.
- Knowing that the Order Management team is responsible for it is metadata.
- Knowing where the data originated and which dashboards consume it is also metadata.
The underlying order rows have not changed. What has changed is how much we know about them.
Different Questions Require Different Metadata
Rather than starting with a taxonomy and trying to memorize different metadata categories, let’s begin with the questions a consumer needs answered.
| Question | Metadata |
|---|---|
| What fields exist and what are their types? | Technical metadata |
What does price actually mean? | Business metadata |
| Where did this dataset come from? | Lineage metadata |
| Who is responsible for it? | Ownership metadata |
| Does it contain sensitive information? | Governance / classification metadata |
| Is the data complete and reliable? | Quality metadata |
| When was it last refreshed? | Operational metadata |
The categories are useful. However, the questions come first.
We should not collect metadata simply because a product allows us to collect it. We should understand what questions need to be answered and what knowledge is required to answer them.
Where Does Metadata Come From?
There would be no single place that necessarily knows everything about customer_orders.
- The database or data warehouse might know the column name, data type, nullability and table structure.
- The application or transformation pipeline might know how the data was calculated.
- The Commerce team might understand what an order represents in the business.
- Finance might define precisely what
pricemeans. - Security or governance teams might determine whether
customer_idis sensitive. - A data-quality system might know whether recent quality checks have passed.
- An orchestration or runtime system might know when the dataset was last successfully refreshed.
An organizational directory might tell us which team or person is responsible for the dataset.
This leads us to an important observation:
The authoritative knowledge about a data asset is itself distributed.

Who Owns the Truth About the Metadata?
‘Source of Truth’ May Be the Wrong Question
We frequently ask:
What is the authoritative source for each kind of knowledge about this data asset?
That might be a reasonable question for the data itself. However, for metadata, it can be too broad.
A better question would be:
What is the authoritative source for each kind of knowledge about this data asset?
| Kind of knowledge | Possible authoritative source |
|---|---|
| Physical schema | Database / warehouse |
| Business definition | Business domain / domain owner |
| Classification | Security / governance |
| Quality status | Data-quality system |
| Runtime freshness | Pipeline / orchestrator |
| Ownership | Organizational process / directory |
Authority may therefore exist at the level of a particular metadata attribute or category, rather than at the level of the entire asset.
This also helps us distinguish between three concepts that can easily be mixed together:
- Ownership is about accountability.
- Authority is about who or what can reliably determine a particular fact.
- Permission is about who is technically allowed to read or modify something.
- A team may own a dataset without being authoritative for its security classification.
- Similarly, a platform administrator may have permission to modify metadata without being the authority on its business meaning.
Metadata Has Producers Too
If we recognise that metadata is information in its own right, somebody or something must produce it.
This leads to another question:
Who actually produces the metadata?
How does that metadata get created?
Is all metadata manually documented by people, or can systems produce and derive metadata too?
Some metadata is produced automatically. A warehouse can expose schemas and column types, a pipeline can expose execution information, and a quality system can produce test results.
Some metadata is supplied by people. A domain expert may describe what a field means, a governance team may establish a classification, and an owner may document expectations around the use of a dataset.
Some metadata may also be derived from other information—for example, relationships or lineage inferred from pipelines and queries.
Metadata is not simply documentation that somebody writes after a data system has been built. It is produced continuously by both systems and people.
Metadata Changes Too
We have already discussed that metadata is produced either automatically by systems or manually by people.
Now another question arises: Once metadata has been created, does it always remain the same?
If somebody or some system changes metadata at its source, how would the consumers know about that change? If those changes are not propagated properly, consumers might continue to use old or stale metadata because the knowledge about the data is distributed across different systems and people.
Let’s take the example of the same customer_orders dataset again.
Imagine that the actual business records remain unchanged. However, somebody or some system changes its metadata—that is, the information or context we maintain about the data has changed.
For example:
| Type of change | Actual change made |
|---|---|
| Ownership change | Ownership of the customer_orders dataset changes |
| Business change | The definition of price is clarified or modified |
| Security / governance change | The classification of customer_id changes |
| Schema change | A new column or field is introduced |
| Operational change | The expected refresh frequency changes from 15 minutes to 10 minutes |
| Dependency change | A new upstream or downstream dependency is added |
| Quality change | Quality expectations or quality rules change |
| Lifecycle change | The dataset is eventually deprecated |
This leads to another question:
If the underlying data has not necessarily changed, but our knowledge about that data has changed, shouldn’t the metadata change as well?
There is another interesting situation to consider.
Suppose customer_orders is eventually retired and the physical table is deleted. Does that mean all the metadata about the dataset immediately becomes useless?
Not necessarily.
Some of that metadata may still be valuable because someone might need to understand its historical lineage, interpret an old dashboard, support an audit, or understand why another system previously depended on it.
This means that the lifecycle of metadata is related to the lifecycle of the data itself, but the two are not necessarily the same.

Who Consumes Metadata
Who actually needs and uses this metadata?
Metadata is not only for humans.
Let’s see how different types of consumers need metadata and use it for different purposes.
Human Consumers
Human consumers might include:
- Data engineers
- Architects
- Developers
- Security and governance teams
- Platform and SRE teams
- Business users
They use metadata to answer questions such as:
- What does this dataset mean?
- Can I trust it?
- Who owns it?
- Is it current?
- Where did it come from?
- Is it sensitive?
System Consumers: Systems also consume metadata for different purposes.
- A pipeline might need schema information.
- A governance system might use classification metadata.
- A data-quality system might use ownership or schema information.
- Discovery and search systems might use descriptions and relationships.
- Automation can use operational metadata.
- AI applications and agents increasingly need context to determine what data means and whether it is appropriate to use.
Therefore, metadata is not only information that helps people understand data. It also provides context that systems can consume programmatically to make decisions and perform different operations.
As data platforms become increasingly automated and AI-driven, this context needs to become more explicit, structured and machine-consumable.
The Metadata Already Exists — So What’s Missing?
The schema may already exist in the warehouse, business definitions in documentation or with domain teams, classifications in governance systems, quality information in data-quality tools, and operational information in pipelines and monitoring systems.
In a small environment with one application, one database and a small team, this distributed knowledge may be perfectly manageable. People know where the data is, what it means, and whom to ask when they need more information.
This leads to an important observation:
Distributed metadata is not inherently a problem.
But what happens when the environment grows to hundreds of applications, thousands of datasets, pipelines, streaming topics, dashboards and ML models distributed across different teams and platforms?
At this scale, the nature of the problem begins to change and raises some new questions:
- Where should a consumer start looking?
- How do they know what data exists?
- How do they determine which dataset is appropriate for their requirement?
We started this discussion by asking why data alone is not enough.
We found that data needs context, that this context comes from different systems and people, that different sources may be authoritative for different kinds of metadata, and that metadata itself has producers, consumers and a lifecycle.
But this leaves us with another architectural question:
If useful metadata already exists across the systems and teams that produce and manage data, why do we need another platform to bring that knowledge together?
Continue the Series
Previous: Beyond Infrastructure: Seeing the Data & AI Platform as a System
Next: Why a Metadata Platform? When Distributed Knowledge Stops Scaling — Coming soon




Leave a comment