Why trusting a metric requires much more than verifying that the number was calculated correctly.

There is a seemingly simple question that can trigger a surprisingly long discussion:

How many active cards do we have?

It sounds like we should be able to query a database and come back with a number.

The problem starts when we ask:

What exactly do we mean by an active card?

A card that hasn't been canceled?

One that is currently enabled for transactions?

One that has had at least one transaction in the last 30 days?

In the last 90?

Does an issued card that has never been used count as active?

Depending on the definition, we can get completely different numbers.

And the interesting part is that all of them can be calculated correctly.

The problem isn't necessarily the data.

It's what we believe the data means.

Two different numbers don't always mean that one of them is wrong. Sometimes they mean we never agreed on what we were measuring.

In practice, part of this problem can be addressed by making the definition explicit.

Documenting it helps. Versioning it does too, because metrics change, and the way an organization understands its business can change as well.

It doesn't eliminate every discussion.

But it allows us to discuss the definition instead of arguing about which of two queries is "the right one."


Data can be calculated correctly and still be wrong for the decision

When we talk about data quality, we tend to think about errors first.

A source that stopped updating.

A status that didn't change when it should have.

Duplicate records.

Missing data.

An incorrect transformation.

These are all real problems. I've encountered many of them over the years.

But there is another, quieter category.

The pipeline works.

The query is correct.

The dashboard is up to date.

The number is mathematically accurate.

And yet two people interpret it differently.

Before trusting a metric, then, it isn't enough to ask:

Is this number correct?

We also need to ask:

What exactly does this number mean?

For decades, research on data quality has argued that evaluating quality based only on accuracy is not enough. Data also needs to be appropriate for the context in which it will be used and understandable to the people consuming it.

That idea remains surprisingly relevant.

Because a technically perfect metric with an ambiguous definition can lead us to a perfectly wrong decision.


The definition is part of the metric

More than once, I've been in a meeting where two teams showed different numbers for what was supposedly the same metric.

In one of those situations, both teams were convinced their number was correct.

The first reaction was to look for the error: review queries, sources, and transformations.

And yes, sometimes we found errors.

But other times, we found something more interesting.

The sources weren't the same.

The filters were different.

One included certain statuses while the other didn't.

One used the creation date while the other used the approval date.

One considered a calendar period while the other looked at the last 30 days.

The error wasn't necessarily in the calculation.

We were using the same name to describe different things.

Let's go back to "active cards."

If Finance considers an active card to be any card that is enabled, while Product considers an active card to be one that has had a transaction in the last 30 days, they will probably get two different numbers.

And both may be right.

The wording of a metric, then, isn't just a documentation detail.

It's part of the metric itself.

A good definition should make it possible to understand what we are measuring, what we are excluding, which source we are using, how the metric is calculated, and what time period it applies to.

Without that, a dashboard can create an illusion of precision.

We have a number with two decimal places.

What we may not have is a shared meaning.


Time changes the meaning of data

In the previous article, I wrote about the importance of preserving the temporal dimension of our systems.

It appears again here, but from a different perspective.

Time is also part of a metric's definition.

Suppose someone asks:

How many customers do we have?

We could answer how many exist today.

How many were active at the end of last month.

How many made a purchase in the last 30 days.

How many had any activity during the last year.

How many have ever been customers.

Those are different questions hidden behind an apparently simple sentence.

This becomes even more important when we compare periods.

If we change the definition of a metric, switch its source, or incorporate information that wasn't previously available, we can break historical comparability without realizing it.

The chart still has a line.

But part of that line may now mean something different from the rest.

A metric without a clear time definition can be correct today and incomparable with itself tomorrow.

When I review a number, I try to understand not only how it was calculated.

But also when.


Averages tell incomplete stories too

There is another area where I try to be particularly careful: means, averages, and aggregations.

Suppose someone says:

"On average, it takes a customer 48 hours to complete the process."

The number may be completely correct.

But we still know very little.

Do most customers take roughly 48 hours?

Or does half of them finish in ten minutes while a small group takes several weeks?

Has the composition of customers changed?

Are there segments with completely different behaviors?

Are we comparing equivalent populations?

An average summarizes.

And precisely because it summarizes, it flattens reality.

It can hide whether observations are relatively concentrated around that value or whether we have a highly skewed distribution, where a small number of extreme cases — a long tail — pushes the average upward.

That doesn't make the average a bad metric. It makes it a metric that needs context.

Sometimes looking at the median, a few percentiles, or simply the full distribution tells a very different story.

Someone once told me something that stuck with me:

The same numbers can support different interpretations.

I think that's true.

Data constrains the stories we can tell, but it rarely eliminates the need for interpretation entirely.

When I see a conclusion built around an average or an aggregation, I want to understand the source, the distribution, the population, the time period, and how the metric was constructed.

Not because I distrust the data.

But because I want to understand what conclusion it can actually support.


Missing data and incorrect data don't carry the same risk

If I had to choose between not having a piece of data and having incorrect data, I'd rather know that I don't have it.

When information is missing, uncertainty is visible.

We can say:

"We don't know."

We can look for another source, build an approximation, postpone a decision, or make it while explicitly acknowledging what information is missing.

Incorrect data has a different characteristic.

It can hide uncertainty.

It allows us to walk into a meeting, look at a dashboard, and make a decision convinced that we have evidence.

That's what makes it dangerous.

When data is missing, we know that we don't know. When data is wrong, we can be convinced that we know something that was never true.

Ambiguous data can have a similar effect.

The number itself may be correct.

But if different people understand different things when they read it, the organization isn't making decisions based on a shared reality either.


Before trusting a metric

Over time, I've developed a few fairly simple questions that I ask when an important number comes up. They aren't a sophisticated framework, nor are they meant to replace a formal data governance process. They're simply a way to understand what I'm looking at before building a conclusion around it.

  1. What exactly are we measuring?
  2. How is it calculated?
  3. What is the source?
  4. Which population does it include, and which does it exclude?
  5. What is the time dimension?
  6. Are we comparing comparable things?
  7. What's underneath this number?

That last question becomes particularly important when looking at averages, aggregations, or significant changes.

We don't always need to answer all of them with the same level of depth. An operational metric we monitor every day may require a different level of scrutiny than a number we are going to use to make a major investment decision.

The level of confidence we require should be proportional to the cost of being wrong.


When two teams have different numbers, the problem isn't always Data

When two teams arrive with different numbers, it's tempting to send the problem to the Data team and ask them to determine which one is "correct."

Sometimes that's exactly what needs to happen.

But if the problem is that the organization never defined what "active customer" means, Data can't solve that technically.

It can show the differences.

It can quantify them.

It can document them.

It can build a single metric once an agreed definition exists.

But someone still needs to make a business decision:

What do we mean when we use this term?

This is where something I consider fundamental comes into play.

Data governance isn't just about permissions, catalogs, and tools.

It's also about building a shared language.

So that when two people say "conversion," "active customer," "sale," "delinquency," or "profitability," they know whether they're actually talking about the same thing.


With AI, the problem changes scale

Until relatively recently, many of these discussions ended at a dashboard.

A person looked at a number.

Interpreted it.

Maybe asked a question.

Maybe noticed that something didn't add up.

Now we're starting to build something different.

Agents that query information.

Systems that interpret results.

Models that recommend actions.

Automations that execute them directly.

And that raises the bar.

Suppose we tell an agent:

"Contact all inactive customers."

It sounds like a perfectly clear instruction.

Until we ask the same question we asked at the beginning:

What does inactive mean?

Or imagine another instruction:

"Prioritize high-risk customers."

High risk according to which model?

Which version?

Using data from when?

What population was the model built on?

Is more recent information available?

And what exactly does "prioritize" mean?

When an analyst encounters ambiguity, they can raise their hand and ask.

We shouldn't assume an agent will do the same.

It may take an incomplete definition, infer what's missing, and continue. And when that system also has the ability to act, ambiguity stops being just an analytical problem.

It becomes an operational one.

A dashboard with ambiguous data can trigger a discussion. An agent with ambiguous data can execute a decision.

This also changes how I think about AI governance.

It's not enough to know which models we're using, how much they're being used, or what information they can access. If we increasingly want to delegate tasks and decisions, we also need to understand what data they are reasoning over, what definitions they receive, and what actions they are authorized to execute.

Agent observability shouldn't end with tokens, latency, or number of executions.

We should also be able to reconstruct the entire path that led to an action. Something similar to the lineage we've spent years trying to build around data, but now applied to decisions:

Decision lineage.

  • What information the system received.
  • How it interpreted that information based on the definitions available.
  • What conclusion or inference it reached.
  • What action it actually executed in the business.

If we want to delegate decisions, we also need to be able to reconstruct how those decisions were made.

This isn't just a theoretical concern. Current AI risk management frameworks increasingly emphasize reliability, transparency, measurement, and governance throughout the lifecycle of these systems.

AI doesn't eliminate our data problems.

It can turn them into actions.


Maybe we need to start with language

For a long time, we thought improving our ability to work with data meant building better pipelines, warehouses, models, and dashboards.

All of that is still necessary.

But I'm increasingly convinced that an important part of the problem is far less technical.

We need to agree on what things mean.

What is a sale?

What is a customer?

What does active mean?

When does a conversion start and end?

Which date do we use?

What exactly does each number we put in front of a decision-maker represent?

In the previous article, I argued that data infrastructure begins long before the data warehouse.

I think there's a complementary idea here:

Data quality also starts before the data. It starts with the definition.

We can have the perfect architecture.

We can have error-free pipelines.

We can update information in real time.

We can have extraordinarily sophisticated models.

But if two people look at the same term and understand different things, we still have a problem.

And as we delegate more decisions to agents and automated systems, that precision will stop being merely a good analytical practice.

It will become a requirement for automating with confidence.

Because in the end, it's not enough for a number to be correct.

We also need to know exactly what it means.


Further reading

  • Beyond Accuracy: What Data Quality Means to Data Consumers — Richard Wang and Diane Strong. A classic work on why data quality goes far beyond whether a number is simply correct.

  • NIST AI Risk Management Framework. A framework for thinking about trustworthiness and risk as we incorporate artificial intelligence into processes and decisions.

  • NIST Generative AI Profile. An extension of the framework focused specifically on the risks and challenges of generative AI systems.