There is a situation I come across quite often.

An organization decides that a particular product is strategic. It sets objectives, allocates budget, builds a team, and starts executing.

Then the data comes in.

And it tells a different story.

Maybe another product converts better. Maybe customers don't value a feature as much as we expected. Maybe a change that seemed obviously beneficial actually makes the outcome worse.

Then something interesting happens: instead of revisiting the decision, we start looking for another metric.

We change the time period. We look at a different segment. We find an explanation.

Until we find a number that confirms what we had already decided.

The company has data. It has dashboards. It probably has a Data team.

But that doesn't mean it is data-driven.

There is a huge difference between using data to make a decision and using data to justify a decision we have already made.

To me, being data-driven has much less to do with how much data we have and much more to do with the method we use to make decisions.

And, above all, with our ability to quickly discover when we are wrong.


The value of data isn't about always being right

For a long time, we have associated data-driven organizations with certain capabilities: data warehouses, dashboards, KPIs, analysts, data scientists, and reports distributed across the company.

All of these can be extremely valuable.

But they are infrastructure.

A company can have all of them and still make decisions exactly the same way it did before.

To me, the logic should look much more like the scientific method:

Hypothesis → Test → Measurement → Decision → Learning

We have an idea about what might happen. We look for a way to test it. We define what outcome we expect and what evidence would make us change our minds. We execute, measure, and learn.

There is one piece of data that I find particularly revealing.

In a paper on large-scale experimentation, Ron Kohavi and other researchers reported that only about one-third of the ideas tested at Microsoft improved the metrics they were intended to improve.

That's an uncomfortable number.

Because if even organizations with enormous technological and experimentation capabilities are frequently wrong about which ideas will work, perhaps the goal shouldn't be to become extraordinarily good at predicting.

Perhaps we should become extraordinarily good at quickly discovering when our prediction was wrong.

One of the greatest benefits of working with data isn't increasing our ability to be right. It's reducing the cost of being wrong.


An opinion can become a hypothesis

This comes up constantly in product decisions.

Suppose we believe that lowering the interest rate on a particular product will increase conversion enough to compensate for the reduction in margin.

Someone may be convinced it will work. Another person, with the same level of experience, may believe exactly the opposite.

We can argue about it for weeks.

Or we can turn the discussion into something that can be tested:

If we reduce the rate from X to Y, we expect conversion to increase by Z, while maintaining certain levels of profitability and risk.

The difference may seem small, but it completely changes the conversation.

We are no longer arguing about who is right.

We are designing a way to find out.

This also changes how I think about intuition.

Someone who has spent fifteen years working in a business has probably built a huge amount of knowledge that no dashboard can fully capture. They recognize patterns, understand exceptions, and may detect changes before they become clearly visible in a metric.

I would never dismiss that intuition.

I would use it to generate hypotheses.

If someone tells me, "Based on my experience, I think that if we do X, Y will happen," my response shouldn't necessarily be, "the data says otherwise."

It should be:

How can we test it?

Experience generates hypotheses. Evidence helps us evaluate them. And the outcome improves our experience for the next decision.


A metric should have a decision behind it

Another situation I often encounter starts like this:

"We need to measure this."

Before building the metric, I like to ask three questions:

What action are you going to take based on it?

What are you going to compare it against?

What does good or bad actually mean?

Suppose we build a dashboard that shows:

Conversion: 18%.

Is that good?

We don't know.

Maybe yesterday it was 12%. Maybe our target was 25%. Maybe one segment converts at 30% while another converts at 4%. Maybe conversion increased because we are accepting customers with significantly higher risk.

Or perhaps it has been at 18% for six months and nobody intends to do anything differently whether it moves to 15% or 21% tomorrow.

In that case, there is an uncomfortable question:

Why are we looking at this metric?

A metric without an associated decision risks becoming nothing more than information.

And accumulating information is not the same as building knowledge.


Improving a metric can make the business worse

This point is particularly important because experiments rarely have only one consequence.

Suppose we lower an interest rate and conversion improves.

Did we win?

We don't know yet.

We may have increased conversion at the expense of reducing margin too much. Or we may have brought in customers with a much higher level of risk.

That is why, in addition to a primary metric, we need to define what we are not willing to deteriorate while trying to improve it.

In experimentation, these are often called guardrail metrics. I prefer to think of them simply as protection metrics.

If our goal is to increase conversion, for example, profitability and risk could serve as protection metrics.

The rule is simple:

If the experiment improves the primary metric but breaks an important protection metric, the experiment didn't win.

This becomes especially important when outcomes appear at different times.

In lending, for example, we can observe changes in conversion almost immediately. But some consequences related to portfolio performance may take months to materialize.

A decision may look extraordinary during the first few weeks and considerably less attractive several months later.

Measuring quickly doesn't necessarily mean learning quickly if we are only measuring the part of the outcome that appears first.


Not everything can — or should — be experimented with

We also need to recognize the limits of this approach.

Not every decision can become an A/B test.

Some organizations or products have volumes so low that waiting for statistical significance may be unrealistic. In those cases, we need other tools: time-based comparisons, qualitative evidence, historical analysis, or simply the best information available.

Some decisions are also difficult to reverse.

An architectural decision, a key hire, or certain regulatory decisions do not always allow for a small experiment before moving forward.

In those situations, the question changes.

It is no longer:

"How do we run an A/B test?"

It becomes:

"What evidence can we gather before making a decision that will be expensive to reverse?"

And then there are long feedback cycles.

We can change something today and quickly observe an improvement in one metric while the cost appears three or six months later.

That is why I don't believe being data-driven means experimenting with absolutely everything.

That can become bureaucracy too.

It means choosing the right instrument for the decision in front of us.

The more reversible a decision is, the cheaper it should be to test and learn.

The less reversible it is, the more important it becomes to understand our assumptions before moving forward.


The problem isn't only cultural. It's also about incentives

There is something important we need to acknowledge.

Saying that an organization should be "willing to change its mind" sounds good.

But it isn't always that simple.

Imagine someone proposed an initiative, secured the budget, convinced leadership, assembled a team, and spent six months executing it.

Now the data shows that the hypothesis was wrong.

Accepting that result doesn't just mean saying, "I was wrong."

It may mean cancelling a project, reallocating people, explaining why a certain budget was spent, and perhaps putting future funding at risk.

There are real incentives to keep going.

So asking people to be objective isn't enough.

We need to design mechanisms that make objectivity more likely.

One simple mechanism is to write down before execution what the hypothesis is, what the primary metric is, which metrics we want to protect, and what result would make us stop or change direction.

Explicit review dates also help. Continuing an initiative shouldn't always be the default while cancelling it requires an extraordinary decision.

And there is probably something even more important:

An organization should be able to recognize stopping an initiative early, after evidence shows it isn't working, as a good outcome.

If cancellation is always interpreted as failure, teams will quickly learn to find arguments for continuing.

We can have all the data infrastructure in the world and still have that problem.


Learning also requires memory

Suppose we test a hypothesis and it doesn't work.

We spend little, measure correctly, and learn that under certain conditions our idea was wrong.

That can be an extraordinarily valuable outcome.

The waste happens when, six months later, another team proposes exactly the same thing because nobody knows it has already been tested.

That's why I believe a data-driven organization should preserve more than tables and dashboards.

It should be able to answer:

What hypothesis did we test?

What did we expect to happen?

What actually happened?

What did we decide?

And what did we learn?

A successful test shouldn't automatically become a permanent truth either.

Customers change. The mix changes. Competitors change. The product changes.

Something that worked a year ago may not produce the same outcome today.

Experimentation also means revisiting whether an effect survived as the context changed.

Seen this way, data stops being only an analytical asset.

It also becomes organizational memory.


Five questions before building another dashboard

I would summarize all of this with five questions:

  1. What do we believe is going to happen?

  2. What is the cheapest way to discover that we are wrong?

  3. What outcome would make us change our decision, and did we define it before looking at the data?

  4. What are we not willing to break while trying to improve this?

  5. Where will we record what we learned?

We don't need a sophisticated experimentation platform to start working this way.

We need to build the habit of turning opinions into hypotheses, hypotheses into evidence, and evidence into learning.


So, what does being data-driven really mean?

I don't think it means having more dashboards.

Or having a data lake, warehouse, or lakehouse.

Or hiring more data scientists.

Or measuring absolutely everything.

All of those things may be necessary.

But they are tools.

In the previous article, I argued that a good technology decision starts by asking:

What problem are we trying to solve?

Data adds a second question:

How will we know if we actually solved it?

At that point, data stops being something we look at after making a decision.

It becomes part of the mechanism through which we make decisions.

So, if I had to define a data-driven organization simply, I would say:

It isn't an organization that has an answer for everything. It's an organization that has built a method to quickly discover when it is wrong.

The goal shouldn't be to be right all the time.

It should be to learn fast enough to make a better decision next time.


Source

Kohavi, R., Deng, A., Frasca, B., Walker, T., Xu, Y. & Pohlmann, N. (2013). Online Controlled Experiments at Large Scale. Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining.