TECTIUS
EN PT

All insights

When the Problem Disappears: Why Historical Data Matters

The goal isn't to collect everything. It is to preserve enough evidence to understand what happened.

There is a particular kind of problem that almost everyone has experienced. You take your car to the garage because something is wrong. Maybe the engine makes a strange noise. Maybe the warning light came on. Maybe the car has been behaving differently for the last few days.

The mechanic asks:

“Can you show me what it is doing?”

You start the engine.

Nothing.

You drive around the block.

Still nothing.

The car is now behaving perfectly. You know something happened. The mechanic has no reason to doubt you. But there is a problem: the evidence is gone.

Modern cars have become much better at keeping that evidence. Their electronic systems can record faults, operating conditions and other events that can later help diagnose a problem. But even then, there are limits. The system can only record the information it was designed to retain. And this is where the analogy with IT systems becomes interesting.

The problem with investigating the past

When an application is slow, a database query suddenly takes longer, or a server experiences an unexpected resource spike, the first reaction is usually to look at what is happening now. That makes sense. Is the CPU high? Are there errors? Is the database responding slowly? Are requests failing?

But many of the most difficult problems are not permanent. They happen intermittently. They happen under specific conditions. They happened three days ago. Or they happened for ten minutes last month and haven’t happened since. By the time somebody starts investigating, everything may be back to normal.

And then comes the inevitable question:

“Do we know what the system looked like when it happened?”

Sometimes the answer is no. Not because the system failed to work. Because we didn’t keep the evidence.

Today’s data can answer tomorrow’s questions

This is one of the less obvious aspects of monitoring and observability. We often collect data because we have a question we want to answer today. Is the service available? Is response time within the expected range? Are there errors? Is the infrastructure running out of capacity? These are important questions.

But historical data allows us to ask a different class of questions. Is performance getting worse? A response time of 800 ms might be perfectly acceptable. Unless it was 200 ms six months ago. Is resource consumption increasing? A database using 70% of available storage may not be a problem. Unless it was using 40% six months ago and the growth rate has not changed.

When did the degradation actually start? Was it after a deployment? A configuration change? A change in transaction volume? A change in user behaviour? Or did it start gradually, long before anybody noticed? Without historical information, these questions become much harder to answer. Sometimes impossible.

A single measurement tells us very little

Consider a simple example. Suppose we measure application response time and discover that a request currently takes 1.2 seconds. Is that good or bad? We don’t know. We need context.

Perhaps it normally takes 1.5 seconds. In that case, the system has actually improved. Perhaps it normally takes 200 milliseconds. Then something is clearly wrong. The value itself is not enough. The history gives the value meaning. This is why trends are often more useful than individual measurements.

Response time

1.2s |                         ●
1.0s |                    ●
0.8s |               ●
0.6s |          ●
0.4s |     ●
0.2s | ●
     +----------------------------> time

The individual measurements might all look reasonable. The trend is not. And that distinction matters operationally. A system that is slowly getting worse can remain inside today’s acceptable thresholds for a surprisingly long time. By the time somebody receives an alert, the underlying problem may have been developing for months.

Monitoring is not only about alerts

There is a tendency to think of monitoring as a mechanism for telling us when something is wrong. That is certainly one of its purposes. But monitoring data can also become historical evidence. It can help us establish baselines, identify trends and understand changes in system behaviour.

This is particularly important for performance. Performance is rarely a binary state of “working” or “not working”. A system can be:

  • 5% slower than last month;
  • consuming progressively more memory;
  • generating increasingly more database traffic;
  • processing fewer transactions per unit of CPU;
  • taking longer to complete background jobs;
  • experiencing more retries without actually generating visible errors.

None of these necessarily produces an immediate incident. But together they can tell us that something is changing. And sometimes the most valuable piece of information is not the value itself. It is the direction in which the value is moving.

But should we keep everything?

No.

That would be the easy answer, and usually the wrong one. Collecting data has a cost. There is storage. There is processing. There is network traffic. There is operational complexity. There is also the problem of noise.

A system producing millions of events per hour does not necessarily become more observable simply because we kept all of them. Sometimes we have simply created a very large pile of data.

The useful question is therefore not:

“How much data can we collect?”

It is:

“What evidence are we likely to need when something goes wrong?”

That requires engineering judgement. We need to consider what should be measured, at what level of detail, and for how long. Some information may only be useful for a few hours. Other information becomes much more valuable when retained for months or years. A detailed diagnostic event might not need long-term retention. A daily performance baseline might. A configuration change that takes two minutes to execute might be almost irrelevant at the time — until someone is investigating a problem six months later. Context matters.

The information we don’t need today may be the most useful tomorrow

This creates an interesting challenge. When designing a system, we know some of the questions we will need to answer. We don’t know all of them. And that is particularly important in operations. The question that matters during an incident is often not the question we expected to ask when the system was designed.

For example:

“Why did this transaction take 12 seconds?”

may eventually become:

“Did this only happen to one type of transaction?”

or:

“Did this start after the database migration?”

or:

“Was the application actually slow, or was it waiting for another service?”

or:

“Was the system already degrading before customers started reporting problems?”

Each question requires different evidence. We cannot predict every future question. But we can design systems so that they preserve enough context to investigate them. That is one of the reasons observability is more than simply putting a dashboard in front of a system.

Technology comes later

There is another important point here. This discussion is not really about a particular monitoring platform, database, logging framework or observability product. Those are implementation choices. The starting point should be the problem.

First we need to understand: What do we need to know? Then: What evidence do we need to retain to know it? Then: How much detail and history do we actually need? Only after that should we decide which technologies are appropriate.

The technology should serve the requirement. Not the other way around. A technically sophisticated monitoring platform does not automatically make a system observable. Just as collecting terabytes of logs does not automatically make an incident easier to investigate. The objective is not to collect more data. The objective is to have the right evidence when we need it.

Back to the garage

This brings us back to the car. The difficult part of diagnosing an intermittent problem is not necessarily finding the problem. Sometimes it is proving that the problem happened at all. The same is true in IT.

A system can be perfectly healthy when we investigate it. The database can be responding normally. The CPU can be quiet. The application can be processing requests within its expected limits. The incident is over. That doesn’t mean nothing happened. It may simply mean that we arrived too late. And if we did not preserve the right information, we may never know exactly what happened.

The best time to decide what evidence we will need to investigate a problem is before the problem happens. Because when the car finally reaches the garage, it may decide that today is a very good day to behave.

The goal isn’t to collect everything. It is to preserve enough evidence to understand what happened.