Independent Submission
Request for Comments: 2610
Category: Informational
I. S. Hudzaifah
Bandung, Indonesia
January 2026

← Section 3, Publications

From logs to spans: what observability actually buys you

Abstract

Logs tell you something happened, metrics tell you how often, spans tell you why it was slow. How I went from one to the next, and the pipeline I run in prod.

1. The Question Logs Can’t Answer

Most systems start with logs, and honestly that’s fine. A service writes lines to a file, someone tails it during an incident, the problem gets found. One service on one server? Works.

It stops working quietly. A second service shows up, then a fifth, then fifteen. Now a request comes in through a firewall, hits an API, calls another service, writes to a queue and reads from two databases (on a good day). When a user says “the page is slow”, you’re no longer asking what happened. You’re asking where did the time go, and no single log file is going to tell you that.

That gap between “something happened” and “here’s why” is basically what observability is for. Over time I’ve come to think of it as three steps, each one sitting on top of the last: logs, then performance, then context.

2. Logs, But in One Place

The first step isn’t glamorous. Get every log line off the servers and into one store where it can be searched.

Two things make it pay off though. First, structure the logs. A line like user 4412 updated attendance in 2534ms is readable for a person but pretty much useless for a query. The same event as fields (user_id, action, duration_ms, http_uri) can be filtered, grouped and charted. Most logging libraries can emit JSON, so just use that.

Second, keep everything and filter later. It’s tempting to log only errors to save space, but the slow request that never errored is often teh one that matters. In our setup the web application firewall alone produces ~19.5 million log rows for one service in two weeks. Sounds like a lot, until you need the one row that explains a complaint. Columnar storage like ClickHouse makes keeping that volume cheap.

Once logs live in one place, a bunch of questions become a query instead of an SSH session. Which endpoints returned 4xx or 5xx in the last hour? What did the firewall block, and from where? Did the error start before or after the last deploy?

That alone is worth the effort. It still answers what though, not why.

3. From Logs to Performance

This is the step people skip. If each request log carries its response time, your logs are already a performance dataset. You don’t need a separate tool to start.

The trick is to stop looking at averages. An average response time of 300 ms can hide a tail where one request in ten takes eight seconds, and those are the requests users actually remember. Percentiles show the tail:

SELECT
  LogAttributes['http_uri']                     AS endpoint,
  count()                                       AS requests,
  quantile(0.50)(toFloat64(LogAttributes['rt'])) AS p50_s,
  quantile(0.90)(toFloat64(LogAttributes['rt'])) AS p90_s,
  quantile(0.99)(toFloat64(LogAttributes['rt'])) AS p99_s
FROM otel_logs
WHERE ServiceName = 'hcportal-waf'
  AND Timestamp > now() - INTERVAL 1 DAY
GROUP BY endpoint
ORDER BY p90_s DESC
LIMIT 20

A query like this turns a vague “it feels slow” into a ranked list: these twenty endpoints, this slow, this often. imo that list is where performance work should start, not with a hunch about which code looks inefficient.

It also changes the conversation after a fix. Instead of “it seems faster” you can show the p90 before and after. When we rebuilt an attendance report whose monthly download took up to 46 minutes, the result (under a minute) was something anyone could see on a chart. Nobody had to take my word for it. That one has its own post.

The list has a limit, though. It tells you which endpoint is slow. Not which part of it.

4. Spans Carry the Context

A slow endpoint is rarely slow everywhere. Usually it’s one step inside it: a query w/o an index, a call to another service, a loop that fetches one row at a time. To see that you need tracing.

A trace follows one request through the whole system. It’s made of spans, and each span is one unit of work, e.g. an HTTP handler, a DB query, a call to another service, a message put on a queue. Each span records when it started, how long it took, and which span it belongs to. Put them together and you get a tree:

GET /api/attendance/report              8.2 s
├─ auth: validate token                 0.01 s
├─ db: load employees                   0.12 s
├─ db: load attendance (per employee)   7.9 s   ← 2,000 small queries
└─ render: build spreadsheet            0.15 s

That tree answers the question logs couldn’t. The time didn’t go into the handler or the rendering. It went into two thousand small queries, one per employee. The fix (one query with the right index, or a precomputed summary) is pretty obvious at that point, and nobody had to guess.

Spans also carry context, and tbh that’s the part I value most. Every span in a trace shares the same trace ID. If you write logs with that trace ID attached, you can jump from a slow span straight to the log lines written during it, across every service the request touched. The trace gives you the shape and the logs give you the detail. You need both.

Context crosses service boundaries too. When service A calls service B it passes the trace ID along in a header. when it publishes a message to Kafka, the ID goes in the message headers, and the consumer (minutes later, on another machine) continues the same trace. A request that used to be fifteen unrelated log files becomes one story.

5. The Pipeline, and Where I’d Start

The setup I run isn’t exotic, and every part of it is replaceable:

services ──OpenTelemetry──▶ Kafka ──▶ ClickHouse     (logs, traces, metrics)
                                  └─▶ Elasticsearch  (full-text search)
                                           │
                                        HyperDX      (search, dashboards, alerts)

OpenTelemetry is the standard way to produce logs, traces and metrics, so services instrument once and the backend can change later w/o touching their code. Kafka sits in the middle as a buffer: if a store is slow or down for maintenance, services keep sending and nothing upstream blocks or drops data. I care about that one a lot, because observability should never become the reason a service is slow. ClickHouse stores the bulk. It’s columnar, compresses well, and answers aggregate questions over hundreds of millions of rows in seconds. HyperDX sits on top for search, traces and dashboards, so whoever’s fixing the problem doesn’t have to write SQL for every question (they still can when they need to).

Funny enough, the same pipeline later turned out to be useful for stuff that was never about errors, like tracking how teams use internal tools, or how engineers adopted an AI coding assistant after training. Once the data is in one place, new questions are cheap.

If a system has none of this today, I wouldn’t try to build it all at once. Roughly in this order:

  1. Ship logs to one place. Structured, with a service name and a timestamp. Nothing else matters until this works.
  2. Add response time and route to every request log. Now you’ve got performance data for free. Build the percentile query above and look at it weekly.
  3. Add a trace ID to every log line. Even before full tracing, this lets you follow one request across services.
  4. Instrument the slowest endpoint with spans. Start where the percentile list points, not everywhere at once.
  5. Propagate context across services and queues. That’s where traces stop being per-service and start telling the the whole story.

Without any of this, every incident starts with guessing and every fix ends with hoping. With it, a complaint turns into a query, the query points at a span, and the span points at the line of code that needs to change. That’s really all I want from it.

HudzaifahInformational[Page 1]