Blog

CPU was not the problem: how a self-hosted PostHog silently dropped a quarter of its events

The question was "CPU is high, should we add a bigger machine?". The real problem was a single-partition ingestion pipeline consuming slower than events arrived, a backlog that hit the 24-hour retention, 22–32% of events lost over five days, and the alert that would have caught it deleted by me ten days earlier.

  • PostHog
  • Capacity
  • Alerting
  • Incident

In mid-September a colleague asked in the team channel: ClickHouse CPU on the PostHog host is very high. Is it affecting the application? Should we scale the machine?

The answer, after looking: CPU was not the problem. The ingestion pipeline was consuming events slower than they arrived. The backlog grew from minutes to 24 hours in five days, hit the retention limit of the message queue, and from then on 22–32% of incoming events were dropped for good. The alert that would have caught this five days earlier had been deleted ten days before. By me.

This post is how “CPU is high” turned into “we are losing data”, the one wrong turn in the middle, and what carries over to other systems.

What the system looks like

This PostHog is self-hosted, the whole stack on one r5.2xlarge (8 vCPU, 64 GiB) using the official docker compose: capture writes events into Redpanda (Kafka-compatible), one Node.js ingestion service consumes them into ClickHouse, with Postgres, Redis, Celery workers and Temporal alongside.

Two structural limits decided everything that followed:

  • The events topic has 1 partition and 1 consumer, single-threaded. Measured consumption ceiling: about 470–500 events/s.
  • Redpanda retention is 24 hours. An event that enters the queue and is not consumed within 24 hours is deleted by retention.

I set the 24 hours in early September. Before that it was 1 hour: a full root disk in early September made Redpanda exit, capture could not accept events, and two days of data were lost for good; the 1-hour retention was the amplifier. Changing it to 24 hours bought “PostHog down for less than 24 hours = data is late, not lost”, at the cost of one more day of events on disk. The same week I built monitoring for the host from nothing: Prometheus, node-exporter, cAdvisor, blackbox, a Grafana dashboard and a dozen alert rules, all in Git and managed by Argo CD.

So by mid-September the machine had monitoring and alerts. The problem was not missing monitoring.

Why CPU was not the problem

When the question came, CloudWatch showed a daily average CPU of 88% with peaks at 99%, a flat line for a week. Breaking it down: ClickHouse used about 2.5 cores, almost all of it background merges and the insert pipeline; user queries used 1.7 CPU-hours per day.

High CPU was real, but it was a result, not the bottleneck. ClickHouse was busy merging because writes were high; writes were high because ingestion was consuming at full speed. Nothing was waiting for CPU.

The bottleneck of a queueing system is arrival rate against service rate, not CPU utilisation. The numbers:

09-06 09-14
Arrival rate 441 events/s 610 events/s (+38%)
Service ceiling 470–500 events/s 470–500 events/s
Queue backlog 2.2 million 52.6 million

Once arrival exceeds service, the backlog can only grow, whatever the CPU percentage says. A bigger machine lowers CPU; a single partition with a single consumer still serves 500/s.

How a backlog became data loss

Daily maximum ingestion delay, replayed from the host’s Prometheus:

Date Daily max delay
09-07 17 minutes
09-08 58 minutes
09-09 2.5 hours
09-10 7.4 hours
09-11 13.3 hours
09-14 24 hours

A delay of 24 hours means the tail of the backlog has reached retention: the oldest events in the queue are deleted before they are consumed. From the afternoon of 09-13, only 68–78% of each hour’s events actually reached ClickHouse. That is 22–32% permanent loss, not lateness.

There is a monitoring trap here. The dashboard’s “age of the newest event” stayed green the whole time, because it looks at insert time: as long as ingestion is running, there is always a freshly inserted event. It does not tell you that the event just inserted happened 24 hours ago. Backlog has to be measured on the queue itself: the distance between the consumer’s CURRENT offset and LOG-START, or the difference between an event’s occurrence time and now.

Why there was no alert

There was one. When I built the monitoring in early September I had a rule for “ingestion delay over 30 minutes”. After going live it fired every night into the morning: the arrival peak in that window was 750–870 events/s, above the service rate, delay climbed past an hour, and the queue drained by noon. I treated it as noise, widened it to 90 minutes, it still fired daily, and on 09-10 I deleted it. The commit message in Git said: the danger line is covered by other rules.

In hindsight it was right every day: the pipeline was losing every morning and catching up every afternoon. It was telling me the safety margin was thin; I heard noise.

And “covered by other rules” was never verified. Looking back, no rule watched the 24-hour line: an event stuck for 15 minutes would fire, Redpanda down would fire, disk full would fire, but “ingestion is running, just not keeping up” fired nothing.

On 09-21 I restored it as two tiers: 3 hours warning, 12 hours critical. Replaying 30 days of data, the 3-hour threshold fired exactly once in those 30 days, during this incident, from 09-09, five days before a human noticed.

The rule for deleting an alert is simple, I just did not follow it: prove coverage, do not claim it. Claiming is cheap; proving is not expensive either. Take the other rules’ expressions and ask, for the line you are about to delete, “in this state, who fires?”.

A wrong turn

In the middle there was a detour. Someone fed ClickHouse’s system.kafka_consumers to an AI and got a recommendation: 17 SYSTEM STOP KAFKA statements, on the grounds that these Kafka consumers were burning CPU.

I did not run them. I looked for a metric that could falsify the claim first. top -H by thread: ClickHouse’s Kafka threads were at 4% CPU. system.query_log: nobody had run STOP KAFKA in seven days. The one thing the AI got right: Redpanda was missing 11 topics and ClickHouse was logging thirty thousand “cannot get assignment” lines a minute, but that was noise, not load.

The same day CPU dropped from 88% to 60%, which looked like the recommendation had worked. The real cause was two orphaned docker compose stats processes I found on the host, each holding 7,885 connections to docker.sock and reading 5 MB/s. Killing them dropped the CPU. Unrelated to the incident; just cleaned up on the way.

Before acting on an AI conclusion in production, find one metric that could falsify it. Here one top -H was enough.

How it ended

Quantified, there were two paths: scale, or reduce input.

I wrote up scaling: partition 1 → 4 plus a second consumer takes the single host to about 1,000/s; beyond that, a 16-vCPU machine with 3–4 consumers for about 2,000/s. But adding partitions is not a hot operation (PostHog partitions by token + distinct_id to keep one person’s events ordered), and a real blue/green is impossible for a single-host ClickHouse, so a machine change means 5–10 minutes down. The ceiling was visible too: at the growth rate of the time, 1,000/s bought about 20 days, 2,000/s about 70.

Reduction came from the other direction: over 98% of this machine’s events came from five instrumentation points in the API layer, added in late August, growing from 300 thousand a day to 45 million a day. The same metrics were already available through another pipeline in the data warehouse. The data team and the developers decided to retire those five events, with the numbers above and the cost of both paths on the table.

The events were retired on the evening of 09-18, the backlog drained at about 650/s, and reached zero on the afternoon of 09-19. After: arrival 610/s → 3.8/s, backlog 52.6 million → 0, host CPU daily average 88% → 17% (CloudWatch), ClickHouse from about 2.5 cores to 0.2–0.35, merges at zero.

To be clear: the bottleneck disappeared, but it was not fixed; the input was removed. Single partition, single consumer, 500/s ceiling are all still there. The next high-volume instrumentation will hit it again. That is in the handover document, with the scaling plan and its cost.

What carries over

  • Watch arrival against service rate, not CPU. High CPU may just mean the service rate is saturated. Before adding a machine, ask: will the service rate change? With a single partition, no.
  • Retention is your tolerance for failure. It is not a disk parameter; it is “how long can the pipeline stop without losing data”. One hour means you tolerate one hour. Set both a time limit and a byte limit: the first decides the tolerance, the second guarantees the disk cannot fill.
  • Check which timestamp an “age of newest event” metric uses. Insert time is always fresh; occurrence time exposes the backlog.
  • Prove coverage before deleting an alert. Take the other rules’ expressions and ask who fires for the line you are removing. Deletion is the right first choice for noise, but this step comes first.
  • An alert that fires every day and resolves every day may be right. It may be measuring something that really happens every day, whose consequences have not accumulated yet.
  • Find a falsifying metric for an AI conclusion. It sees the one table you pasted, not top -H.
  • Record “problem gone” and “problem fixed” separately. Which one is in the handover decides whether the next person hits it again.

What I would do differently

The alert that would have given five days’ notice was deleted on the strength of an unverified sentence.

Starting over, I would make backlog the single most important metric of this machine on day one, measured from queue offsets or event occurrence time, with a threshold set as a fraction of retention (say one eighth) rather than starting from an intuitive “30 minutes” and widening it. False positives were a symptom of a threshold on the wrong quantity, not of a threshold set too low.

← All posts