Most messaging systems assume a message is a unique event. It arrives, somebody consumes it, it ages out. Every retention knob you get is about time or size.
Kafka has that mode too. The interesting one is cleanup.policy=compact, where retention is per key: keep the latest value for every key, drop the superseded ones, and do it forever in bounded space. The topic is sized by how many distinct keys exist, not by how long you’ve been running.
That turns a topic into a system of record you consume as a stream.
What that buys you
Read a compacted topic from offset 0 and you have every key that currently exists, at its current value. A new consumer bootstraps by reading from the beginning and arrives at correct state. Nobody writes a backfill. There’s no separate full-dump path sitting alongside the incremental path, quietly disagreeing with it.
Keeping another system in sync stops being a project and becomes a consumer.
The usual arrangement is a nightly full export, an incremental feed, and a reconciliation job that tells you the two disagree without telling you which one is right. Drift detection is a whole genre of work that exists because the bootstrap path and the steady-state path are different code reading different sources.
Then you find the bug — an enrichment that was wrong all month, a field mapped to the wrong column. Reset the offset, replay, rebuild the projection. Same code, same data, correct answer, and no special one-off migration written under pressure.
The shape it replaces
Take anything you’d normally put behind a key-value API. Service A owns the users’ preferred shipping addresses, service B needs one on a hot path. What gets built is an endpoint on A, then a cache in B because that call can’t be on the hot path, then invalidation for the cache, then a fallback for when A is unavailable, then retries and timeouts around all of it.
Every one of those pieces exists to paper over the same thing: B is asking a remote system a question it needs answered locally.
A compacted topic inverts the direction. B consumes it into a local KV and reads its own copy. Updates arrive when they happen, so nothing needs invalidating. A going down stops the flow of new values and leaves B’s reads working. The cache, the invalidation, the fallback and the retry logic all collapse into a consumer and a local store.
The assumption underneath
All of this works because keys get written many times, which means the key space has to be bounded, or at least growing slower than the writes.
Orders fail that test. Every order is a new key, the count only ever goes up, and each one takes a handful of status changes before going cold. A compacted orders topic is about the same size as the full history, so you’ve gained nothing.
A user’s preferred shipping address is the opposite. One key per user, rewritten whenever they move, and the user count grows slowly. Same shape for an asset’s configuration, a meter’s latest reading, a tariff, a feature flag, an entitlement. The same few million keys, updated forever.
Systems built around unique event identifiers can’t do anything useful with this. Compacting a topic where every key appears once is a no-op, so the feature never gets built, and the architecture that would have used it never gets considered. The assumption gets baked in early and everything downstream inherits it.
Where the guarantee actually ends
Compaction preserves current state. It does not preserve history.
Once the cleaner has run over a segment, the intermediate values for a key are gone. If a row changed five times in August and once in September, August is not recoverable from a compacted topic. Rebuilding a projection works because the projection only needs the latest value. Auditing what a value was on the 12th does not.
Worth knowing too: cleanup.policy=compact,delete gives up the guarantee entirely. Records go if either policy says they can, so a key whose last write is old enough gets deleted outright and the topic no longer holds at least one record per key. If you want the snapshot property, compaction has to be the only policy.
Compaction also has no timing guarantee. It runs on inactive segments when the cleaner gets to them, so duplicates for a key can exist in the log at any moment. Consumers have to be idempotent, which they should be regardless.
Kafka isn’t the only one
Pulsar compacts topics, producing a separate compacted ledger consumers can read instead of the full log. NATS JetStream goes further and exposes the same idea directly as a KV store with per-subject limits and compare-and-set.
The capability is around. What’s missing is anyone treating it as the foundation. It gets documented as a way to shrink a topic, filed next to retention tuning, when it’s the difference between a message bus and a system of record.