Skip to main content

From Kafka PoC to Production: The Checklist

· 9 min read
OSO Engineering
The team behind OSO Kafka Backup

A Kafka PoC and a Kafka production cluster frequently run the exact same broker version, on the exact same defaults. The gap between them isn't code — it's everything nobody explicitly configured. Replication factor, backup, DR targets, monitoring, security, cluster topology, and retention all ship with defaults that are harmless for a three-week proof of concept and dangerous for a system carrying real traffic.

This is the checklist for closing that gap. It doesn't teach you anything new — it maps each silent default to the guide that already covers the fix.

Kafka Cross-Region Replication Patterns: Hub-and-Spoke, Mesh, and When to Use Neither

· 11 min read
OSO Engineering
The team behind OSO Kafka Backup

Kafka cross-region replication comes down to a small set of named topologies: active-passive, active-active, hub-and-spoke, and mesh. Most teams need one of the first two, and both already have a full implementation guide on this site. This post covers the other two — the ones that show up as a single sentence in every comparison article and never get built out: hub-and-spoke fan-in with real MirrorMaker 2 configuration, and an honest account of what mesh actually costs before you reach for it.

Kafka for Database Replication (CDC) — and What It Means for Backup

· 10 min read
OSO Engineering
The team behind OSO Kafka Backup

Kafka database replication almost never means what the words suggest. Searchers using that phrase are usually not asking about Kafka's own broker-to-broker replication — they're asking how to get changes out of a database and into Kafka, so downstream systems can react to them. That's change data capture (CDC), and Kafka Connect with a connector like Debezium is the standard way to do it.

This post untangles the terms, walks through the CDC architecture that "Kafka database replication" actually describes, and covers something that architecture's own documentation rarely does: what happens to those topics when you need to recover from a failure.

Kafka Streams Disaster Recovery: State Stores and Changelogs

· 9 min read
OSO Engineering
The team behind OSO Kafka Backup

Kafka Streams disaster recovery means recovering state stores, not just topics. Every Streams task keeps its aggregation, join, or window state on local disk — Kafka's replication never touches that disk. The only durable copy of that state Kafka manages is the changelog topic behind it, and that is the thing your recovery plan actually needs to protect.

This post explains why a Streams app needs a different recovery plan than a plain consumer, compares the three ways to get its state back, and works out what state-store size does to your recovery time.

MirrorMaker 1 Is Gone: How to Migrate Legacy Mirroring to MirrorMaker 2

· 10 min read
OSO Engineering
The team behind OSO Kafka Backup

The original Kafka MirrorMaker (kafka.tools.MirrorMaker, run through the kafka-mirror-maker.sh script) was deprecated in Kafka 3.0.0 under KIP-720 and fully removed in Kafka 4.0.0. If a cluster is still running it, moving that job onto MirrorMaker 2 isn't a config patch — it's a new config file format, a topic-naming decision, and, for the first time, real consumer-offset translation between clusters.

This post is the flag-by-flag migration guide: what changes, what the old options map to, and the order of operations that keeps a rollback path open until the cutover is proven.

Kafka MirrorMaker 2 Architecture: How the Connectors, Topics, and Clusters Fit Together

· 10 min read
OSO Engineering
The team behind OSO Kafka Backup

The Kafka MirrorMaker architecture is three Kafka Connect connectors run by one or more worker processes, coordinating cross-cluster replication through two distinct layers of internal Kafka topics. MirrorMaker 2 ships no bespoke server of its own — it inherits everything about how it runs from the Connect framework underneath it.

That inheritance explains most of what confuses people operating MirrorMaker 2 in production: why it needs a Connect cluster at all, what its internal topics actually are, how one config file scales past two clusters, and what changed when exactly-once support arrived.

Kafka Mirroring vs Replication: What's the Difference?

· 9 min read
OSO Engineering
The team behind OSO Kafka Backup

Kafka mirroring vs replication is a vocabulary question before it is an architecture question. In Kafka's own terms, replication copies partitions between brokers inside a single cluster. Mirroring copies topics from one cluster to another — the job MirrorMaker is named for. Both mechanisms produce copies of your data. They run in different places, protect against different failures, and in production you usually need both.

This post pins down where each term comes from and compares the two mechanisms side by side. It also shows why the copy neither of them makes — immutable, outside every cluster — still has to come from somewhere else.

Kafka Point-in-Time Recovery, Explained

· 10 min read
OSO Engineering
The team behind OSO Kafka Backup

Kafka point-in-time recovery (PITR) restores topic data to a specific moment — down to the millisecond — by filtering backed-up records on the timestamps they already carry. Kafka itself has no rewind button. Retention deletes forward, and replication copies corruption as faithfully as it copies good data. PITR therefore needs two things the platform does not provide: an immutable copy of the data outside the cluster, and a restore engine that filters that copy by time.

This post explains the mechanism end to end: why native Kafka features cannot rewind topic state, how timestamp filtering works, what happens to consumer groups, and the caveats — clock skew, transactions, compaction — that decide whether the restore you get matches the restore you planned.

Running Kafka in KRaft Mode with Docker Compose

· 11 min read
OSO Engineering
The team behind OSO Kafka Backup

Running Kafka in KRaft mode with Docker Compose takes one service and a handful of environment variables. The official apache/kafka image runs KRaft out of the box — no ZooKeeper container, no migration steps, and docker run -p 9092:9092 apache/kafka:4.3.1 alone gives you a working broker. A short Compose file makes that broker reproducible, and a slightly longer one gives you a 3-controller, 3-broker cluster that mirrors production topology.

This post walks through both files, the env-var conventions behind them, and how to keep data across restarts. For what KRaft mode actually is, see our KRaft guide.

KRaft Controller Quorum Explained: Voters, Observers, and kafka-metadata-quorum

· 11 min read
OSO Engineering
The team behind OSO Kafka Backup

The KRaft controller quorum is the small group of Kafka nodes that replicate the cluster metadata log and elect the active controller. A majority of these voters must acknowledge every metadata change before it commits. Each Kafka KRaft controller is either the active leader or a hot standby, and every broker follows the same log as a non-voting observer.

Our KRaft mode guide covers the architecture. This post covers running the quorum: how elections work, how to read kafka-metadata-quorum.sh output, how static and dynamic quorums differ, and how to add or remove controllers without breaking the majority.