Skip to main content

27 posts tagged with "Replication"

Cross-cluster and cross-datacenter Kafka replication patterns.

View All Tags

Kafka Disaster Recovery in Practice: Real-World Stories from Production

· 12 min read
OSO Engineering
The team behind OSO Kafka Backup

Theory says your Kafka DR plan will work. Production has other ideas. These Kafka disaster recovery stories walk through four failures that actually happen — a region outage, an accidental topic deletion, a multi-datacenter partition, and a Kubernetes broker cascade — and the choices that decided whether recovery took minutes or hours. Learning from someone else's outage is cheaper than living through your own.

Kafka MirrorMaker Best Practices: Production-Ready Cross-Cluster Replication

· 10 min read
OSO Engineering
The team behind OSO Kafka Backup

MirrorMaker 2 is the standard tool for Kafka cross-cluster replication, but the default configuration is rarely production-ready. These Kafka MirrorMaker best practices close the gap between a working tutorial and a replication flow you can trust for disaster recovery and migration. The short version: match tasks to partitions, sync offsets and configs, alert on lag before it hurts, and never run MirrorMaker on your broker nodes.

Kafka Backup Tools Compared: MirrorMaker 2, Connect S3, Replicator, and Point-in-Time Backup

· 11 min read
OSO Engineering
The team behind OSO Kafka Backup

Choosing the right Kafka backup tool is the difference between a five-minute recovery and a five-hour scramble. There is no single winner. Each tool solves a specific problem: MirrorMaker 2 and MSK Replicator keep a warm cluster for regional failover, Kafka Connect S3 Sink archives topics to cheap object storage, and point-in-time backup restores a topic to the moment before a bad deploy. This guide compares all four, shows where each fits, and gives you a decision matrix so you can match tools to requirements instead of the other way around.

MSK Replicator: Cross-Region Kafka DR on AWS (Complete Guide)

· 11 min read
OSO Engineering
The team behind OSO Kafka Backup

MSK Replicator is a fully managed feature of Amazon MSK that copies topic data, consumer group offsets, and topic configurations between MSK clusters — across regions or across accounts — without you running any replication infrastructure. AWS manages the brokers, but cross-region disaster recovery for your streams is still your responsibility. MSK Replicator closes that gap. This guide covers its architecture, how to set it up for cross-region DR, failover procedures, and when to choose it over MirrorMaker 2.

Kafka Disaster Recovery: Active-Passive vs Active-Active Architectures

· 12 min read
OSO Engineering
The team behind OSO Kafka Backup

Kafka disaster recovery is the practice of keeping a second, independent copy of your streaming data and cluster metadata so you can resume service after a region outage, a bad deploy, or human error. Kafka's built-in replication — replication factor 3 with min.insync.replicas=2 — survives broker failure, but it does not survive a lost region or a deleted topic. The right DR design depends on three numbers: your recovery time objective (RTO), your recovery point objective (RPO), and your budget. This guide compares active-passive and active-active architectures and gives you a framework to choose.

Kafka Geo Replication: Multi-Region and Cross-Datacenter Patterns

· 6 min read
OSO Engineering
The team behind OSO Kafka Backup

Kafka geo replication copies topics between clusters in different regions or datacenters, so a regional outage does not take your streaming platform with it. In-cluster replication (RF=3) protects against broker loss inside one failure domain; geo replication protects against losing the domain itself. This guide compares the four patterns, shows what MirrorMaker 2 setup looks like over a WAN, and covers the failure mode replication cannot solve.

Kafka Replication Across Data Centers: Architecture Patterns for Multi-DC

· 10 min read
OSO Engineering
The team behind OSO Kafka Backup

Kafka replication across data centers comes down to one architectural decision: run a single stretched cluster spanning your datacenters, or run independent clusters connected by asynchronous replication. Stretched clusters give you zero data loss but demand low-latency links. Separate clusters tolerate WAN conditions but accept a non-zero recovery point. This guide covers both designs, when each wins, and how to build the second one with MirrorMaker 2.