# OSO Kafka Backup Documentation High-performance backup and restore for Apache Kafka > This file contains the complete documentation in a single file for LLM consumption. > For a lightweight index, see https://kafkabackup.com/llms.txt > For individual markdown files, download https://kafkabackup.com/markdown.zip --- title: Intro source_url: html: https://kafkabackup.com/intro md: https://kafkabackup.com/intro.md --- --- title: Frequently Asked Questions description: Common questions about OSO Kafka Backup architecture, operations, and best practices source_url: html: https://kafkabackup.com/well-architected/faq md: https://kafkabackup.com/well-architected/faq.md --- # Frequently Asked Questions Common questions about OSO Kafka Backup, organized by category. * * * OSO Kafka Backup is an open-source, high-performance backup and restore tool for Apache Kafka, written in Rust. It provides point-in-time recovery (PITR) for Kafka topics and consumer group offsets, supports multi-cloud storage backends (S3, Azure Blob Storage, GCS), and ships as a single static binary. The project is licensed under the MIT License. MirrorMaker 2 is a **replication** tool designed to mirror data between live Kafka clusters in real time. OSO Kafka Backup is a **backup and recovery** tool designed to create durable, versioned copies of your Kafka data in external object storage. Key differences: | Capability | MirrorMaker 2 | OSO Kafka Backup | | --- | --- | --- | | Primary purpose | Cross-cluster replication | Backup and restore | | Storage target | Another Kafka cluster | Object storage (S3, GCS, Azure Blob) | | Point-in-time recovery | No | Yes | | Offset recovery | Limited | Full consumer group offset restore | | Independent of Kafka | No (requires target cluster) | Yes (stores to object storage) | Use MirrorMaker 2 for active-active or active-passive cluster topologies. Use OSO Kafka Backup for disaster recovery, compliance archival, and point-in-time restore scenarios. Running a standby cluster with MirrorMaker 2 or Confluent Replicator means paying for two full Kafka clusters around the clock, plus the operational burden of maintaining the replication pipeline. Mirror setups are fragile — they can break during Kafka version upgrades, require ongoing certificate and configuration management, and add significant infrastructure cost. More critically, mirroring **propagates deletes**. If data is accidentally deleted or a bad producer pushes destructive events at 9:00 AM, that deletion is faithfully replicated to your standby cluster. You have no way to go back. OSO Kafka Backup stores immutable, incremental snapshots in object storage (S3, GCS, Azure Blob). This gives you: - **Point-in-time recovery** — restore to any previous backup, not just "current state" - **Delete protection** — deletions are never propagated to your backups - **Dramatic cost reduction** — object storage is 10–100× cheaper than running a second Kafka cluster - **Simpler operations** — no MirrorMaker to configure, monitor, or fix after upgrades If you need active-active replication for low-latency failover, mirroring is the right tool. If your goal is disaster recovery with the ability to restore within hours, backup to object storage is simpler, safer, and cheaper. Unlike Confluent Replicator, OSO Kafka Backup: - **Stores backups in external object storage** rather than requiring a destination Kafka cluster - **Supports point-in-time recovery (PITR)** to restore data to any arbitrary timestamp - **Recovers consumer group offsets** so applications resume from the correct position after restore - **Ships as a single binary** with no dependencies on the Confluent Platform or Connect framework - **Is open source** under the MIT License, with no per-broker licensing costs Yes. OSO Kafka Backup is built in Rust for memory safety and high performance. In production environments it achieves throughput exceeding 100 MB/s and operates with less than 500 MB of memory. It includes built-in checkpointing for crash resilience, Prometheus metrics for observability, and has been validated across enterprise workloads. OSO Kafka Backup supports any Kafka cluster that implements the Kafka protocol version 0.10 or later. This includes clusters running in both ZooKeeper mode and KRaft mode. The tool uses the standard Kafka consumer and producer APIs, so it is compatible with all Kafka distributions that adhere to the protocol. OSO Kafka Backup works with all major managed Kafka services, including: - **Amazon MSK** (both provisioned and serverless) - **Confluent Cloud** - **Aiven for Apache Kafka** - **Redpanda** (Kafka API-compatible) - **Azure Event Hubs for Kafka** (Kafka protocol endpoint) Any service that exposes a standard Kafka protocol endpoint is supported. OSO Kafka Backup also works with Kafka-compatible platforms that implement the Kafka wire protocol, including: - **AutoMQ** (cloud-native Kafka with tiered storage) - **WarpStream** (Kafka-compatible, zero-disk architecture) If the platform speaks the Kafka protocol, OSO Kafka Backup can back it up. Yes. The open source and Enterprise editions use the same backup storage format. You can start with the OSS edition, build your backup infrastructure, and upgrade to Enterprise at any time without migrating or re-creating existing backups. Typical triggers for upgrading to Enterprise: - You adopt a **Schema Registry** and need schema backup and restore - You need **data masking** or **GDPR compliance tools** (field-level redaction, right to be forgotten) - You require **RBAC** to control who can perform backup and restore operations - You want **priority support with SLAs** and a dedicated Slack channel The Enterprise licence is applied as a configuration change — no re-deployment or data migration is required. **OSS edition** includes: - Full backup and restore functionality - Point-in-time recovery (PITR) - Compression (zstd, lz4, none) - Prometheus metrics and monitoring - Consumer group offset backup and restore **Enterprise edition** adds: - Confluent CSFLE and DEK Registry metadata backup - Confluent and Apicurio Schema Registry backup and restore - Confluent RBAC metadata backup and restore - MSK ZooKeeper to KRaft migration workflows - Priority support with SLAs * * * There are two primary approaches: **Kubernetes CronJob** -- Run `kafka-backup backup` with `stop_at_current_offsets: true` on a schedule: ``` apiVersion: batch/v1 kind: CronJob metadata: name: kafka-backup-scheduled spec: schedule: "0 */6 * * *" # Every 6 hours jobTemplate: spec: template: spec: containers: - name: kafka-backup image: ghcr.io/osodevops/kafka-backup:latest args: ["backup", "--config", "/etc/kafka-backup/config.yaml"] restartPolicy: OnFailure ``` **Kafka Backup Operator** -- Use a `KafkaBackup` resource with `spec.schedule` to define schedules declaratively: ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: daily-backup spec: schedule: "0 0 2 * * * *" stopAtCurrentOffsets: true kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 credentialsSecret: name: s3-credentials ``` It depends on your **Recovery Point Objective (RPO)** — how much data you can afford to lose in a disaster: | Approach | RPO | Best for | | --- | --- | --- | | **Scheduled (cron)** — e.g., hourly | Up to 1 hour of data loss | Most disaster recovery scenarios | | **Continuous** | Near-zero data loss | Mission-critical streams where every message matters | For most organisations, **hourly scheduled backups** provide a practical balance between protection and simplicity. If your Kafka cluster fails, you lose at most one hour of data and can restore the rest from the last backup. Use continuous mode when your data stream is the source of truth (e.g., event-sourced architectures) and any data loss is unacceptable. Both modes use the same incremental mechanism — only new messages since the last checkpoint are captured — so the operational cost difference is minimal. Unlike cluster mirroring, OSO Kafka Backup **does not propagate deletes** to your backups. Backups are immutable snapshots stored in object storage. If a bad producer pushes destructive events, a topic is accidentally deleted, or a compaction policy removes data unexpectedly, your backups remain intact. To recover from an accidental deletion: 1. Identify the timestamp just before the deletion occurred 2. Configure a point-in-time restore with `time_window_end` set to that timestamp 3. Restore to the original or a new cluster This is a fundamental advantage over mirrored standby clusters, where deletes are faithfully replicated and there is no way to roll back. All incremental backup snapshots are retained in your object storage bucket indefinitely by default. You can restore to **any previous backup point**, not just the latest. For Kubernetes operator deployments, retention can be enforced per `KafkaBackup` with opt-in `spec.retention`. When enabled, the operator prunes complete backup sets after successful backup runs and can keep a minimum number of newest backups: ``` spec: retention: enabled: true maxAgeDays: 90 keepLast: 7 dryRun: true ``` Storage provider lifecycle policies are also supported and remain the right choice for backend-native retention controls. For example, you can configure S3 Lifecycle Rules to: - Transition older backups to cheaper storage tiers (e.g., S3 Glacier after 90 days) - Automatically delete backups older than a defined retention period (e.g., 1 year) OSO Kafka Backup does not delete previous snapshots when creating new ones unless operator-managed retention is explicitly enabled. Your backup history grows incrementally by default, and you control how long it is kept. Yes. Use `topics.include` and `topics.exclude` with wildcard patterns: ``` topics: include: - "orders.*" - "payments.*" - "inventory.updates" exclude: - "*.test" - "*.staging" ``` Patterns use glob-style matching. If `include` is not specified, all topics are backed up. The `exclude` list takes precedence over `include`. OSO Kafka Backup uses checkpoint-based incremental backups. A local SQLite database tracks the last committed offset for each topic-partition. On each backup run, the tool resumes consuming from the last checkpointed offset, so only new messages are read and stored. This makes subsequent backup runs significantly faster and reduces storage costs. **How to enable it:** - **Continuous mode** (`continuous: true`): Incremental tracking is automatic — the offset store is created and managed for you. - **One-shot or snapshot mode** (v0.13.5+): Add the `offset_storage` section to your config to enable incremental behavior: ``` offset_storage: db_path: /data/offsets.db sync_interval_secs: 30 ``` Without `offset_storage`, one-shot and snapshot backups start from `start_offset` (default: `earliest`) on every run, producing a full backup each time. With it, each run resumes from the last checkpoint. See the [Incremental Backups Guide](https://kafkabackup.com/guides/incremental-backups.md) for a step-by-step walkthrough. The checkpoint mechanism ensures crash resilience. If a backup run fails or is interrupted, the checkpoint database retains the last successfully committed offset for each partition. The next backup run automatically resumes from that point. No data is lost and no duplicate data is written to storage. Every backed-up message carries its source offset as an `x-original-offset` header by default (`include_offset_headers: true`; `restore.strip_offset_headers` removes it again if you need verbatim records). During restore, a three-phase process recovers consumer group positions: 1. **Restore messages** to the target cluster 2. **Plan offset reset** using `kafka-backup offset-reset plan` to compute the mapping between original and new offsets 3. **Execute offset reset** using `kafka-backup offset-reset execute` to commit the mapped offsets to the target cluster's consumer groups Yes. You can run multiple instances of OSO Kafka Backup concurrently, provided each instance is configured to back up a different set of topics. Use non-overlapping `topics.include` patterns to partition the workload. Do not configure multiple instances to back up the same topic-partition, as this will result in duplicate data in storage. Use the built-in validation command: ``` kafka-backup validate --deep --config /path/to/config.yaml ``` The `--deep` flag performs a full integrity check, verifying that all segments are present, checksums are valid, and the manifest is consistent with the stored data. The maximum message size is governed by the Kafka cluster's `max.message.bytes` configuration, which defaults to 1 MB. OSO Kafka Backup has been tested with messages up to 10 MB. If your cluster uses a non-default maximum, ensure the backup tool's consumer configuration matches (via `message.max.bytes` in the consumer properties). * * * Use the `time_window_start` and `time_window_end` parameters in your restore configuration, specified in epoch milliseconds: ``` restore: time_window_start: 1742817600000 # 2026-03-24 12:00:00 UTC time_window_end: 1742846400000 # 2026-03-24 20:00:00 UTC source: storage: type: s3 bucket: my-kafka-backups target: bootstrap_servers: "target-kafka:9092" ``` Only messages with timestamps within the specified window will be restored. **Bash:** ``` date -d "2026-03-24 12:00:00 UTC" +%s%3N # Output: 1742817600000 ``` **Python:** ``` from datetime import datetime int(datetime(2026, 3, 24, 12).timestamp() * 1000) # Output: 1742817600000 ``` **macOS (BSD date):** ``` date -j -u -f "%Y-%m-%d %H:%M:%S" "2026-03-24 12:00:00" +%s000 ``` Yes. You can restore any backup — or a subset of it — to a completely different cluster. This is useful for: - **Reproducing production bugs** in a staging environment with real data - **Populating test environments** with representative datasets - **Data analysis** on a separate cluster without impacting production Use `topic_mapping` to restore to different topic names and `source_partitions` to restore only specific partitions. You can also use `time_window_start` and `time_window_end` to restore a specific time slice of data rather than the full history. ``` restore: time_window_start: 1742817600000 time_window_end: 1742846400000 topic_mapping: "articles.production": "articles.staging" source: storage: type: s3 bucket: prod-kafka-backups target: bootstrap_servers: "staging-kafka:9092" ``` OSO Kafka Backup handles large, unbounded topics efficiently thanks to its incremental backup mechanism. Only new messages since the last checkpoint are read and stored on each backup run, regardless of the total topic size. Real-world characteristics: - **Tested with topics exceeding 800 GB and 20+ million messages** - **Single-partition topics are fully supported**, including those where message ordering is critical - **Restore time scales with data volume** — under optimal conditions, an 800 GB topic restores in approximately 20 minutes, though actual performance depends on storage backend throughput, network bandwidth, and target cluster write capacity For very large topics, consider: - Using `zstd` compression to reduce storage footprint (5–7× compression for JSON data) - Deploying the backup tool in the same region and availability zone as your Kafka cluster and storage backend - Monitoring the `kafka_backup_consumer_lag` metric to ensure backups keep pace with producers Yes. Specify the target cluster's `bootstrap_servers` in your restore configuration. The source and target clusters are completely independent. This is a core use case for disaster recovery -- restoring data to a standby cluster in a different region or cloud provider. Yes. Use the `topic_mapping` configuration to remap topic names during restore: ``` restore: topic_mapping: "orders.production": "orders.restored" "payments.production": "payments.restored" ``` Use the two-step offset reset workflow: ``` # Step 1: Generate the offset mapping plan kafka-backup offset-reset plan \ --path s3://kafka-backups/production \ --backup-id production-20260719 \ --groups order-service,payment-service \ --bootstrap-servers target-kafka:9092 \ --format json > offset-plan.json # Step 2: Review it, then execute with the same inputs kafka-backup offset-reset execute \ --path s3://kafka-backups/production \ --backup-id production-20260719 \ --groups order-service,payment-service \ --bootstrap-servers target-kafka:9092 ``` The plan maps original offsets to the corresponding offsets in the restored topic, accounting for any gaps or reordering. Restore duration depends on the data volume, storage backend read throughput, network bandwidth, and target cluster write capacity. Under optimal conditions, OSO Kafka Backup achieves approximately 100 MB/s restore throughput. For example, restoring 1 TB of data takes roughly 2.5 to 3 hours. Yes. Use the `source_partitions` configuration to specify which partitions to restore: ``` restore: source_partitions: [0, 1, 2, 5] ``` Only the specified partitions will be restored from the backup. OSO Kafka Backup **appends** data to the target topic; it does not overwrite or truncate existing data. If you need a clean restore, create a new topic (or use `topic_mapping` to restore to a different topic name) to avoid mixing existing and restored data. * * * OSO Kafka Backup supports: - **Amazon S3** - **Azure Blob Storage** - **Google Cloud Storage (GCS)** - **S3-compatible storage** (MinIO, Ceph, Wasabi, DigitalOcean Spaces) - **Local filesystem** (for testing and development) Estimate storage as: ``` storage_required = raw_data_size / compression_ratio ``` Compression ratios vary by data type: | Data Type | Compression (zstd) | Example | | --- | --- | --- | | JSON | 5:1 to 7:1 | 1 TB raw ≈ 200-300 GB compressed | | Avro | 2:1 to 3:1 | 1 TB raw ≈ 350-500 GB compressed | | Protobuf | 2:1 to 3:1 | 1 TB raw ≈ 350-500 GB compressed | | Already compressed | ~1:1 | No significant reduction | Backups are organized as follows: ``` backup-root/ ├── manifest.json ├── state/ │ └── offsets.db └── topics/ └── {topic-name}/ └── partition={id}/ ├── segment-000000000000.zst ├── segment-000000001000.zst └── ... ``` - `manifest.json` -- Metadata about the backup (topics, partitions, offsets, timestamps) - `offsets.db` -- SQLite offset store tracking backup progress (only with `continuous: true` or an `offset_storage:` section) - `topics/{topic}/partition={id}/segment-<20-digit start offset>.bin.zst` -- Compressed data segments Yes. Use the `describe` command to inspect backup metadata: ``` kafka-backup describe --config /path/to/config.yaml ``` You can also directly access objects in S3 (or other storage) using standard tools such as the AWS CLI, `gsutil`, or `az storage blob`. Segment files are compressed with the configured algorithm (e.g., zstd) and contain Kafka records in a binary format. Yes. Since backups are stored as standard objects, you can copy them between backends using tools like `aws s3 sync`, `gsutil rsync`, `azcopy`, or `rclone`. After copying, update your restore configuration to point to the new storage location. Yes. Configure the `endpoint` URL to point to your S3-compatible service: ``` storage: type: s3 bucket: my-backups region: us-east-1 endpoint: "https://minio.internal:9000" force_path_style: true ``` This works with MinIO, Ceph Object Gateway, Wasabi, DigitalOcean Spaces, and other S3-compatible services. * * * Under optimal conditions, OSO Kafka Backup achieves 100+ MB/s for both backup and restore operations. Actual throughput depends on: - Network bandwidth between Kafka, the backup tool, and storage - Storage backend write/read latency - Compression algorithm and level - Message size (larger messages achieve higher throughput) - Number of partitions being processed concurrently Typical memory usage is under 500 MB when processing 4 partitions concurrently. Memory consumption scales with the number of concurrent partitions and the configured segment size. For high-concurrency workloads, monitor RSS via the `process_resident_memory_bytes` Prometheus metric and adjust `segment_max_bytes` or concurrency settings accordingly. Refer to **PE-01: Throughput Optimisation** in the Performance Efficiency pillar. Key tuning parameters: - **Segment size:** Increase `segment_max_bytes` to reduce the number of storage write operations - **Fetch size:** Increase `fetch.max.bytes` and `max.partition.fetch.bytes` in the consumer config - **Compression level:** Use a lower zstd compression level (e.g., 1-3) for faster compression at the cost of slightly larger files - **Co-location:** Deploy the backup tool in the same region and availability zone as the Kafka cluster and storage backend Minimal. OSO Kafka Backup operates as a standard Kafka consumer. It does not require any broker restarts, plugins, or configuration changes. The impact is equivalent to adding another consumer to the cluster. For latency-sensitive workloads, consider configuring a dedicated consumer group and using rack-aware replica fetching. The [kafka-backup-demos](https://github.com/osodevops/kafka-backup-demos) repository includes a benchmark suite that generates synthetic workloads and measures backup/restore throughput under various configurations. Use it to establish baselines for your environment before deploying to production. * * * **Server-side encryption:** All major cloud storage providers offer server-side encryption (SSE-S3, SSE-KMS, Azure Storage Service Encryption, GCS default encryption). Enable this on your storage bucket for encryption at rest. **Confluent CSFLE metadata (Enterprise):** Enterprise backs up KEKs, encrypted DEKs, encrypted-subject details, and Schema Registry encryption rules from an existing Confluent CSFLE deployment. It does not encrypt backup segment files or plaintext records. See [Confluent CSFLE Metadata Backup](https://kafkabackup.com/enterprise/encryption.md). Set the security protocol and certificate paths in your configuration: ``` kafka: bootstrap_servers: "kafka:9093" security_protocol: "SSL" # or "SASL_SSL" for SASL + TLS ssl_ca_location: "/certs/ca.pem" ssl_certificate_location: "/certs/client.pem" ssl_key_location: "/certs/client-key.pem" ``` For mTLS, provide both the client certificate and key. The CA certificate is used to verify the broker's identity. OSO Kafka Backup supports the following SASL mechanisms: - **PLAIN** -- Username and password (use with TLS) - **SCRAM-SHA-256** -- Salted Challenge Response Authentication - **SCRAM-SHA-512** -- Salted Challenge Response Authentication (stronger hash) ``` kafka: security_protocol: "SASL_SSL" sasl_mechanism: "SCRAM-SHA-512" sasl_username: "backup-user" sasl_password: "${KAFKA_SASL_PASSWORD}" ``` The **Enterprise edition** provides GDPR compliance tools: - **Data masking:** Redact or mask personally identifiable information (PII) during backup - **Right to be forgotten:** Delete specific records from backups by key - **Field-level redaction:** Selectively redact fields within messages while preserving the rest of the record These features enable compliance with data protection regulations without sacrificing backup completeness. Multiple layers of access control are available: - **Enterprise RBAC:** Define roles (backup-operator, restore-operator, admin) with fine-grained permissions - **IAM policies:** Restrict access to storage buckets using AWS IAM, Azure RBAC, or GCP IAM - **Kubernetes RBAC:** Limit which service accounts can create `KafkaRestore` custom resources * * * Install the Kafka Backup Operator via Helm: ``` helm repo add oso https://charts.oso.sh helm repo update helm install kafka-backup-operator oso/kafka-backup-operator \ --namespace kafka-backup \ --create-namespace ``` Then create backup and restore resources using CRDs: ``` apiVersion: kafkabackup.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: production-backup spec: configRef: name: backup-config ``` OSO Kafka Backup Operator requires Kubernetes 1.24 or later. It is tested against the latest three minor versions of Kubernetes. Yes. You can deploy the Kafka Backup Operator into multiple Kubernetes namespaces, with each instance managing backups for a specific team, product line, or environment. This provides: - **Blast-radius isolation** — a misconfiguration in one namespace does not affect others - **Fine-grained IAM** — bind each operator to a dedicated IAM role with access to only its S3 bucket or prefix - **Independent lifecycle management** — each team can manage their own backup schedules and retention policies ``` # Team A — news platform backups apiVersion: kafkabackup.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: news-backup namespace: team-news spec: configRef: name: news-backup-config --- # Team B — streaming platform backups apiVersion: kafkabackup.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: streaming-backup namespace: team-streaming spec: configRef: name: streaming-backup-config ``` Alternatively, a single operator instance can manage multiple backup configurations in one namespace if you prefer centralised management. Yes. Store your `KafkaBackup` and `KafkaRestore` manifests in a Git repository. Configure an ArgoCD `Application` or Flux `Kustomization` pointing to the manifests directory: ``` apiVersion: argoproj.io/v1alpha1 kind: Application metadata: name: kafka-backup spec: source: repoURL: https://github.com/myorg/k8s-manifests path: kafka-backup/ targetRevision: main destination: server: https://kubernetes.default.svc namespace: kafka-backup ``` Yes. OSO Kafka Backup ships as a standalone static binary that runs on bare metal, virtual machines, and Docker containers. No Kubernetes or container orchestration is required: ``` # Download the binary curl -LO https://github.com/osodevops/kafka-backup/releases/latest/download/kafka-backup-linux-amd64 # Run directly ./kafka-backup-linux-amd64 backup --config /etc/kafka-backup/config.yaml ``` The operator exposes Prometheus metrics on port 8080. Create a `ServiceMonitor` to scrape them: ``` apiVersion: monitoring.coreos.com/v1 kind: ServiceMonitor metadata: name: kafka-backup-metrics spec: selector: matchLabels: app: kafka-backup endpoints: - port: metrics interval: 15s ``` Pair this with the provided Grafana dashboards from the [kafka-backup-demos](https://github.com/osodevops/kafka-backup-demos) repository for comprehensive visibility. * * * The Enterprise edition extends the OSS version with: - **Confluent CSFLE metadata backup** for KEKs, DEKs, encrypted subjects, and schema rules - **Confluent and Apicurio Schema Registry backup and restore** - **Confluent RBAC metadata backup and restore** - **MSK ZooKeeper to KRaft migration workflows** - **Priority support** with defined SLAs Contact the OSO sales team at [oso.sh](https://oso.sh) to discuss your requirements and obtain a licence key. Yes. The enterprise binary starts a 14-day trial automatically with no signup. Contact the sales team if you need an extended 30-day evaluation. Support is included with the Enterprise licence and delivered by OSO's own engineers in support hours (08:00 to 18:00 CET, Monday to Friday): - **Critical issues (P1):** response within 60 minutes and a workaround or restore path within 4 hours - **Other issues (P2 to P4):** response within 4 hours, 1 business day or 2 business days by priority - **Shared Slack Connect or Microsoft Teams channel** for direct communication with the engineering team - **Monthly performance report and quarterly service review** - **Out-of-hours P1 call-out** available as a priced option; there is no staffed 24x7 desk The full model, including escalation and the supported-version window, is on the [support page](https://kafkabackup.com/troubleshooting/support.md#enterprise-support). * * * Use the `-v` flag for debug-level logging or `-vv` for trace-level: ``` # Debug logging kafka-backup -v backup --config /path/to/config.yaml # Trace logging (very verbose) kafka-backup -vv backup --config /path/to/config.yaml ``` Alternatively, set the `RUST_LOG` environment variable: ``` RUST_LOG=debug kafka-backup backup --config /path/to/config.yaml ``` Check the following, in order: 1. **Network latency:** Measure latency between the backup tool and both the Kafka cluster and storage backend 2. **Storage write latency:** Monitor the `kafka_backup_storage_write_duration_seconds` Prometheus metric 3. **Compression overhead:** Try a faster compression level or algorithm (e.g., lz4 instead of zstd) 4. **Resource utilisation:** Check CPU and memory usage on the host running the backup 5. **Consumer lag:** Monitor `kafka_backup_consumer_lag` to see if the tool is keeping up with producers Verify the following: 1. **Bootstrap servers:** Ensure the `bootstrap_servers` address is correct and resolvable 2. **TLS certificates:** Verify certificates are valid, not expired, and the CA chain is complete 3. **Network connectivity:** Confirm the backup tool can reach the Kafka brokers on the configured port (e.g., `telnet kafka-broker 9093`) 4. **Firewall rules:** Check that security groups, NACLs, or firewall rules allow traffic on the Kafka port 5. **Kafka ACLs:** Ensure the backup user has `READ` and `DESCRIBE` permissions on the target topics and consumer group Follow these steps: 1. **Validate the backup** first: `kafka-backup validate --deep --config /path/to/config.yaml` 2. **Check target connectivity:** Verify the restore tool can reach the target Kafka cluster 3. **Verify IAM/storage permissions:** Ensure the restore process has read access to the backup storage location 4. **Check disk space:** Ensure sufficient local disk space for temporary decompression buffers 5. **Review error logs:** Enable debug logging (`-v`) and check for specific error messages - **Bug reports:** Open an issue on GitHub at [github.com/osodevops/kafka-backup/issues](https://github.com/osodevops/kafka-backup/issues) - **Community support:** Start a discussion at [GitHub Discussions](https://github.com/osodevops/kafka-backup/discussions) - **Enterprise support:** Use your dedicated Slack channel or contact the support team directly --- title: Index source_url: html: https://kafkabackup.com/getting-started/index md: https://kafkabackup.com/getting-started/index.md --- --- title: 5-Minute Quickstart description: Get your first Kafka backup running in under 5 minutes using Docker source_url: html: https://kafkabackup.com/getting-started/quickstart md: https://kafkabackup.com/getting-started/quickstart.md --- # 5-Minute Quickstart This guide will have you backing up and restoring Kafka topics in under 5 minutes using Docker. - Docker and Docker Compose installed - A running Kafka cluster (we'll provide a test setup) If you don't have a Kafka cluster, start one with Docker Compose: docker-compose.yml ``` version: '3.8' services: kafka: image: confluentinc/cp-kafka:7.5.0 hostname: kafka ports: - "9092:9092" environment: KAFKA_NODE_ID: 1 KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: CONTROLLER:PLAINTEXT,PLAINTEXT:PLAINTEXT KAFKA_LISTENERS: PLAINTEXT://0.0.0.0:9092,CONTROLLER://0.0.0.0:9093 KAFKA_ADVERTISED_LISTENERS: PLAINTEXT://kafka:9092 KAFKA_CONTROLLER_LISTENER_NAMES: CONTROLLER KAFKA_CONTROLLER_QUORUM_VOTERS: 1@kafka:9093 KAFKA_PROCESS_ROLES: broker,controller KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR: 1 CLUSTER_ID: MkU3OEVBNTcwNTJENDM2Qk ``` ``` docker-compose up -d ``` ``` # Create a topic docker exec -it kafka kafka-topics --create \ --topic test-topic \ --bootstrap-server localhost:9092 \ --partitions 3 \ --replication-factor 1 # Produce some test messages docker exec -it kafka bash -c 'for i in {1..100}; do echo "message-$i"; done | kafka-console-producer --broker-list localhost:9092 --topic test-topic' # Verify messages docker exec -it kafka kafka-console-consumer \ --bootstrap-server localhost:9092 \ --topic test-topic \ --from-beginning \ --max-messages 5 ``` Create a backup configuration file: backup.yaml ``` mode: backup backup_id: "quickstart-backup" source: bootstrap_servers: - kafka:9092 topics: include: - test-topic storage: backend: filesystem path: "/data/backups" backup: compression: zstd compression_level: 3 ``` ``` # Create backup directory mkdir -p ./backups # Run the backup docker run --rm \ --network host \ -v $(pwd)/backup.yaml:/config/backup.yaml \ -v $(pwd)/backups:/data/backups \ ghcr.io/osodevops/kafka-backup:latest \ backup --config /config/backup.yaml ``` You should see output like: ``` [INFO] Starting backup: quickstart-backup [INFO] Connected to Kafka cluster [INFO] Backing up topic: test-topic (3 partitions) [INFO] Partition 0: 34 records backed up [INFO] Partition 1: 33 records backed up [INFO] Partition 2: 33 records backed up [INFO] Backup completed successfully [INFO] Total: 100 records, 2.4 KB compressed ``` List the backup: ``` docker run --rm \ -v $(pwd)/backups:/data/backups \ ghcr.io/osodevops/kafka-backup:latest \ list --path /data/backups ``` Output: ``` Available Backups: quickstart-backup Created: 2024-12-03T10:00:00Z Topics: 1 Records: 100 Size: 2.4 KB (compressed) ``` Get detailed backup info: ``` docker run --rm \ -v $(pwd)/backups:/data/backups \ ghcr.io/osodevops/kafka-backup:latest \ describe --path /data/backups --backup-id quickstart-backup ``` ``` # Delete the topic docker exec -it kafka kafka-topics --delete \ --topic test-topic \ --bootstrap-server localhost:9092 # Verify it's gone docker exec -it kafka kafka-topics --list \ --bootstrap-server localhost:9092 ``` Create a restore configuration: restore.yaml ``` mode: restore backup_id: "quickstart-backup" target: bootstrap_servers: - kafka:9092 storage: backend: filesystem path: "/data/backups" restore: dry_run: false ``` Run the restore: ``` docker run --rm \ --network host \ -v $(pwd)/restore.yaml:/config/restore.yaml \ -v $(pwd)/backups:/data/backups \ ghcr.io/osodevops/kafka-backup:latest \ restore --config /config/restore.yaml ``` ``` # Check the topic exists docker exec -it kafka kafka-topics --list \ --bootstrap-server localhost:9092 # Verify the messages docker exec -it kafka kafka-console-consumer \ --bootstrap-server localhost:9092 \ --topic test-topic \ --from-beginning \ --max-messages 5 ``` You should see your original messages restored. In this quickstart, you: 1. Created a Kafka topic with test data 2. Backed up the topic to local storage 3. Verified the backup contents 4. Simulated a disaster by deleting the topic 5. Restored the topic from backup 6. Verified the restored data - **[CLI Basics](https://kafkabackup.com/getting-started/cli-basics.md)** - Learn all available commands - **[First Backup Tutorial](https://kafkabackup.com/getting-started/first-backup.md)** - More detailed walkthrough - **[Backup to S3](https://kafkabackup.com/guides/backup-to-s3.md)** - Use cloud storage - **[Point-in-Time Recovery](https://kafkabackup.com/guides/restore-pitr.md)** - Restore to specific timestamps - **[Kubernetes Operator](https://kafkabackup.com/operator.md)** - Automated backups with CRDs ``` # Stop and remove containers docker-compose down -v # Remove backup files rm -rf ./backups ``` --- title: CLI Basics description: Learn the OSO Kafka Backup command-line interface - commands, flags, and common operations source_url: html: https://kafkabackup.com/getting-started/cli-basics md: https://kafkabackup.com/getting-started/cli-basics.md --- # CLI Basics OSO Kafka Backup provides a powerful command-line interface with 12 commands for backup, restore, and offset management operations. All commands support these global options: | Option | Description | | --- | --- | | `-v, --verbose` | Enable verbose logging (use `-vv` for trace level) | | `--help` | Show help for any command | | `--version` | Show version information | ``` # Enable debug logging kafka-backup -v backup --config backup.yaml # Enable trace logging (very verbose) kafka-backup -vv backup --config backup.yaml ``` | Command | Description | | --- | --- | | `backup` | Run a backup operation | | `restore` | Restore data from a backup | | `list` | List available backups | | `status` | Show backup status and statistics | | `describe` | Show detailed backup manifest | | `validate` | Validate backup integrity | | `validate-restore` | Validate restore configuration (dry-run) | | `show-offset-mapping` | Display offset mapping for consumer groups | | `offset-reset` | Generate or execute offset reset plans | | `offset-reset-bulk` | Parallel bulk offset reset | | `offset-rollback` | Snapshot and rollback consumer offsets | | `three-phase-restore` | Complete restore with automatic offset reset | Run a backup operation from a Kafka cluster to storage. ``` kafka-backup backup --config ``` **Options:** | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-c, --config` | string | Yes | Path to backup configuration YAML | **Example:** ``` kafka-backup backup --config /etc/kafka-backup/backup.yaml ``` Restore data from a backup to a Kafka cluster. ``` kafka-backup restore --config ``` **Options:** | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-c, --config` | string | Yes | Path to restore configuration YAML | **Example:** ``` kafka-backup restore --config /etc/kafka-backup/restore.yaml ``` List available backups in a storage location. ``` kafka-backup list --path [--backup-id ] ``` **Options:** | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-p, --path` | string | Yes | Path to storage location | | `-b, --backup-id` | string | No | Show details for specific backup | **Examples:** ``` # List all backups kafka-backup list --path /var/lib/kafka-backup/data # List all backups from S3 kafka-backup list --path s3://my-bucket/backups # Show details for a specific backup kafka-backup list --path /data --backup-id daily-backup-001 ``` Show status and statistics of a backup. Supports two modes: 1. **Static inspection**: Inspect stored backup artifacts 2. **Live monitoring**: Monitor a running backup in real-time ``` # Static inspection kafka-backup status --path --backup-id [--db-path ] # Live monitoring kafka-backup status --config [--watch] [--interval ] ``` **Options:** | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-p, --path` | string | No\* | Path to storage location (static inspection) | | `-b, --backup-id` | string | No\* | Backup ID to show status for (static inspection) | | `--db-path` | string | No | Path to offset database for progress tracking | | `-c, --config` | string | No\* | Path to config file (live monitoring) | | `--watch` | flag | No | Enable continuous watch mode (requires --config) | | `--interval` | int | No | Refresh interval in seconds (default: 2) | \*Either `--config` OR both `--path` and `--backup-id` are required. **Examples:** ``` # Static inspection of stored backup kafka-backup status --path /data --backup-id backup-001 # Live monitoring (one-shot) kafka-backup status --config backup.yaml # Live monitoring with continuous watch kafka-backup status --config backup.yaml --watch # Custom refresh interval kafka-backup status --config backup.yaml --watch --interval 5 ``` **Static inspection output includes:** - Manifest information (created date, source cluster) - Topic, partition, and segment counts - Record and size statistics - Compression ratio - Offset tracking status per partition **Live monitoring output includes:** - Real-time progress (records, bytes, throughput) - Consumer lag per partition - Component health status - Compression ratio and error count Show detailed backup manifest information. ``` kafka-backup describe --path --backup-id [--format ] ``` **Options:** | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-p, --path` | string | Yes | Path to storage location | | `-b, --backup-id` | string | Yes | Backup ID to describe | | `-f, --format` | string | No | Output format: `text` (default), `json`, `yaml` | **Examples:** ``` # Text output (default) kafka-backup describe --path /data --backup-id backup-001 # JSON output for scripting kafka-backup describe --path /data --backup-id backup-001 --format json # YAML output kafka-backup describe --path /data --backup-id backup-001 -f yaml ``` Validate backup integrity. ``` kafka-backup validate --path --backup-id [--deep] ``` **Options:** | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-p, --path` | string | Yes | Path to storage location | | `-b, --backup-id` | string | Yes | Backup ID to validate | | `--deep` | bool | No | Perform deep validation (read all segments) | **Examples:** ``` # Quick validation (check existence and metadata) kafka-backup validate --path /data --backup-id backup-001 # Deep validation (read and verify all segments) kafka-backup validate --path /data --backup-id backup-001 --deep ``` **Validation Report:** - Segments checked/valid/missing/corrupted - Records validated - Issues found with details - Final result: VALID or INVALID Validate a restore configuration without executing it. ``` kafka-backup validate-restore --config [--format ] ``` **Options:** | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-c, --config` | string | Yes | Path to restore configuration | | `-f, --format` | string | No | Output format: `text`, `json`, `yaml` | **Example:** ``` kafka-backup validate-restore --config restore.yaml --format json ``` **Report includes:** - Restore status (VALID or INVALID) - Topics and segments to process - Records and bytes to restore - Time range - Errors and warnings Display offset mapping for a backup (useful for consumer group reset). ``` kafka-backup show-offset-mapping --path --backup-id [--format ] ``` **Options:** | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-p, --path` | string | Yes | Path to storage location | | `-b, --backup-id` | string | Yes | Backup ID | | `-f, --format` | string | No | Output: `text`, `json`, `yaml`, `csv` | **Example:** ``` kafka-backup show-offset-mapping --path /data --backup-id backup-001 --format csv ``` Generate or execute consumer group offset reset plans. **Subcommands:** ``` # Generate a reset plan kafka-backup offset-reset plan \ --path \ --backup-id \ --groups \ --bootstrap-servers # Execute a reset plan kafka-backup offset-reset execute \ --path \ --backup-id \ --groups \ --bootstrap-servers # Generate a shell script for manual execution kafka-backup offset-reset script \ --path \ --backup-id \ --groups \ --bootstrap-servers \ --output reset.sh ``` Execute bulk parallel offset reset (50x faster than sequential). ``` kafka-backup offset-reset-bulk \ --path \ --backup-id \ --groups \ --bootstrap-servers \ [--max-concurrent ] \ [--max-retries ] ``` **Options:** | Flag | Type | Default | Description | | --- | --- | --- | --- | | `--max-concurrent` | int | 50 | Maximum concurrent requests | | `--max-retries` | int | 3 | Maximum retry attempts | | `--security-protocol` | string | PLAINTEXT | Security protocol | Snapshot and rollback consumer group offsets. **Subcommands:** ``` # Create a snapshot before making changes kafka-backup offset-rollback snapshot \ --path /data/snapshots \ --groups my-group \ --bootstrap-servers localhost:9092 \ --description "Before migration" # List available snapshots kafka-backup offset-rollback list --path /data/snapshots # Show snapshot details kafka-backup offset-rollback show \ --path /data/snapshots \ --snapshot-id snapshot-20241203-100000 # Rollback to a previous snapshot kafka-backup offset-rollback rollback \ --path /data/snapshots \ --snapshot-id snapshot-20241203-100000 \ --bootstrap-servers localhost:9092 # Verify offsets match a snapshot kafka-backup offset-rollback verify \ --path /data/snapshots \ --snapshot-id snapshot-20241203-100000 \ --bootstrap-servers localhost:9092 # Delete a snapshot kafka-backup offset-rollback delete \ --path /data/snapshots \ --snapshot-id snapshot-20241203-100000 ``` Run a complete three-phase restore with automatic offset reset. ``` kafka-backup three-phase-restore --config ``` This command performs: 1. **Phase 1**: Collect offset headers from source backup 2. **Phase 2**: Restore data to target cluster 3. **Phase 3**: Reset consumer group offsets ``` #!/bin/bash # Verify yesterday's backup BACKUP_PATH="/var/lib/kafka-backup/data" BACKUP_ID="daily-$(date -d yesterday +%Y%m%d)" # Quick validation kafka-backup validate --path "$BACKUP_PATH" --backup-id "$BACKUP_ID" # Get backup statistics kafka-backup describe --path "$BACKUP_PATH" --backup-id "$BACKUP_ID" --format json ``` ``` #!/bin/bash # Snapshot offsets before migration kafka-backup offset-rollback snapshot \ --path /data/snapshots \ --groups "app-consumer,analytics-consumer" \ --bootstrap-servers broker-1:9092,broker-2:9092 \ --description "Pre-migration snapshot" ``` ``` #!/bin/bash # Full disaster recovery restore # 1. Validate the backup kafka-backup validate --path s3://backups/kafka --backup-id latest --deep # 2. Validate restore config kafka-backup validate-restore --config dr-restore.yaml # 3. Execute three-phase restore kafka-backup three-phase-restore --config dr-restore.yaml ``` - **[First Backup Tutorial](https://kafkabackup.com/getting-started/first-backup.md)** - Step-by-step backup walkthrough - **[CLI Reference](https://kafkabackup.com/reference/cli-reference.md)** - Complete command documentation - **[Configuration Reference](https://kafkabackup.com/reference/config-yaml.md)** - All configuration options --- title: Your First Backup description: Complete step-by-step guide to backing up and restoring your first Kafka topic source_url: html: https://kafkabackup.com/getting-started/first-backup md: https://kafkabackup.com/getting-started/first-backup.md --- # Your First Backup This tutorial walks you through a complete backup and restore cycle, explaining each step in detail. - How to create a backup configuration - How to run a backup operation - How to verify backup integrity - How to restore from backup - How to handle consumer group offsets - OSO Kafka Backup installed ([Installation Guide](https://kafkabackup.com/deployment.md)) - Access to a Kafka cluster - Storage location (local path or cloud bucket) Before creating a backup, decide: 1. **Which topics to back up?** You can use explicit names or patterns 2. **Where to store backups?** Local filesystem, S3, Azure, or GCS 3. **What compression to use?** Zstd (best ratio), LZ4 (fastest), or none 4. **Starting offset?** From earliest (full backup) or latest (incremental) Create a file named `backup.yaml`: backup.yaml ``` # Backup mode mode: backup # Unique identifier for this backup backup_id: "production-backup-001" # Source Kafka cluster source: bootstrap_servers: - broker-1.kafka.svc:9092 - broker-2.kafka.svc:9092 - broker-3.kafka.svc:9092 # Optional: Security configuration # security: # security_protocol: SASL_SSL # sasl_mechanism: SCRAM-SHA256 # sasl_username: backup-user # sasl_password: ${KAFKA_PASSWORD} # Environment variable # Topics to back up topics: include: - orders # Explicit topic name - payments # Another topic - "events-*" # Wildcard pattern exclude: - "__consumer_offsets" # Internal topics - "_schemas" # Schema registry topic # Storage destination storage: backend: filesystem # Options: filesystem, s3, azure, gcs path: "/var/lib/kafka-backup/data" # For S3: # backend: s3 # bucket: my-backup-bucket # region: us-west-2 # prefix: kafka-backups/production # Backup settings backup: # Compression compression: zstd # Options: zstd, lz4, none compression_level: 3 # 1-22 for zstd (higher = better ratio, slower) # Starting point start_offset: earliest # Options: earliest, latest # Segment settings segment_max_bytes: 134217728 # 128 MB per segment segment_max_interval_ms: 60000 # Force segment roll every 60s # Checkpointing for resumable backups checkpoint_interval_secs: 30 # Include offset headers (required for offset reset) include_offset_headers: true # Source cluster identifier (for tracking) source_cluster_id: "production-cluster" # Optional: Snapshot mode (capture current offsets and exit when caught up) # stop_at_current_offsets: true # Optional: Performance tuning # max_concurrent_partitions: 8 # Parallel partition tasks (default: 8) # poll_interval_ms: 100 # Poll delay in ms (default: 100) ``` | Setting | Purpose | | --- | --- | | `backup_id` | Unique name for this backup; used for restore | | `bootstrap_servers` | Kafka broker addresses | | `topics.include` | Topics to back up (names or patterns) | | `topics.exclude` | Topics to skip | | `compression` | Reduce storage size and costs | | `start_offset: earliest` | Back up all data from beginning | | `checkpoint_interval_secs` | How often to save progress | | `include_offset_headers` | Add `x-original-offset` / `x-original-timestamp` headers to every archived record for consumer offset reset (default `true`) | | `stop_at_current_offsets` | Snapshot mode: exit after catching up to current offsets | | `max_concurrent_partitions` | Number of partitions to process in parallel | Execute the backup: ``` kafka-backup backup --config backup.yaml ``` With verbose logging: ``` kafka-backup -v backup --config backup.yaml ``` ``` [2024-12-03T10:00:00Z INFO] Starting backup: production-backup-001 [2024-12-03T10:00:00Z INFO] Connecting to Kafka cluster... [2024-12-03T10:00:01Z INFO] Connected to cluster: production-cluster [2024-12-03T10:00:01Z INFO] Discovered topics matching patterns: - orders (6 partitions) - payments (3 partitions) - events-clickstream (12 partitions) - events-pageviews (12 partitions) [2024-12-03T10:00:02Z INFO] Starting backup of 4 topics, 33 partitions [2024-12-03T10:00:02Z INFO] Topic: orders Partition 0: 150,234 records (earliest: 0, latest: 150233) Partition 1: 148,892 records ... [2024-12-03T10:05:32Z INFO] Backup completed successfully [2024-12-03T10:05:32Z INFO] Summary: Topics: 4 Partitions: 33 Records: 2,456,789 Uncompressed: 1.2 GB Compressed: 245 MB Compression ratio: 4.9x Duration: 5m 30s Throughput: 7,445 records/sec ``` > [!TIP] > > [!NOTE] > > Backup Modes > > [!NOTE] > > By default, backup runs in **one-shot mode** (backs up all data from `start_offset` to current high watermark, then exits). For **continuous streaming replication**, set `continuous: true`. For **snapshot mode** (v0.8.0+), set `stop_at_current_offsets: true` — this captures high watermarks at the start and exits once all partitions reach them, providing a consistent point-in-time snapshot ideal for scheduled DR backups. > [!TIP] > > [!NOTE] > > Incremental Backups (v0.13.5+) > > [!NOTE] > > To make one-shot or snapshot backups **incremental** (resume from where the last run stopped instead of re-backing up everything), add `offset_storage` to your config: > > > > ``` > > offset_storage: > > db_path: /data/offsets.db > > sync_interval_secs: 30 > > ``` > > > > With this, the first run backs up all data and saves progress. Subsequent runs with the same `backup_id` pick up from the last saved offset — only new messages are backed up. This is ideal for scheduled backups (e.g. hourly cron jobs) where you want each run to be fast and avoid duplicating work. See the [Incremental Backups Guide](https://kafkabackup.com/guides/incremental-backups.md) for details. > [!TIP] > > [!NOTE] > > Graceful Shutdown > > [!NOTE] > > You can safely stop a running backup with `Ctrl+C` (or `SIGTERM`). The process will flush in-progress segments and save a checkpoint before exiting, so it can resume from where it left off. ``` kafka-backup list --path /var/lib/kafka-backup/data ``` ``` Available Backups: ───────────────────────────────────────────────────────────── production-backup-001 Created: 2024-12-03T10:00:00Z Source: production-cluster Topics: 4 Partitions: 33 Records: 2,456,789 Size: 245 MB (compressed) ───────────────────────────────────────────────────────────── ``` ``` kafka-backup describe --path /var/lib/kafka-backup/data --backup-id production-backup-001 ``` ``` Backup: production-backup-001 ════════════════════════════════════════════════════════════ Metadata: Created: 2024-12-03T10:00:00Z Source Cluster: production-cluster Compression: zstd (level 3) Statistics: Topics: 4 Partitions: 33 Segments: 156 Records: 2,456,789 Uncompressed: 1.2 GB Compressed: 245 MB Compression Ratio: 4.9x Time Range: Earliest Message: 2024-11-01T00:00:00Z Latest Message: 2024-12-03T09:59:59Z Topics: orders 6 partitions 523,456 records payments 3 partitions 234,567 records events-click 12 partitions 890,123 records events-pages 12 partitions 808,643 records ``` ``` # Quick validation kafka-backup validate --path /var/lib/kafka-backup/data --backup-id production-backup-001 # Deep validation (reads all data) kafka-backup validate --path /var/lib/kafka-backup/data --backup-id production-backup-001 --deep ``` ``` Validation Report: production-backup-001 ════════════════════════════════════════════════════════════ Segments: Checked: 156 Valid: 156 Missing: 0 Corrupted: 0 Records Validated: 2,456,789 Result: ✓ VALID ``` Create `restore.yaml`: restore.yaml ``` mode: restore backup_id: "production-backup-001" # Target Kafka cluster (can be different from source) target: bootstrap_servers: - dr-broker-1.kafka.svc:9092 - dr-broker-2.kafka.svc:9092 storage: backend: filesystem path: "/var/lib/kafka-backup/data" restore: # Optional: Point-in-time recovery # time_window_start: 1701417600000 # Unix ms timestamp # time_window_end: 1701504000000 # Optional: Topic remapping # topic_mapping: # orders: orders_restored # payments: payments_dr # Consumer offset handling consumer_group_strategy: skip # Options: skip, header-based, timestamp-based, cluster-scan, manual # Add x-original-* / x-source-partition headers to restored records (default: false). # The backup already archived x-original-offset / x-original-timestamp by default; # strip_offset_headers: true (v0.19.0+) restores records without any of them. include_original_offset_header: true # Dry run first to validate dry_run: false ``` Always validate a restore configuration before executing: ``` kafka-backup validate-restore --config restore.yaml ``` ``` Restore Validation Report ════════════════════════════════════════════════════════════ Status: ✓ VALID Backup: ID: production-backup-001 Source: production-cluster Target Cluster: Brokers: dr-broker-1.kafka.svc:9092, dr-broker-2.kafka.svc:9092 Connection: ✓ OK Topics to Restore: orders → orders (6 partitions) payments → payments (3 partitions) events-* → events-* (24 partitions) Data: Segments: 156 Records: 2,456,789 Estimated Size: 1.2 GB (uncompressed) Warnings: - Topic 'orders' exists on target with 6 partitions (will append data) ``` ``` kafka-backup restore --config restore.yaml ``` ``` [2024-12-03T11:00:00Z INFO] Starting restore from: production-backup-001 [2024-12-03T11:00:01Z INFO] Connected to target cluster [2024-12-03T11:00:02Z INFO] Restoring 4 topics, 33 partitions [2024-12-03T11:00:02Z INFO] Topic: orders Partition 0: Restoring 150,234 records... Partition 0: ✓ Complete ... [2024-12-03T11:08:45Z INFO] Restore completed successfully [2024-12-03T11:08:45Z INFO] Summary: Records Restored: 2,456,789 Duration: 8m 43s Throughput: 4,698 records/sec ``` After restore, consumer groups need their offsets updated. There are several strategies: ``` # View offset mapping from backup kafka-backup show-offset-mapping \ --path /var/lib/kafka-backup/data \ --backup-id production-backup-001 \ --format text ``` ``` Offset Mapping: production-backup-001 ───────────────────────────────────────────────────────────── Topic Partition Source Start Source End Records ───────────────────────────────────────────────────────────── orders 0 0 150233 150234 orders 1 0 148891 148892 payments 0 0 78234 78235 ... To reset consumer groups, use: kafka-consumer-groups --bootstrap-server \ --group \ --topic orders:0 \ --reset-offsets --to-offset 150233 --execute ``` ``` # Generate a reset plan kafka-backup offset-reset plan \ --path /var/lib/kafka-backup/data \ --backup-id production-backup-001 \ --groups my-consumer-group,analytics-group \ --bootstrap-servers dr-broker-1:9092 # Execute the reset kafka-backup offset-reset execute \ --path /var/lib/kafka-backup/data \ --backup-id production-backup-001 \ --groups my-consumer-group \ --bootstrap-servers dr-broker-1:9092 ``` For disaster recovery, use the three-phase restore which handles everything: dr-restore.yaml ``` mode: restore backup_id: "production-backup-001" target: bootstrap_servers: - dr-broker-1:9092 storage: backend: filesystem path: "/var/lib/kafka-backup/data" restore: consumer_group_strategy: header-based reset_consumer_offsets: true consumer_groups: - my-consumer-group - analytics-group ``` ``` kafka-backup three-phase-restore --config dr-restore.yaml ``` You've learned how to: 1. **Create a backup configuration** with topic selection, compression, and checkpointing 2. **Execute and monitor a backup** operation 3. **Verify backup integrity** with validation commands 4. **Plan and execute a restore** with optional PITR 5. **Handle consumer group offsets** after restore - **[Backup to S3](https://kafkabackup.com/guides/backup-to-s3.md)** - Store backups in cloud storage - **[Point-in-Time Recovery](https://kafkabackup.com/guides/restore-pitr.md)** - Restore to specific timestamps - **[Offset Management](https://kafkabackup.com/guides/offset-management.md)** - Advanced consumer offset handling - **[Performance Tuning](https://kafkabackup.com/guides/performance-tuning.md)** - Optimize throughput - **[Kubernetes Operator](https://kafkabackup.com/operator.md)** - Automated scheduled backups --- title: Disaster Recovery description: Use OSO Kafka Backup for zero-downtime disaster recovery source_url: html: https://kafkabackup.com/use-cases/disaster-recovery md: https://kafkabackup.com/use-cases/disaster-recovery.md --- # Disaster Recovery Implement robust disaster recovery for your Kafka infrastructure with OSO Kafka Backup. Kafka replication (MirrorMaker, Confluent Replicator) provides high availability but has limitations: | Challenge | Impact | | --- | --- | | **Active-active complexity** | Dual writes, conflict resolution | | **Data corruption propagates** | Replication copies bad data too | | **No point-in-time recovery** | Can't go back to "before the incident" | | **Topic deletion is permanent** | Accidentally deleted topics can't be recovered | | **High cross-region costs** | Continuous replication is expensive | OSO Kafka Backup provides true disaster recovery: - **Point-in-time recovery**: Restore to any moment - **Isolated backups**: Corruption doesn't propagate - **Cost-effective**: Backup storage vs. active replication - **Flexible recovery**: Full or partial restore ``` Production Region DR Region ┌─────────────────┐ ┌─────────────────┐ │ │ │ │ │ Kafka Cluster │───Backup───────▶│ S3 Bucket │ │ │ │ │ │ ┌───────────┐ │ │ ┌───────────┐ │ │ │ Topic A │ │ │ │ Backups │ │ │ │ Topic B │ │ │ └───────────┘ │ │ │ Topic C │ │ │ │ │ └───────────┘ │ └────────┬────────┘ │ │ │ └─────────────────┘ │ │ Restore ▼ ┌─────────────────┐ │ DR Kafka │ │ Cluster │ │ ┌───────────┐ │ │ │ Topic A │ │ │ │ Topic B │ │ │ │ Topic C │ │ │ └───────────┘ │ └─────────────────┘ ``` **Situation**: Entire production region becomes unavailable. **Recovery**: 1. Activate DR Kafka cluster 2. Restore from latest backup 3. Reset consumer offsets 4. Redirect applications to DR cluster **RTO**: 30 minutes - 2 hours (depending on data volume) **RPO**: Last backup (typically 15 minutes - 1 hour) **Situation**: Bad data published to topics, affecting consumers. **Recovery**: 1. Identify corruption timestamp 2. PITR restore to just before corruption 3. Continue from clean state ``` restore: time_window_end: 1701234500000 # Just before corruption ``` **Situation**: Critical topic accidentally deleted. **Recovery**: 1. Identify which backup contains the topic 2. Restore specific topic 3. Resume operations ``` restore: topics: - accidentally-deleted-topic ``` **Situation**: Kafka cluster compromised, data encrypted. **Recovery**: 1. Provision new cluster (isolated) 2. Restore from clean backup 3. Validate data integrity 4. Cut over to new cluster dr-backup.yaml ``` mode: backup backup_id: "production-${TIMESTAMP}" source: bootstrap_servers: - kafka-prod-1:9092 - kafka-prod-2:9092 - kafka-prod-3:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA256 sasl_username: backup-service sasl_password: ${KAFKA_PASSWORD} topics: include: - "*" exclude: - "__consumer_offsets" - "_schemas" storage: backend: s3 bucket: kafka-dr-backups region: us-east-1 # DR region prefix: production/hourly backup: compression: zstd compression_level: 3 checkpoint_interval_secs: 30 include_offset_headers: true source_cluster_id: "prod-us-west-2" ``` **Kubernetes Operator**: ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: dr-backup spec: schedule: "0 * * * *" # Hourly kafkaCluster: bootstrapServers: - kafka-prod-1:9092 topics: - "*" storage: storageType: s3 s3: bucket: kafka-dr-backups region: us-east-1 ``` ``` #!/bin/bash # dr-restore.sh - Disaster Recovery Runbook BACKUP_ID="${1:-latest}" DR_CLUSTER="kafka-dr-1:9092,kafka-dr-2:9092,kafka-dr-3:9092" echo "=== Kafka Disaster Recovery ===" echo "Backup ID: $BACKUP_ID" echo "Target: $DR_CLUSTER" # 1. Validate backup echo "Step 1: Validating backup..." kafka-backup validate \ --path s3://kafka-dr-backups/production/hourly \ --backup-id "$BACKUP_ID" \ --deep # 2. Validate restore configuration echo "Step 2: Validating restore config..." kafka-backup validate-restore --config dr-restore.yaml # 3. Execute restore echo "Step 3: Executing restore..." kafka-backup three-phase-restore --config dr-restore.yaml # 4. Verify echo "Step 4: Verifying restore..." kafka-topics --bootstrap-server "$DR_CLUSTER" --list echo "=== DR Complete ===" ``` Create a DR test schedule: | Test Type | Frequency | Description | | --- | --- | --- | | Backup validation | Daily | Verify backup integrity | | Restore test | Weekly | Restore to test cluster | | Full DR drill | Quarterly | Complete failover simulation | Time to restore service: | Data Volume | RTO Estimate | | --- | --- | | 1 TB | 2-4 hours | **Factors affecting RTO**: - Network bandwidth to DR region - Target cluster capacity - Number of topics/partitions - Consumer offset reset time Maximum acceptable data loss: | Backup Frequency | RPO | | --- | --- | | Continuous | ~1 minute | | Every 15 minutes | 15 minutes | | Hourly | 1 hour | | Daily | 24 hours | **Choose based on**: - Data criticality - Backup costs - Compliance requirements | Approach | Monthly Cost (100 GB/day) | | --- | --- | | Active-Active (MirrorMaker) | $$$$ (2x infrastructure + transfer) | | Cross-region replication | $$$ (transfer costs) | | Hourly backup to S3 | $ (storage + occasional transfer) | | Daily backup to S3 | $ (storage + rare transfer) | ``` Monthly storage = daily_backup_size × retention_days × storage_cost Example: 10 GB/day × 30 days × $0.023/GB = $6.90/month (S3 Standard) With 4x compression: $1.73/month With Glacier after 30 days: Even less ``` 1. **Multiple backup frequencies** - Hourly for critical data - Daily for full backups - Weekly for long-term retention 2. **Cross-region storage** - Store backups in DR region - Consider multi-region buckets 3. **Encryption** - Enable storage encryption - Use customer-managed keys 1. **Pre-provision DR cluster** - Keep cluster running (minimal) - Or use auto-scaling on demand 2. **Automate everything** - Scripted restore process - Infrastructure as code 3. **Regular testing** - Monthly restore tests - Quarterly full DR drills After restore, consumers need attention: ``` kafka-consumer-groups \ --bootstrap-server kafka-dr:9092 \ --group my-consumer \ --reset-offsets \ --to-earliest \ --execute ``` Use three-phase restore with offset reset: ``` restore: consumer_group_strategy: header-based reset_consumer_offsets: true consumer_groups: - order-processor - payment-service ``` ``` kafka-consumer-groups \ --bootstrap-server kafka-dr:9092 \ --group my-consumer \ --reset-offsets \ --to-datetime 2024-12-01T10:00:00.000 \ --execute ``` - [Point-in-Time Recovery](https://kafkabackup.com/guides/restore-pitr.md) - PITR implementation - [Offset Management](https://kafkabackup.com/guides/offset-management.md) - Consumer recovery - [Kubernetes Operator](https://kafkabackup.com/operator.md) - Automated DR backups --- title: Compliance & Audit description: Use OSO Kafka Backup for regulatory compliance and audit requirements source_url: html: https://kafkabackup.com/use-cases/compliance-audit md: https://kafkabackup.com/use-cases/compliance-audit.md --- # Compliance & Audit Meet regulatory requirements and audit needs with OSO Kafka Backup. Organizations face various compliance requirements: | Regulation | Requirement | | --- | --- | | **GDPR** | Data retention, right to erasure, audit trails | | **SOX** | Financial data integrity, audit trails | | **HIPAA** | Healthcare data protection, 6-year retention | | **PCI-DSS** | Payment data security, 1-year retention | | **SOC 2** | Data availability, integrity, security | | **MiFID II** | 5-7 year retention for financial communications | Maintain historical data for required periods: ``` # Long-term backup configuration backup: compression: zstd compression_level: 9 # Maximum compression for archival storage: backend: s3 bucket: compliance-archives prefix: kafka/2024 ``` With S3 lifecycle policies: ``` { "Rules": [ { "ID": "compliance-retention", "Status": "Enabled", "Transitions": [ { "Days": 90, "StorageClass": "GLACIER" }, { "Days": 365, "StorageClass": "DEEP_ARCHIVE" } ], "Expiration": { "Days": 2555 } // 7 years } ] } ``` Create immutable audit trails: ``` # Audit-specific backup mode: backup backup_id: "audit-${DATE}" source: topics: include: - audit-events - user-actions - transactions backup: include_offset_headers: true source_cluster_id: "production" ``` Retrieve data as it existed at any specific moment: ``` restore: time_window_start: 1701388800000 # Investigation start time_window_end: 1701475200000 # Investigation end ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: compliance-backup spec: schedule: "0 0 * * *" # Daily at midnight kafkaCluster: bootstrapServers: - kafka:9092 topics: - "transactions" - "audit-log" - "user-events" storage: storageType: s3 s3: bucket: compliance-backups region: us-east-1 prefix: production/daily compression: zstd compressionLevel: 9 ``` Use S3 Object Lock for immutable backups: ``` # Enable Object Lock on bucket aws s3api put-object-lock-configuration \ --bucket compliance-backups \ --object-lock-configuration '{ "ObjectLockEnabled": "Enabled", "Rule": { "DefaultRetention": { "Mode": "COMPLIANCE", "Years": 7 } } }' ``` Store backups in multiple regions for redundancy: ``` # Primary backup storage: backend: s3 bucket: compliance-us-east-1 region: us-east-1 --- # Cross-region replication in S3 # Automatically replicates to compliance-eu-west-1 ``` **Request**: "Provide all transactions for customer X during Q3 2024" **Solution**: ``` # 1. Identify relevant backups kafka-backup list --path s3://compliance-backups/production/daily # 2. PITR restore for Q3 kafka-backup restore --config audit-restore.yaml ``` audit-restore.yaml ``` mode: restore backup_id: "compliance-q3-2024" restore: time_window_start: 1719792000000 # Jul 1, 2024 time_window_end: 1727740800000 # Oct 1, 2024 topic_mapping: transactions: audit-investigation-transactions target: bootstrap_servers: - audit-kafka:9092 ``` ``` # 3. Query restored data kafka-console-consumer \ --bootstrap-server audit-kafka:9092 \ --topic audit-investigation-transactions \ --from-beginning \ | grep "customer-id:X" ``` **Request**: "What data was visible to user Y on December 1st?" **Solution**: ``` restore: time_window_start: 1701388800000 # Dec 1, 00:00 time_window_end: 1701475199000 # Dec 1, 23:59 topics: - user-views - access-log ``` **Request**: "Preserve all communications for legal hold" **Solution**: ``` # Create immutable backup kafka-backup backup --config legal-hold.yaml # Apply S3 Object Lock aws s3api put-object-retention \ --bucket compliance-backups \ --key legal-hold-case-123/manifest.json \ --retention '{"Mode": "COMPLIANCE", "RetainUntilDate": "2030-01-01"}' ``` Regularly validate backup integrity: ``` # Weekly deep validation kafka-backup validate \ --path s3://compliance-backups/production/daily \ --backup-id compliance-daily-20241201 \ --deep ``` Document backup chain: ``` { "backup_id": "compliance-daily-20241201", "created_at": "2024-12-01T00:15:32Z", "created_by": "kafka-backup-operator", "source_cluster": "production-us-east-1", "checksum": "sha256:abc123...", "storage_location": "s3://compliance-backups/production/daily/", "encryption": "AES-256 (SSE-KMS)", "retention_policy": "7-years-compliance" } ``` ``` # Get backup manifest with checksums kafka-backup describe \ --path s3://compliance-backups/production/daily \ --backup-id compliance-daily-20241201 \ --format json | jq '.checksums' ``` ``` storage: backend: s3 bucket: compliance-backups # S3 SSE-KMS encryption # Configure in bucket settings ``` Enable encryption: ``` aws s3api put-bucket-encryption \ --bucket compliance-backups \ --server-side-encryption-configuration '{ "Rules": [{ "ApplyServerSideEncryptionByDefault": { "SSEAlgorithm": "aws:kms", "KMSMasterKeyID": "arn:aws:kms:us-east-1:123456789:key/compliance-key" } }] }' ``` ``` source: security: security_protocol: SASL_SSL ssl_ca_location: /certs/ca.crt storage: backend: s3 # All S3 traffic is HTTPS by default ``` ``` { "Version": "2012-10-17", "Statement": [ { "Sid": "ComplianceBackupWrite", "Effect": "Allow", "Action": [ "s3:PutObject", "s3:PutObjectTagging" ], "Resource": "arn:aws:s3:::compliance-backups/*", "Condition": { "StringEquals": { "s3:x-amz-server-side-encryption": "aws:kms" } } }, { "Sid": "ComplianceBackupRead", "Effect": "Allow", "Action": [ "s3:GetObject", "s3:GetObjectTagging" ], "Resource": "arn:aws:s3:::compliance-backups/*", "Condition": { "StringEquals": { "aws:PrincipalTag/Department": "Compliance" } } } ] } ``` Enable S3 access logging: ``` aws s3api put-bucket-logging \ --bucket compliance-backups \ --bucket-logging-status '{ "LoggingEnabled": { "TargetBucket": "compliance-access-logs", "TargetPrefix": "s3-access-logs/" } }' ``` | Data Type | Hot (S3 Standard) | Warm (S3-IA) | Cold (Glacier) | Archive | | --- | --- | --- | --- | --- | | Audit logs | 30 days | 90 days | 1 year | 7 years | | Transactions | 90 days | 1 year | 3 years | 7 years | | User data | 30 days | \- | On deletion | Per GDPR | For non-compliance data, implement cleanup: ``` # Lifecycle policy for normal backups { "Rules": [ { "ID": "cleanup-old-backups", "Status": "Enabled", "Filter": { "Prefix": "production/hourly/" }, "Expiration": { "Days": 30 } } ] } ``` Generate compliance reports: ``` #!/bin/bash # Monthly compliance report echo "=== Kafka Backup Compliance Report ===" echo "Period: $(date -d 'last month' +%Y-%m)" echo "" # List all backup IDs (the command's output is human-readable) kafka-backup list --path s3://compliance-backups/production/daily # Validate the backup IDs selected for this report, one per line while IFS= read -r backup; do [ -n "$backup" ] || continue echo "Validating: $backup" kafka-backup validate --path s3://compliance-backups/production/daily --backup-id "$backup" done < backup-ids.txt ``` Track compliance metrics: ``` # Backup success rate sum(increase(kafka_backup_operator_backups_total{outcome="success"}[30d])) / sum(increase(kafka_backup_operator_backups_total[30d])) # Last completed backup size, by resource kafka_backup_operator_backup_size_bytes ``` 1. **Automate everything** - Manual processes fail audits 2. **Test restores regularly** - Prove data is recoverable 3. **Document retention policies** - Clear policy documentation 4. **Use immutable storage** - Prevent tampering 5. **Enable comprehensive logging** - Audit all access 6. **Encrypt everywhere** - At rest and in transit 7. **Separate compliance data** - Dedicated buckets/accounts OSO Kafka Backup includes automated backup validation with compliance evidence generation. Instead of manually running `kafka-backup validate` and documenting the results, the validation suite: - Runs configurable checks (message counts, offset ranges, consumer groups) against a restored cluster - Generates signed JSON and PDF evidence reports with SHA-256 checksums - Maps validation results to compliance frameworks (SOX ITGC, CMMC RE.3.139, GDPR Article 32) - Sends Slack/PagerDuty notifications on pass or failure - Stores evidence in object storage with configurable retention (default: 7 years for SOX) ``` # Run validation and generate signed evidence $ kafka-backup validation run --config validation.yaml --triggered-by "weekly-sox-check" ``` See the [Compliance Evidence Generation](https://kafkabackup.com/use-cases/compliance-evidence.md) use case and the [Backup Validation Guide](https://kafkabackup.com/guides/validation-compliance.md) for full details. - [Compliance Evidence Generation](https://kafkabackup.com/use-cases/compliance-evidence.md) - Automated validation and signed evidence reports - [Backup Validation Guide](https://kafkabackup.com/guides/validation-compliance.md) - Step-by-step validation setup - [SOX Compliance Example](https://kafkabackup.com/examples/compliance-evidence-sox.md) - End-to-end SOX scenario - [Security Setup](https://kafkabackup.com/guides/security-setup.md) - Encryption configuration - [Enterprise Features](https://kafkabackup.com/enterprise.md) - Audit logging, RBAC - [Kubernetes Operator](https://kafkabackup.com/operator.md) - Automated compliance backups --- title: Compliance Evidence Generation description: Close the gap between paper compliance and operational reality for Kafka backups source_url: html: https://kafkabackup.com/use-cases/compliance-evidence md: https://kafkabackup.com/use-cases/compliance-evidence.md --- # Compliance Evidence Generation Replace screenshots in Word documents with cryptographically signed, machine-readable evidence that auditors can verify independently. Auditors are no longer accepting assertions. They want proof. | What auditors ask for | What most teams provide | | --- | --- | | Evidence of tested backups | A bucket listing in a PDF | | Restore test results | "We tested it last quarter" | | Tamper-proof documentation | Screenshots in a Word document | | Repeatable validation | Manual, ad-hoc procedures | | Signed evidence chain | Email attachments | Approximately 40% of organisations now receive requests for live backup restoration evidence during audits. The trend is accelerating under GDPR, SOX, and CMMC. 1. **Proof that backups exist** and contain the expected data 2. **Proof that backups can be restored** — not just that they were taken 3. **Tamper-evident documentation** — checksums, signatures, chain of custody 4. **Repeatable process** — the same validation produces the same evidence 5. **Framework-specific controls** — which check satisfies which requirement IT General Controls require backup integrity monitoring, restore testing, and 7-year evidence retention. The validation suite satisfies these with SHA-256 checksums (accepted by SOX auditors as tamper-evident documentation) and configurable retention. Requires organisations handling CUI to "regularly perform and test data back-ups" with documented evidence. Every validation check maps directly to RE.3.139. Requires "regularly testing, assessing, and evaluating the effectiveness of technical and organisational measures" for data protection. The evidence report demonstrates restore capability with performance metrics (RTO). - **NIS2** — regular recovery tests with documented results - **ISO 27001** — documented RTOs/RPOs and recovery test evidence - **Cyber Insurance** — monthly compliance reports on backup health ``` Backup (existing) Restore (existing) Validate (new) │ │ │ ▼ ▼ ▼ ┌─────────┐ ┌──────────┐ ┌──────────────┐ │ S3/GCS/ │ │ Restored │ │ Validation │ │ Azure │─────────▶│ Kafka │─────────▶│ Runner │ │ Storage │ manifest │ Cluster │ checks │ │ └─────────┘ └──────────┘ └──────┬───────┘ │ ┌─────────────┼─────────────┐ ▼ ▼ ▼ ┌──────────┐ ┌─────────┐ ┌──────────┐ │ JSON │ │ PDF │ │ .sig │ │ Report │ │ Report │ │Signature │ └──────────┘ └─────────┘ └──────────┘ │ ┌─────┴──────┐ ▼ ▼ ┌────────┐ ┌──────────┐ │ Slack │ │PagerDuty │ └────────┘ └──────────┘ ``` Run every Sunday at 02:00 UTC via cron: ``` # /etc/cron.d/kafka-backup-validation 0 2 * * 0 kafka-backup validation run --config /etc/kafka-backup/sox-validation.yaml ``` The evidence report is automatically uploaded to S3 with 7-year retention. Your SOX auditor gets a URL to 52 weeks of uninterrupted evidence. When an auditor asks "prove you can restore data from March 15th": ``` $ kafka-backup validation run \ --config validation.yaml \ --pitr 1710460800000 \ --triggered-by "External auditor KPMG - Q1 2026 review" ``` The `--triggered-by` string appears in the evidence report, establishing chain of custody. ``` notifications: pagerduty: integration_key: "your-key" severity: critical ``` On failure, the SRE on-call receives a PagerDuty alert with the evidence report URL, enabling investigation before the next audit. | Feature | OSO Kafka Backup | Kannika Armory | Confluent Platform | Velero | | --- | --- | --- | --- | --- | | Automated restore-and-validate | Yes | Manual kubectl steps | No | Manual | | Machine-readable evidence (JSON) | Yes, signed | No | No | No | | Auditor-ready PDF report | Yes, branded | No | No | No | | Cryptographic signing | ECDSA-P256-SHA256 | No | No | No | | SOX/CMMC/GDPR mapping | Explicit in report | No | No | No | | Prometheus validation metrics | 8 dedicated metrics | Backup progress only | Partial | Partial | | Slack/PagerDuty on result | Yes | No | No | No | | Evidence retention (7yr default) | Yes, configurable | No | No | No | | OSS / free | MIT | Commercial | Commercial | Apache 2.0 | - [Backup Validation Guide](https://kafkabackup.com/guides/validation-compliance.md) — step-by-step setup - [Evidence Signing Guide](https://kafkabackup.com/guides/evidence-signing.md) — key management deep-dive - [SOX Compliance Example](https://kafkabackup.com/examples/compliance-evidence-sox.md) — complete SOX scenario - [GDPR Compliance Example](https://kafkabackup.com/examples/compliance-evidence-gdpr.md) — GDPR Article 32 scenario - [Evidence Report Schema](https://kafkabackup.com/reference/evidence-report-schema.md) — JSON schema reference --- title: Migration description: Use OSO Kafka Backup for Kafka cluster migrations source_url: html: https://kafkabackup.com/use-cases/migration md: https://kafkabackup.com/use-cases/migration.md --- # Migration Migrate Kafka data between clusters, versions, or cloud providers using OSO Kafka Backup. | Scenario | Example | | --- | --- | | **Version upgrade** | Kafka 2.x → Kafka 3.x | | **Cloud migration** | On-premises → AWS MSK | | **Provider switch** | AWS MSK → Confluent Cloud | | **Region migration** | us-east-1 → eu-west-1 | | **Environment cloning** | Production → Staging | | **Cluster consolidation** | Multiple clusters → One | | **ZooKeeper to KRaft** | MSK ZK-mode → MSK KRaft-mode | > [!TIP] > > [!NOTE] > > Enterprise feature — plan and precheck are free > > [!NOTE] > > Apache Kafka 4.0 removes ZooKeeper. AWS MSK requires a new KRaft cluster — there is no in-place upgrade. kafka-backup Enterprise handles the full migration: topic replication, ACL migration, consumer group offset translation, and cryptographic evidence. > > > > **Try the free tools now** — `plan` generates a migration runbook, cost estimate, and IAM policies. `precheck` verifies your clusters are ready. No license needed. > > > > [Full MSK KRaft Migration Guide →](https://kafkabackup.com/enterprise/msk-kraft-migration.md) | Feature | MirrorMaker | Kafka Backup | | --- | --- | --- | | Live sync | Yes | No | | PITR | No | Yes | | Data transformation | Limited | Plugin support | | Topic remapping | Yes | Yes | | Offset preservation | Complex | Built-in | | Network requirement | Continuous | One-time | **Choose Kafka Backup when**: - Clusters can't communicate directly - You need data transformation during migration - You want to select specific time ranges - Minimizing cross-region bandwidth is important ``` # List topics kafka-topics --bootstrap-server source-kafka:9092 --list # Get topic details kafka-topics --bootstrap-server source-kafka:9092 --describe # Check data volume kafka-log-dirs --bootstrap-server source-kafka:9092 --describe ``` **Consider**: - Total data volume - Number of topics and partitions - Consumer groups to migrate - Acceptable downtime - Topic naming changes migration-backup.yaml ``` mode: backup backup_id: "migration-$(date +%Y%m%d)" source: bootstrap_servers: - source-kafka-1:9092 - source-kafka-2:9092 topics: include: - "*" exclude: - "__consumer_offsets" - "_schemas" storage: backend: s3 bucket: kafka-migration region: us-west-2 prefix: migration/source-cluster backup: compression: zstd include_offset_headers: true source_cluster_id: "source-cluster" ``` ``` kafka-backup backup --config migration-backup.yaml ``` ``` # List backup kafka-backup list --path s3://kafka-migration/migration/source-cluster # Describe backup kafka-backup describe \ --path s3://kafka-migration/migration/source-cluster \ --backup-id "migration-20241203" # Deep validation kafka-backup validate \ --path s3://kafka-migration/migration/source-cluster \ --backup-id "migration-20241203" \ --deep ``` migration-restore.yaml ``` mode: restore backup_id: "migration-20241203" target: bootstrap_servers: - target-kafka-1:9092 - target-kafka-2:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA256 sasl_username: admin sasl_password: ${KAFKA_PASSWORD} storage: backend: s3 bucket: kafka-migration region: us-west-2 prefix: migration/source-cluster restore: # Optional: Remap topic names topic_mapping: old-topic-name: new-topic-name # Optional: Change partition count # (must be >= source partitions) include_original_offset_header: true consumer_group_strategy: skip ``` ``` # Validate first kafka-backup validate-restore --config migration-restore.yaml # Execute restore kafka-backup restore --config migration-restore.yaml ``` ``` # Option 1: Reset to beginning (reprocess all) kafka-consumer-groups \ --bootstrap-server target-kafka:9092 \ --group my-consumer \ --reset-offsets \ --to-earliest \ --all-topics \ --execute # Option 2: Use offset mapping from backup kafka-backup offset-reset execute \ --path s3://kafka-migration/migration/source-cluster \ --backup-id "migration-20241203" \ --groups my-consumer \ --bootstrap-servers target-kafka:9092 # Option 3: Three-phase restore (automated) kafka-backup three-phase-restore --config migration-restore.yaml ``` ``` # Compare topic counts echo "Source topics:" kafka-topics --bootstrap-server source-kafka:9092 --list | wc -l echo "Target topics:" kafka-topics --bootstrap-server target-kafka:9092 --list | wc -l # Verify record counts kafka-backup describe \ --path s3://kafka-migration/migration/source-cluster \ --backup-id "migration-20241203" \ --format json | jq '.statistics.records' # Sample data verification kafka-console-consumer \ --bootstrap-server target-kafka:9092 \ --topic my-topic \ --from-beginning \ --max-messages 10 ``` 1. **Stop producers** to source cluster 2. **Final backup** (capture last changes) 3. **Final restore** to target 4. **Update consumer offsets** on target 5. **Redirect applications** to target cluster 6. **Monitor** for issues 7. **Decommission** source cluster Clone production data to dev/staging with data masking: ``` mode: restore backup_id: "production-backup" target: bootstrap_servers: - dev-kafka:9092 restore: topic_mapping: orders: dev-orders users: dev-users # Enterprise: Data masking via plugins # plugins: # - name: pii-masker # config: # fields_to_mask: # - email # - ssn ``` Migrate from multiple source clusters: ``` # Backup cluster A kafka-backup backup --config cluster-a-backup.yaml # Backup cluster B kafka-backup backup --config cluster-b-backup.yaml # Restore both to target (with prefix) kafka-backup restore --config cluster-a-restore.yaml kafka-backup restore --config cluster-b-restore.yaml ``` cluster-a-restore.yaml ``` restore: topic_mapping: orders: cluster-a/orders payments: cluster-a/payments ``` AWS MSK → Confluent Cloud: msk-backup.yaml ``` source: bootstrap_servers: - b-1.msk-cluster.abc123.kafka.us-east-1.amazonaws.com:9092 security: security_protocol: SASL_SSL sasl_mechanism: AWS_MSK_IAM # Uses AWS credentials storage: backend: s3 bucket: migration-data region: us-east-1 ``` confluent-restore.yaml ``` target: bootstrap_servers: - pkc-xxxxx.us-west-2.aws.confluent.cloud:9092 security: security_protocol: SASL_SSL sasl_mechanism: PLAIN sasl_username: ${CONFLUENT_API_KEY} sasl_password: ${CONFLUENT_API_SECRET} storage: backend: s3 bucket: migration-data region: us-east-1 ``` Migrate only specific topics: ``` source: topics: include: - orders - payments - "events-*" exclude: - events-debug ``` Change partition count during migration: 1. Backup with original partitions 2. Create topics on target with desired partitions 3. Restore (data redistributes across partitions) ``` # Create topic with more partitions kafka-topics --bootstrap-server target-kafka:9092 \ --create --topic orders \ --partitions 12 \ --replication-factor 3 # Restore will use existing topic configuration kafka-backup restore --config restore.yaml ``` 1. Set up continuous backup from source 2. Start consumers on target (dual-read) 3. Gradually migrate producers 4. Disable source consumers 5. Stop source backup 1. Initial full backup/restore 2. Delta backup during cutover window 3. Delta restore 4. Switch applications 1. Stop all applications 2. Full backup 3. Full restore 4. Restart applications pointing to target Always have a rollback strategy: rollback-backup.yaml ``` # Before migration, backup target cluster state mode: backup backup_id: "pre-migration-rollback" source: bootstrap_servers: - target-kafka:9092 topics: include: - "*" storage: backend: s3 bucket: kafka-migration prefix: rollback ``` 1. **Test with subset first** - Migrate few topics as pilot 2. **Validate at each step** - Don't proceed with errors 3. **Document everything** - Record all steps taken 4. **Have rollback ready** - Plan for failure 5. **Monitor closely** - Watch for data discrepancies 6. **Communicate cutover** - Coordinate with stakeholders - [Disaster Recovery](https://kafkabackup.com/use-cases/disaster-recovery.md) - DR planning - [Offset Management](https://kafkabackup.com/guides/offset-management.md) - Consumer migration - [Performance Tuning](https://kafkabackup.com/guides/performance-tuning.md) - Optimize migration speed --- title: Feature Comparison description: Compare OSO Kafka Backup with MirrorMaker, Confluent Replicator, and other solutions source_url: html: https://kafkabackup.com/use-cases/vs-alternatives md: https://kafkabackup.com/use-cases/vs-alternatives.md --- # Feature Comparison Compare OSO Kafka Backup with alternative approaches for Kafka data protection. | Solution | Type | Best For | | --- | --- | --- | | **OSO Kafka Backup** | Backup/Restore | DR, compliance, migration | | **MirrorMaker 2** | Replication | Active-active, geo-distribution | | **Confluent Replicator** | Replication | Enterprise replication | | **Tiered Storage** | Offload | Cost reduction, infinite retention | | **Custom Scripts** | DIY | Simple use cases | | Feature | OSO Kafka Backup | MirrorMaker 2 | | --- | --- | --- | | **Purpose** | Backup/restore | Real-time replication | | **Point-in-time recovery** | Yes | No | | **Data transformation** | Yes (plugins) | Limited | | **Storage backends** | S3, Azure, GCS, local | Kafka only | | **Offset preservation** | Built-in | Requires configuration | | **Network requirement** | One-time transfer | Continuous connection | | **Compression** | Zstd, LZ4 | Kafka default | | **Cost** | Storage costs | Double infrastructure | | **Recovery time** | Minutes-hours | Instant (already replicated) | | **Data loss (RPO)** | Last backup | Near-zero | **Choose OSO Kafka Backup when:** - You need point-in-time recovery - Clusters can't communicate directly - Cost optimization is important - Compliance requires immutable backups **Choose MirrorMaker 2 when:** - Near-zero RPO is required - Active-active is needed - Real-time geo-distribution - Network allows continuous sync | Feature | OSO Kafka Backup | Confluent Replicator | | --- | --- | --- | | **Licensing** | Open source (MIT) | Commercial | | **Point-in-time recovery** | Yes | No | | **Schema Registry sync** | Enterprise | Yes | | **Offset sync** | Built-in | Yes | | **Data transformation** | Plugins | SMT (Connect) | | **Monitoring** | Prometheus | Control Center | | **Support** | Community/Enterprise | Confluent support | **Choose OSO Kafka Backup when:** - Open source is preferred - PITR is required - Budget constraints exist - Not using Confluent Platform **Choose Confluent Replicator when:** - Already using Confluent Platform - Need Schema Registry sync - Prefer integrated tooling - Have Confluent support contract | Feature | OSO Kafka Backup | Tiered Storage | | --- | --- | --- | | **Purpose** | Backup/restore | Cost reduction | | **Point-in-time recovery** | Yes | Limited | | **Independent from Kafka** | Yes | No (Kafka feature) | | **Cross-cluster restore** | Yes | No | | **Compression** | Additional | Kafka default | | **Availability** | Any Kafka | Kafka 3.0+ / Confluent | | **Broker dependency** | None | Requires running brokers | **Choose OSO Kafka Backup when:** - Cross-cluster recovery needed - Independent disaster recovery - Using older Kafka versions - Compliance requires separate backups **Choose Tiered Storage when:** - Primary goal is cost reduction - Data stays in same cluster - Using compatible Kafka version - Simpler operational model preferred | Feature | OSO Kafka Backup | Custom Scripts | | --- | --- | --- | | **Development effort** | None | High | | **Maintenance** | Vendor managed | Self-maintained | | **Performance** | Optimized (Rust) | Variable | | **Features** | Complete | What you build | | **Reliability** | Production-tested | Depends | | **Offset management** | Built-in | Must implement | | **Cloud storage** | Native support | Must implement | **Choose OSO Kafka Backup when:** - Don't want to build from scratch - Need production-ready solution - Value ongoing development - Time to market matters **Choose Custom Scripts when:** - Very simple requirements - Unique constraints - Learning exercise - Full control required | Feature | OSO Backup | MM2 | Replicator | Tiered | | --- | --- | --- | --- | --- | | Backup to object storage | Yes | No | No | Yes | | Point-in-time recovery | Yes | No | No | Limited | | Cross-cluster restore | Yes | Yes | Yes | No | | Incremental backup | Yes | N/A | N/A | N/A | | Compression | Zstd/LZ4 | Kafka | Kafka | Kafka | | Topic selection | Patterns | Patterns | Patterns | All | | Topic remapping | Yes | Yes | Yes | No | | Feature | OSO Backup | MM2 | Replicator | Tiered | | --- | --- | --- | --- | --- | | Kubernetes operator | Yes | No | Yes | No | | GitOps support | CRDs | No | CRDs | No | | Prometheus metrics | Yes | JMX | JMX | JMX | | CLI tool | Yes | No | No | No | | Scheduled operations | Yes | N/A | N/A | N/A | | Dry-run validation | Yes | No | No | No | | Feature | OSO Backup | MM2 | Replicator | Tiered | | --- | --- | --- | --- | --- | | Confluent CSFLE metadata backup | Enterprise | No | No | No | | Data masking | Enterprise | No | SMT | No | | Audit logging | Enterprise | No | Yes | No | | RBAC | Enterprise | No | Yes | No | | Schema Registry | Enterprise | No | Yes | No | ``` ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ Kafka │───▶│ Backup │───▶│ Object │ │ Cluster │ │ Service │ │ Storage │ └─────────────┘ └─────────────┘ └─────────────┘ │ ▼ ┌─────────────┐ ┌─────────────┐ │ Target │◀───│ Restore │ │ Cluster │ │ Service │ └─────────────┘ └─────────────┘ ``` **Pros:** - Decoupled storage - Independent recovery - Cost-effective long-term storage **Cons:** - Not real-time - Recovery takes time ``` ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ Source │───▶│ MM2 / │───▶│ Target │ │ Cluster │ │ Replicator│ │ Cluster │ └─────────────┘ └─────────────┘ └─────────────┘ ``` **Pros:** - Real-time sync - Active-active possible - Near-zero RPO **Cons:** - No PITR - Continuous infrastructure cost - Network dependency ``` ┌─────────────────────────────────────────┐ │ Kafka Cluster │ │ ┌───────────┐ ┌───────────────┐ │ │ │ Hot │ ───▶ │ Cold │ │ │ │ Tier │ │ Tier (S3) │ │ │ │ (Local) │ │ │ │ │ └───────────┘ └───────────────┘ │ └─────────────────────────────────────────┘ ``` **Pros:** - Transparent to clients - Infinite retention - Cost optimization **Cons:** - No cross-cluster recovery - Broker dependency - Limited PITR | Solution | Infrastructure | Storage | Network | Total | | --- | --- | --- | --- | --- | | **OSO Backup** (hourly) | $0 | ~$70 | ~$10 | ~$80 | | **OSO Backup** (daily) | $0 | ~$70 | ~$1 | ~$71 | | **MM2** | ~$500 | $0 | ~$50 | ~$550 | | **Replicator** | ~$500+ | $0 | ~$50 | ~$550+ license | | **Tiered Storage** | $0 | ~$70 | $0 | ~$70 | _Costs are illustrative and vary by region and provider._ **OSO Kafka Backup**: - One-time transfer costs - Object storage (cheaper than block) - No additional compute during backup window **Replication solutions**: - 2x Kafka infrastructure - Continuous network transfer - Ongoing compute costs If currently using MM2 for backup purposes: 1. Set up OSO Backup alongside MM2 2. Validate backups match replicated data 3. Disable MM2 (keep target cluster temporarily) 4. Rely on OSO Backup for recovery 5. Decommission MM2 target cluster 1. Audit current script capabilities 2. Map to OSO Backup configuration 3. Run parallel backups 4. Validate backup contents match 5. Retire custom scripts - Point-in-time recovery is important - Compliance requires immutable backups - Cost optimization is a priority - Cross-region/cross-cloud recovery needed - Kubernetes-native operations preferred - Open source is preferred - Near-zero RPO is required - Active-active architecture needed - Real-time geo-distribution required - Network allows continuous sync - Immediate failover is critical - Defense in depth required - Different RPO/RTO for different scenarios - Compliance requires multiple protection methods - Active-active + disaster recovery needed - [Getting Started](https://kafkabackup.com/getting-started/quickstart.md) - Try OSO Kafka Backup - [Disaster Recovery](https://kafkabackup.com/use-cases/disaster-recovery.md) - DR planning - [Migration Guide](https://kafkabackup.com/use-cases/migration.md) - Migrate from alternatives --- title: Index source_url: html: https://kafkabackup.com/deployment/index md: https://kafkabackup.com/deployment/index.md --- --- title: Bare Metal Installation description: Install OSO Kafka Backup on bare metal servers and VMs source_url: html: https://kafkabackup.com/deployment/bare-metal md: https://kafkabackup.com/deployment/bare-metal.md --- # Bare Metal Installation Install OSO Kafka Backup directly on Linux servers or virtual machines. - Linux (x86\_64 or ARM64) or macOS - Network access to Kafka brokers - Storage space for backups (local or mounted) > [!TIP] > > [!NOTE] > > Enterprise Edition > > [!NOTE] > > For enterprise features (Schema Registry backup, Confluent RBAC backup), install the enterprise binary instead. Download URLs use `kafka-backup-enterprise` in place of `kafka-backup`: > > > > ``` > > https://github.com/osodevops/kafka-backup-enterprise/releases/latest/download/kafka-backup-x86_64-linux.tar.gz > > ``` > > > > Or install via Homebrew: `brew install osodevops/tap/kafka-backup-enterprise` > > > > See the full [Enterprise Installation Guide](https://kafkabackup.com/enterprise/installation.md). ``` # Download latest release curl -L -o kafka-backup.tar.gz \ https://github.com/osodevops/kafka-backup/releases/latest/download/kafka-backup-linux-amd64.tar.gz # Extract tar xzf kafka-backup.tar.gz # Move to PATH sudo mv kafka-backup /usr/local/bin/ # Verify installation kafka-backup --version ``` ``` curl -L -o kafka-backup.tar.gz \ https://github.com/osodevops/kafka-backup/releases/latest/download/kafka-backup-linux-arm64.tar.gz tar xzf kafka-backup.tar.gz sudo mv kafka-backup /usr/local/bin/ ``` ``` # Intel Mac curl -L -o kafka-backup.tar.gz \ https://github.com/osodevops/kafka-backup/releases/latest/download/kafka-backup-darwin-amd64.tar.gz # Apple Silicon (M1/M2) curl -L -o kafka-backup.tar.gz \ https://github.com/osodevops/kafka-backup/releases/latest/download/kafka-backup-darwin-arm64.tar.gz tar xzf kafka-backup.tar.gz sudo mv kafka-backup /usr/local/bin/ ``` Create directories for configuration and data: ``` # Create directories sudo mkdir -p /etc/kafka-backup sudo mkdir -p /var/lib/kafka-backup/data sudo mkdir -p /var/log/kafka-backup # Set permissions (create dedicated user) sudo useradd -r -s /bin/false kafka-backup sudo chown -R kafka-backup:kafka-backup /var/lib/kafka-backup sudo chown -R kafka-backup:kafka-backup /var/log/kafka-backup ``` Create a backup configuration file: ``` sudo tee /etc/kafka-backup/backup.yaml << 'EOF' mode: backup backup_id: "daily-backup" source: bootstrap_servers: - kafka-1.example.com:9092 - kafka-2.example.com:9092 - kafka-3.example.com:9092 topics: include: - "*" exclude: - "__consumer_offsets" - "_schemas" storage: backend: filesystem path: "/var/lib/kafka-backup/data" backup: compression: zstd compression_level: 3 checkpoint_interval_secs: 30 include_offset_headers: true EOF ``` ``` # Run as the kafka-backup user sudo -u kafka-backup kafka-backup backup --config /etc/kafka-backup/backup.yaml # With verbose logging sudo -u kafka-backup kafka-backup -v backup --config /etc/kafka-backup/backup.yaml ``` Create a systemd service for automated backups: ``` sudo tee /etc/systemd/system/kafka-backup.service << 'EOF' [Unit] Description=OSO Kafka Backup Service After=network.target [Service] Type=oneshot User=kafka-backup Group=kafka-backup ExecStart=/usr/local/bin/kafka-backup backup --config /etc/kafka-backup/backup.yaml StandardOutput=append:/var/log/kafka-backup/backup.log StandardError=append:/var/log/kafka-backup/backup.log # Security hardening NoNewPrivileges=yes ProtectSystem=strict ProtectHome=yes ReadWritePaths=/var/lib/kafka-backup /var/log/kafka-backup PrivateTmp=yes [Install] WantedBy=multi-user.target EOF ``` > [!TIP] > > [!NOTE] > > Graceful Shutdown > > [!NOTE] > > OSO Kafka Backup handles `SIGTERM` and `SIGINT` for graceful shutdown (v0.8.1+). When `systemctl stop kafka-backup` is run, the process flushes in-progress segments and saves a checkpoint before exiting. The default systemd stop timeout (90s) is sufficient for most workloads. Create a timer for scheduled execution: ``` sudo tee /etc/systemd/system/kafka-backup.timer << 'EOF' [Unit] Description=Run Kafka Backup daily at 2 AM [Timer] OnCalendar=*-*-* 02:00:00 Persistent=true RandomizedDelaySec=300 [Install] WantedBy=timers.target EOF ``` Enable and start the timer: ``` sudo systemctl daemon-reload sudo systemctl enable kafka-backup.timer sudo systemctl start kafka-backup.timer # Check timer status sudo systemctl list-timers kafka-backup.timer ``` Alternative to systemd timer: ``` # Edit crontab for kafka-backup user sudo -u kafka-backup crontab -e # Add daily backup at 2 AM 0 2 * * * /usr/local/bin/kafka-backup backup --config /etc/kafka-backup/backup.yaml >> /var/log/kafka-backup/backup.log 2>&1 ``` Configure log rotation: ``` sudo tee /etc/logrotate.d/kafka-backup << 'EOF' /var/log/kafka-backup/*.log { daily rotate 14 compress delaycompress missingok notifempty create 640 kafka-backup kafka-backup } EOF ``` ``` # List backups kafka-backup list --path /var/lib/kafka-backup/data # Describe latest backup kafka-backup describe --path /var/lib/kafka-backup/data --backup-id daily-backup # Validate backup integrity kafka-backup validate --path /var/lib/kafka-backup/data --backup-id daily-backup ``` ``` # Follow backup logs tail -f /var/log/kafka-backup/backup.log # Check for errors grep -i error /var/log/kafka-backup/backup.log ``` ``` # Check backup storage usage du -sh /var/lib/kafka-backup/data/* # Monitor with df df -h /var/lib/kafka-backup ``` For shared or network storage: ``` # Mount NFS volume sudo mount -t nfs nfs-server:/kafka-backups /var/lib/kafka-backup/data # Add to /etc/fstab for persistence echo "nfs-server:/kafka-backups /var/lib/kafka-backup/data nfs defaults 0 0" | sudo tee -a /etc/fstab # Update configuration storage: backend: filesystem path: "/var/lib/kafka-backup/data" ``` For cloud storage: ``` # Set AWS credentials export AWS_ACCESS_KEY_ID="AKIA..." export AWS_SECRET_ACCESS_KEY="..." export AWS_REGION="us-west-2" # Or use instance profile (recommended) # No credentials needed if running on EC2 with IAM role ``` Update configuration: ``` storage: backend: s3 bucket: my-kafka-backups region: us-west-2 prefix: production/daily ``` Create a script to manage backup retention: ``` sudo tee /usr/local/bin/kafka-backup-rotate << 'EOF' #!/bin/bash BACKUP_PATH="/var/lib/kafka-backup/data" RETENTION_DAYS=30 # List backups older than retention period find "$BACKUP_PATH" -maxdepth 1 -type d -mtime +$RETENTION_DAYS -name "backup-*" | while read dir; do echo "Removing old backup: $dir" rm -rf "$dir" done # Log disk usage echo "Current disk usage:" du -sh "$BACKUP_PATH"/* EOF chmod +x /usr/local/bin/kafka-backup-rotate # Add to cron (run daily at 3 AM) echo "0 3 * * * /usr/local/bin/kafka-backup-rotate >> /var/log/kafka-backup/rotation.log 2>&1" | sudo -u kafka-backup crontab - ``` ``` # Check file ownership ls -la /var/lib/kafka-backup/ # Fix permissions sudo chown -R kafka-backup:kafka-backup /var/lib/kafka-backup/ ``` ``` # Test Kafka connectivity nc -zv kafka-1.example.com 9092 # Check firewall sudo iptables -L -n | grep 9092 ``` ``` # Check disk usage df -h /var/lib/kafka-backup # Find large backups du -sh /var/lib/kafka-backup/data/* # Remove old backups kafka-backup list --path /var/lib/kafka-backup/data rm -rf /var/lib/kafka-backup/data/old-backup-001 ``` - [Configuration Reference](https://kafkabackup.com/reference/config-yaml.md) - All configuration options - [Security Setup](https://kafkabackup.com/guides/security-setup.md) - TLS and SASL configuration - [Performance Tuning](https://kafkabackup.com/guides/performance-tuning.md) - Optimize backup speed --- title: Docker Deployment description: Deploy OSO Kafka Backup using Docker and Docker Compose source_url: html: https://kafkabackup.com/deployment/docker md: https://kafkabackup.com/deployment/docker.md --- # Docker Deployment Run OSO Kafka Backup in Docker containers for development, testing, or production use. - Docker 20.10+ - Docker Compose (optional) - Network access to Kafka brokers > [!TIP] > > [!NOTE] > > Enterprise Edition > > [!NOTE] > > For enterprise features (Schema Registry backup, Confluent RBAC backup), use the enterprise image instead: > > > > ``` > > docker pull osodevops/kafka-backup-enterprise:latest > > ``` > > > > The enterprise image is a drop-in replacement — all commands and configuration are identical. See the [Enterprise Installation Guide](https://kafkabackup.com/enterprise/installation.md). ``` docker pull ghcr.io/osodevops/kafka-backup:latest ``` ``` docker run --rm ghcr.io/osodevops/kafka-backup:latest --version docker run --rm ghcr.io/osodevops/kafka-backup:latest --help ``` ``` # Create backup directory mkdir -p ./backups # Create config file cat > backup.yaml << 'EOF' mode: backup backup_id: "my-backup" source: bootstrap_servers: - kafka:9092 topics: include: - my-topic storage: backend: filesystem path: "/data/backups" backup: compression: zstd EOF # Run backup docker run --rm \ --network host \ -v $(pwd)/backup.yaml:/config/backup.yaml:ro \ -v $(pwd)/backups:/data/backups \ ghcr.io/osodevops/kafka-backup:latest \ backup --config /config/backup.yaml ``` ``` docker run --rm \ -e AWS_ACCESS_KEY_ID="${AWS_ACCESS_KEY_ID}" \ -e AWS_SECRET_ACCESS_KEY="${AWS_SECRET_ACCESS_KEY}" \ -e AWS_REGION="us-west-2" \ -v $(pwd)/backup.yaml:/config/backup.yaml:ro \ ghcr.io/osodevops/kafka-backup:latest \ backup --config /config/backup.yaml ``` ``` docker run --rm \ --network host \ -v $(pwd)/restore.yaml:/config/restore.yaml:ro \ -v $(pwd)/backups:/data/backups:ro \ ghcr.io/osodevops/kafka-backup:latest \ restore --config /config/restore.yaml ``` Complete environment with Kafka and backup service: docker-compose.yml ``` version: '3.8' services: kafka: image: confluentinc/cp-kafka:7.5.0 hostname: kafka ports: - "9092:9092" environment: KAFKA_NODE_ID: 1 KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: CONTROLLER:PLAINTEXT,PLAINTEXT:PLAINTEXT KAFKA_LISTENERS: PLAINTEXT://0.0.0.0:9092,CONTROLLER://0.0.0.0:9093 KAFKA_ADVERTISED_LISTENERS: PLAINTEXT://kafka:9092 KAFKA_CONTROLLER_LISTENER_NAMES: CONTROLLER KAFKA_CONTROLLER_QUORUM_VOTERS: 1@kafka:9093 KAFKA_PROCESS_ROLES: broker,controller KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR: 1 CLUSTER_ID: MkU3OEVBNTcwNTJENDM2Qk healthcheck: test: ["CMD", "kafka-topics", "--bootstrap-server", "localhost:9092", "--list"] interval: 10s timeout: 5s retries: 5 kafka-backup: image: ghcr.io/osodevops/kafka-backup:latest depends_on: kafka: condition: service_healthy volumes: - ./config:/config:ro - ./backups:/data/backups command: ["backup", "--config", "/config/backup.yaml"] # For one-shot runs, use: # profiles: ["backup"] volumes: backups: ``` docker-compose.prod.yml ``` version: '3.8' services: kafka-backup: image: ghcr.io/osodevops/kafka-backup:latest restart: unless-stopped environment: - AWS_ACCESS_KEY_ID=${AWS_ACCESS_KEY_ID} - AWS_SECRET_ACCESS_KEY=${AWS_SECRET_ACCESS_KEY} - AWS_REGION=${AWS_REGION} volumes: - ./config:/config:ro - kafka-backup-data:/data networks: - kafka-network deploy: resources: limits: cpus: '2' memory: 1G reservations: cpus: '0.5' memory: 256M healthcheck: test: ["CMD", "kafka-backup", "--help"] interval: 30s timeout: 10s retries: 3 logging: driver: "json-file" options: max-size: "100m" max-file: "5" networks: kafka-network: external: true volumes: kafka-backup-data: ``` OSO Kafka Backup handles `SIGTERM` and `SIGINT` for graceful shutdown (v0.8.1+). When you stop a container, the backup engine flushes in-progress segments and saves a checkpoint before exiting. ``` # Graceful stop (sends SIGTERM, waits for shutdown) docker stop kafka-backup # With custom timeout (default: 10s) docker stop --time 60 kafka-backup ``` For Docker Compose, the default stop timeout is 10 seconds. Increase it for large backups: ``` services: kafka-backup: image: ghcr.io/osodevops/kafka-backup:latest stop_grace_period: 60s ``` Create a wrapper script: run-backup.sh ``` #!/bin/bash docker-compose -f docker-compose.yml run --rm kafka-backup \ backup --config /config/backup.yaml ``` Add to crontab: ``` # Daily backup at 2 AM 0 2 * * * /path/to/run-backup.sh >> /var/log/kafka-backup.log 2>&1 ``` docker-compose.yml ``` version: '3.8' services: ofelia: image: mcuadros/ofelia:latest depends_on: - kafka-backup command: daemon --docker volumes: - /var/run/docker.sock:/var/run/docker.sock:ro labels: ofelia.job-run.kafka-backup.schedule: "0 2 * * *" ofelia.job-run.kafka-backup.container: "kafka-backup" kafka-backup: image: ghcr.io/osodevops/kafka-backup:latest container_name: kafka-backup volumes: - ./config:/config:ro - ./backups:/data/backups command: ["backup", "--config", "/config/backup.yaml"] labels: ofelia.enabled: "true" ``` | Variable | Description | | --- | --- | | `AWS_ACCESS_KEY_ID` | AWS access key for S3 | | `AWS_SECRET_ACCESS_KEY` | AWS secret key for S3 | | `AWS_REGION` | AWS region | | `AZURE_STORAGE_ACCOUNT` | Azure storage account | | `AZURE_STORAGE_KEY` | Azure storage key | | `GOOGLE_APPLICATION_CREDENTIALS` | Path to GCP service account JSON | | `RUST_LOG` | Logging level (info, debug, trace) | | Container Path | Purpose | Mode | | --- | --- | --- | | `/config` | Configuration files | Read-only | | `/data/backups` | Backup storage | Read-write | | `/certs` | TLS certificates | Read-only | The image runs as non-root by default (UID 1000): ``` services: kafka-backup: image: ghcr.io/osodevops/kafka-backup:latest user: "1000:1000" ``` The `/tmp` tmpfs mount is **required** — the backup engine stores its offset tracking database there for continuous and incremental backups. ``` services: kafka-backup: image: ghcr.io/osodevops/kafka-backup:latest read_only: true tmpfs: - /tmp volumes: - ./backups:/data/backups ``` To use a custom path for the offset database (e.g. for persistence), add `offset_storage.db_path` to your config and mount an additional volume: ``` services: kafka-backup: read_only: true tmpfs: - /tmp volumes: - ./backups:/data/backups - kafka-backup-state:/data/state # In your backup config: # offset_storage: # db_path: /data/state/offsets.db ``` Use Docker secrets for sensitive data: ``` services: kafka-backup: image: ghcr.io/osodevops/kafka-backup:latest secrets: - kafka_password - aws_credentials environment: - KAFKA_PASSWORD_FILE=/run/secrets/kafka_password secrets: kafka_password: file: ./secrets/kafka_password.txt aws_credentials: file: ./secrets/aws_credentials ``` The image supports multiple architectures: ``` # AMD64 (Intel/AMD) docker pull ghcr.io/osodevops/kafka-backup:latest --platform linux/amd64 # ARM64 (Apple Silicon, Graviton) docker pull ghcr.io/osodevops/kafka-backup:latest --platform linux/arm64 ``` ``` docker run --rm \ -e RUST_LOG=debug \ -v $(pwd)/backup.yaml:/config/backup.yaml:ro \ ghcr.io/osodevops/kafka-backup:latest \ -v backup --config /config/backup.yaml ``` ``` docker run -it --rm \ --entrypoint /bin/sh \ ghcr.io/osodevops/kafka-backup:latest ``` ``` docker logs kafka-backup docker logs -f kafka-backup # Follow logs ``` ``` # Test connectivity from container docker run --rm --network host \ ghcr.io/osodevops/kafka-backup:latest \ sh -c "nc -zv kafka 9092" ``` ``` # Check volume permissions ls -la ./backups # Fix permissions sudo chown -R 1000:1000 ./backups ``` ``` # Use host networking for DNS docker run --rm --network host \ -v $(pwd)/backup.yaml:/config/backup.yaml:ro \ ghcr.io/osodevops/kafka-backup:latest \ backup --config /config/backup.yaml ``` - [Kubernetes Deployment](https://kafkabackup.com/deployment/kubernetes.md) - Deploy on K8s - [Configuration Reference](https://kafkabackup.com/reference/config-yaml.md) - All options - [Backup to S3](https://kafkabackup.com/guides/backup-to-s3.md) - Cloud storage guide --- title: Kubernetes Deployment description: Deploy OSO Kafka Backup on Kubernetes manually or with the Operator source_url: html: https://kafkabackup.com/deployment/kubernetes md: https://kafkabackup.com/deployment/kubernetes.md --- # Kubernetes Deployment Deploy OSO Kafka Backup on Kubernetes clusters. This guide covers manual deployment; for automated management, see the [Kubernetes Operator](https://kafkabackup.com/operator.md). | Method | Use Case | Management | | --- | --- | --- | | **CronJob** | Scheduled backups | Manual | | **Job** | One-time backup/restore | Manual | | **Operator** | Production automation | GitOps/CRDs | > [!TIP] > > [!NOTE] > > Enterprise Edition > > [!NOTE] > > For enterprise features, swap the image reference in all manifests below: > > > > ``` > > # OSS > > image: ghcr.io/osodevops/kafka-backup:latest > > # Enterprise > > image: osodevops/kafka-backup-enterprise:latest > > ``` > > > > You will also need to mount a license file. See the [Helm Chart guide](https://kafkabackup.com/enterprise/helm-chart.md) for the recommended deployment method, or the [Enterprise Installation Guide](https://kafkabackup.com/enterprise/installation.md#kubernetes-helm-chart) for manual setup. - Kubernetes 1.21+ - kubectl configured - Kafka cluster accessible from K8s - Storage (PVC or cloud credentials) ``` kubectl create namespace kafka-backup ``` configmap.yaml ``` apiVersion: v1 kind: ConfigMap metadata: name: kafka-backup-config namespace: kafka-backup data: backup.yaml: | mode: backup backup_id: "k8s-daily-backup" source: bootstrap_servers: - kafka-0.kafka.kafka.svc:9092 - kafka-1.kafka.kafka.svc:9092 - kafka-2.kafka.kafka.svc:9092 topics: include: - "*" exclude: - "__consumer_offsets" - "_schemas" storage: backend: filesystem path: "/data/backups" backup: compression: zstd compression_level: 3 checkpoint_interval_secs: 30 include_offset_headers: true ``` ``` kubectl apply -f configmap.yaml ``` pvc.yaml ``` apiVersion: v1 kind: PersistentVolumeClaim metadata: name: kafka-backup-storage namespace: kafka-backup spec: accessModes: - ReadWriteOnce resources: requests: storage: 100Gi storageClassName: standard # Adjust for your cluster ``` ``` kubectl apply -f pvc.yaml ``` cronjob.yaml ``` apiVersion: batch/v1 kind: CronJob metadata: name: kafka-backup namespace: kafka-backup spec: schedule: "0 2 * * *" # Daily at 2 AM concurrencyPolicy: Forbid successfulJobsHistoryLimit: 3 failedJobsHistoryLimit: 3 jobTemplate: spec: backoffLimit: 2 template: spec: restartPolicy: OnFailure terminationGracePeriodSeconds: 60 # Allow graceful shutdown securityContext: runAsNonRoot: true runAsUser: 1000 fsGroup: 1000 containers: - name: kafka-backup image: ghcr.io/osodevops/kafka-backup:latest args: - backup - --config - /config/backup.yaml resources: requests: cpu: 500m memory: 512Mi limits: cpu: 2000m memory: 2Gi volumeMounts: - name: config mountPath: /config readOnly: true - name: data mountPath: /data/backups - name: tmp mountPath: /tmp securityContext: allowPrivilegeEscalation: false readOnlyRootFilesystem: true capabilities: drop: - ALL volumes: - name: config configMap: name: kafka-backup-config - name: data persistentVolumeClaim: claimName: kafka-backup-storage - name: tmp emptyDir: {} ``` ``` kubectl apply -f cronjob.yaml ``` ``` # Check CronJob kubectl get cronjob -n kafka-backup # List Jobs kubectl get jobs -n kafka-backup # Check recent Pod logs kubectl logs -n kafka-backup -l job-name=kafka-backup- ``` aws-secret.yaml ``` apiVersion: v1 kind: Secret metadata: name: aws-credentials namespace: kafka-backup type: Opaque stringData: AWS_ACCESS_KEY_ID: "AKIA..." AWS_SECRET_ACCESS_KEY: "..." ``` Or using IAM Roles for Service Accounts (IRSA): serviceaccount.yaml ``` apiVersion: v1 kind: ServiceAccount metadata: name: kafka-backup namespace: kafka-backup annotations: eks.amazonaws.com/role-arn: arn:aws:iam::123456789:role/kafka-backup-role ``` ``` data: backup.yaml: | mode: backup backup_id: "k8s-daily-backup" source: bootstrap_servers: - kafka-0.kafka.kafka.svc:9092 topics: include: - "*" storage: backend: s3 bucket: my-kafka-backups region: us-west-2 prefix: production/daily backup: compression: zstd ``` ``` spec: jobTemplate: spec: template: spec: serviceAccountName: kafka-backup # For IRSA containers: - name: kafka-backup envFrom: - secretRef: name: aws-credentials # Or use IRSA ``` - AKS cluster with Workload Identity enabled - Azure Key Vault for secret management - Azure CSI Secrets Store Driver installed ``` # Add Helm repo helm repo add csi-secrets-store-provider-azure https://azure.github.io/secrets-store-csi-driver-provider-azure/charts helm repo update # Install the driver helm install csi-secrets-store-provider-azure csi-secrets-store-provider-azure/csi-secrets-store-provider-azure \ --namespace kube-system ``` ``` # Create managed identity for Kafka Backup az identity create \ --name kafka-backup-identity \ --resource-group # Get identity client ID CLIENT_ID=$(az identity show \ --name kafka-backup-identity \ --resource-group \ --query clientId -o tsv) # Assign Storage Blob Data Contributor role az role assignment create \ --role "Storage Blob Data Contributor" \ --assignee $CLIENT_ID \ --scope /subscriptions//resourceGroups//providers/Microsoft.Storage/storageAccounts/ # Grant Key Vault access az keyvault set-policy \ --name \ --object-id $(az identity show --name kafka-backup-identity --resource-group --query principalId -o tsv) \ --secret-permissions get list ``` ``` # Get AKS OIDC issuer URL AKS_OIDC_ISSUER=$(az aks show \ --name \ --resource-group \ --query oidcIssuerProfile.issuerUrl -o tsv) # Create federated credential az identity federated-credential create \ --name kafka-backup-federated \ --identity-name kafka-backup-identity \ --resource-group \ --issuer $AKS_OIDC_ISSUER \ --subject system:serviceaccount:kafka-backup:kafka-backup \ --audience api://AzureADTokenExchange ``` serviceaccount.yaml ``` apiVersion: v1 kind: ServiceAccount metadata: name: kafka-backup namespace: kafka-backup annotations: azure.workload.identity/client-id: labels: azure.workload.identity/use: "true" ``` Use the Azure CSI Secrets Store Driver to sync secrets from Azure Key Vault: secretproviderclass.yaml ``` apiVersion: secrets-store.csi.x-k8s.io/v1 kind: SecretProviderClass metadata: name: kafka-backup-secrets namespace: kafka-backup spec: provider: azure parameters: usePodIdentity: "false" clientID: keyvaultName: tenantId: objects: | array: - | objectName: kafka-sasl-username objectType: secret - | objectName: kafka-sasl-password objectType: secret - | objectName: azure-storage-account-key objectType: secret secretObjects: - secretName: kafka-backup-secrets type: Opaque data: - objectName: kafka-sasl-username key: KAFKA_SASL_USERNAME - objectName: kafka-sasl-password key: KAFKA_SASL_PASSWORD - objectName: azure-storage-account-key key: AZURE_STORAGE_KEY ``` configmap-azure.yaml ``` apiVersion: v1 kind: ConfigMap metadata: name: kafka-backup-config namespace: kafka-backup data: backup.yaml: | mode: backup backup_id: "k8s-daily-backup" source: bootstrap_servers: - :9092 security: security_protocol: SASL_SSL sasl_mechanism: PLAIN sasl_username: ${KAFKA_SASL_USERNAME} sasl_password: ${KAFKA_SASL_PASSWORD} topics: include: - "*" exclude: - "__consumer_offsets" - "_schemas" storage: backend: azure account_name: container_name: kafka-backups account_key: ${AZURE_STORAGE_KEY} prefix: production/daily backup: compression: zstd compression_level: 3 ``` cronjob-azure.yaml ``` apiVersion: batch/v1 kind: CronJob metadata: name: kafka-backup namespace: kafka-backup spec: schedule: "0 2 * * *" concurrencyPolicy: Forbid jobTemplate: spec: template: metadata: labels: azure.workload.identity/use: "true" spec: serviceAccountName: kafka-backup restartPolicy: OnFailure containers: - name: kafka-backup image: ghcr.io/osodevops/kafka-backup:latest args: - backup - --config - /config/backup.yaml env: - name: KAFKA_SASL_USERNAME valueFrom: secretKeyRef: name: kafka-backup-secrets key: KAFKA_SASL_USERNAME - name: KAFKA_SASL_PASSWORD valueFrom: secretKeyRef: name: kafka-backup-secrets key: KAFKA_SASL_PASSWORD - name: AZURE_STORAGE_KEY valueFrom: secretKeyRef: name: kafka-backup-secrets key: AZURE_STORAGE_KEY volumeMounts: - name: config mountPath: /config readOnly: true - name: secrets-store mountPath: /mnt/secrets-store readOnly: true volumes: - name: config configMap: name: kafka-backup-config - name: secrets-store csi: driver: secrets-store.csi.k8s.io readOnly: true volumeAttributes: secretProviderClass: kafka-backup-secrets ``` For keyless authentication using Workload Identity: configmap-azure-workload-identity.yaml ``` apiVersion: v1 kind: ConfigMap metadata: name: kafka-backup-config namespace: kafka-backup data: backup.yaml: | mode: backup backup_id: "k8s-daily-backup" source: bootstrap_servers: - :9092 topics: include: - "*" storage: backend: azure account_name: container_name: kafka-backups use_workload_identity: true prefix: production/daily backup: compression: zstd ``` kafka-secret.yaml ``` apiVersion: v1 kind: Secret metadata: name: kafka-credentials namespace: kafka-backup type: Opaque stringData: username: backup-user password: your-password ``` Update ConfigMap: ``` data: backup.yaml: | source: bootstrap_servers: - kafka:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA256 sasl_username: ${KAFKA_USERNAME} sasl_password: ${KAFKA_PASSWORD} ``` Update CronJob: ``` containers: - name: kafka-backup env: - name: KAFKA_USERNAME valueFrom: secretKeyRef: name: kafka-credentials key: username - name: KAFKA_PASSWORD valueFrom: secretKeyRef: name: kafka-credentials key: password ``` tls-secret.yaml ``` apiVersion: v1 kind: Secret metadata: name: kafka-tls namespace: kafka-backup type: Opaque data: ca.crt: client.crt: client.key: ``` Mount in CronJob: ``` volumeMounts: - name: tls mountPath: /certs readOnly: true volumes: - name: tls secret: secretName: kafka-tls ``` For one-time backup or restore: backup-job.yaml ``` apiVersion: batch/v1 kind: Job metadata: name: kafka-backup-manual namespace: kafka-backup spec: ttlSecondsAfterFinished: 86400 # Clean up after 24h template: spec: restartPolicy: Never containers: - name: kafka-backup image: ghcr.io/osodevops/kafka-backup:latest args: - backup - --config - /config/backup.yaml volumeMounts: - name: config mountPath: /config - name: data mountPath: /data/backups volumes: - name: config configMap: name: kafka-backup-config - name: data persistentVolumeClaim: claimName: kafka-backup-storage ``` ``` kubectl apply -f backup-job.yaml # Watch job progress kubectl logs -n kafka-backup -f job/kafka-backup-manual ``` restore-job.yaml ``` apiVersion: batch/v1 kind: Job metadata: name: kafka-restore namespace: kafka-backup spec: template: spec: restartPolicy: Never containers: - name: kafka-backup image: ghcr.io/osodevops/kafka-backup:latest args: - restore - --config - /config/restore.yaml volumeMounts: - name: config configMap: name: kafka-restore-config items: - key: restore.yaml path: restore.yaml - name: data mountPath: /data/backups volumes: - name: config configMap: name: kafka-restore-config - name: data persistentVolumeClaim: claimName: kafka-backup-storage ``` For long-running continuous backups with metrics exposure, use a Deployment instead of a CronJob. > [!NOTE] > > [!NOTE] > > Metrics in Kubernetes > > [!NOTE] > > When running in Kubernetes, you **must** set `bind_address: "0.0.0.0"` in the metrics configuration. The default `127.0.0.1` binding only allows localhost access, which prevents the Service from routing traffic to the metrics endpoint. configmap-continuous.yaml ``` apiVersion: v1 kind: ConfigMap metadata: name: kafka-backup-config namespace: kafka-backup data: backup.yaml: | mode: backup backup_id: "k8s-continuous-backup" source: bootstrap_servers: - kafka-0.kafka.kafka.svc:9092 - kafka-1.kafka.kafka.svc:9092 - kafka-2.kafka.kafka.svc:9092 topics: include: - "*" exclude: - "__consumer_offsets" - "_schemas" storage: backend: s3 bucket: my-kafka-backups region: us-west-2 prefix: production/continuous backup: compression: zstd compression_level: 3 continuous: true checkpoint_interval_secs: 30 include_offset_headers: true # Enable metrics endpoint for Prometheus scraping metrics: enabled: true port: 8080 bind_address: "0.0.0.0" # Required for K8s Service routing path: "/metrics" logging: level: info format: json ``` deployment.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: name: kafka-backup namespace: kafka-backup labels: app: kafka-backup spec: replicas: 1 selector: matchLabels: app: kafka-backup template: metadata: labels: app: kafka-backup spec: serviceAccountName: kafka-backup securityContext: runAsNonRoot: true runAsUser: 1000 fsGroup: 1000 containers: - name: kafka-backup image: ghcr.io/osodevops/kafka-backup:latest args: - backup - --config - /config/backup.yaml ports: - name: metrics containerPort: 8080 protocol: TCP livenessProbe: httpGet: path: /health port: metrics initialDelaySeconds: 10 periodSeconds: 30 timeoutSeconds: 5 readinessProbe: httpGet: path: /health port: metrics initialDelaySeconds: 5 periodSeconds: 10 timeoutSeconds: 3 resources: requests: cpu: 500m memory: 512Mi limits: cpu: 2000m memory: 2Gi volumeMounts: - name: config mountPath: /config readOnly: true - name: tmp mountPath: /tmp securityContext: allowPrivilegeEscalation: false readOnlyRootFilesystem: true capabilities: drop: - ALL volumes: - name: config configMap: name: kafka-backup-config - name: tmp emptyDir: {} ``` > [!WARNING] > > [!NOTE] > > Required: `/tmp` volume with `readOnlyRootFilesystem` > > [!NOTE] > > When using `readOnlyRootFilesystem: true`, you **must** mount a writable `/tmp` volume. The backup engine stores its offset tracking database in `$TMPDIR` (defaults to `/tmp`) for continuous and incremental backups. Without this mount, the backup will fail with: > > > > ``` > > Error: Storage error: Backend error: error returned from database: (code: 14) unable to open database file > > ``` > > > > See [GitHub Issue #62](https://github.com/osodevops/kafka-backup/issues/62) for details. The Service exposes the metrics endpoint so Prometheus can scrape it: service.yaml ``` apiVersion: v1 kind: Service metadata: name: kafka-backup-metrics namespace: kafka-backup labels: app: kafka-backup spec: selector: app: kafka-backup ports: - name: metrics port: 8080 targetPort: metrics protocol: TCP type: ClusterIP ``` > [!TIP] > > [!NOTE] > > Why is the Service necessary? > > [!NOTE] > > Without a Service, there's no stable endpoint to scrape metrics from. Pod IPs are ephemeral and change when pods restart. The Service provides a stable DNS name (`kafka-backup-metrics.kafka-backup.svc`) that Prometheus can use. ``` kubectl apply -f configmap-continuous.yaml kubectl apply -f deployment.yaml kubectl apply -f service.yaml # Verify metrics are accessible kubectl port-forward -n kafka-backup svc/kafka-backup-metrics 8080:8080 curl http://localhost:8080/metrics ``` pdb.yaml ``` apiVersion: policy/v1 kind: PodDisruptionBudget metadata: name: kafka-backup-pdb namespace: kafka-backup spec: minAvailable: 0 selector: matchLabels: app: kafka-backup ``` The ServiceMonitor requires a Service to discover endpoints. If you haven't created one already: service.yaml ``` apiVersion: v1 kind: Service metadata: name: kafka-backup-metrics namespace: kafka-backup labels: app: kafka-backup spec: selector: app: kafka-backup ports: - name: metrics port: 8080 targetPort: metrics protocol: TCP type: ClusterIP ``` servicemonitor.yaml ``` apiVersion: monitoring.coreos.com/v1 kind: ServiceMonitor metadata: name: kafka-backup namespace: kafka-backup labels: release: prometheus # Match your Prometheus Operator's serviceMonitorSelector spec: selector: matchLabels: app: kafka-backup endpoints: - port: metrics interval: 15s path: /metrics namespaceSelector: matchNames: - kafka-backup ``` If you're not using the Prometheus Operator, add a scrape config: prometheus.yml ``` scrape_configs: - job_name: 'kafka-backup' kubernetes_sd_configs: - role: endpoints namespaces: names: - kafka-backup relabel_configs: - source_labels: [__meta_kubernetes_service_label_app] regex: kafka-backup action: keep - source_labels: [__meta_kubernetes_endpoint_port_name] regex: metrics action: keep ``` OSO Kafka Backup handles `SIGTERM` and `SIGINT` signals for graceful shutdown (v0.8.1+). When a pod is terminated, the backup engine: 1. Stops consuming new records 2. Flushes in-progress segments to storage 3. Saves a final checkpoint 4. Exits with code 0 Set `terminationGracePeriodSeconds` to give the process enough time to flush (default Kubernetes value is 30s, 60s is recommended): ``` spec: template: spec: terminationGracePeriodSeconds: 60 ``` > [!TIP] > > [!NOTE] > > Snapshot Backups with CronJob > > [!NOTE] > > For scheduled DR backups, use `stop_at_current_offsets: true` in your backup config. This captures a consistent point-in-time snapshot and exits cleanly when all partitions catch up — ideal for CronJob-based backups. | Cluster Size | CPU Request | Memory Request | CPU Limit | Memory Limit | | --- | --- | --- | --- | --- | | Small | 250m | 256Mi | 1000m | 1Gi | | Medium | 500m | 512Mi | 2000m | 2Gi | | Large | 1000m | 1Gi | 4000m | 4Gi | ``` kubectl get jobs -n kafka-backup kubectl describe job kafka-backup-manual -n kafka-backup ``` ``` kubectl logs -n kafka-backup -l job-name=kafka-backup- --tail=100 ``` ``` kubectl run -n kafka-backup debug --rm -it \ --image=ghcr.io/osodevops/kafka-backup:latest \ --restart=Never -- /bin/sh ``` - [Kubernetes Operator](https://kafkabackup.com/operator.md) - Automated CRD-based management - [Security Setup](https://kafkabackup.com/guides/security-setup.md) - TLS and SASL configuration - [AWS S3 Setup](https://kafkabackup.com/deployment/cloud-setup/aws-s3.md) - S3 storage configuration --- title: AWS S3 Setup description: Configure AWS S3 as storage backend for OSO Kafka Backup source_url: html: https://kafkabackup.com/deployment/cloud-setup/aws-s3 md: https://kafkabackup.com/deployment/cloud-setup/aws-s3.md --- # AWS S3 Setup Configure Amazon S3 or S3-compatible storage for Kafka backups. - AWS account with S3 access - AWS CLI configured (optional, for testing) - IAM permissions to create buckets and policies ``` # Create bucket aws s3 mb s3://my-kafka-backups --region us-west-2 # Enable versioning (recommended) aws s3api put-bucket-versioning \ --bucket my-kafka-backups \ --versioning-configuration Status=Enabled # Enable server-side encryption aws s3api put-bucket-encryption \ --bucket my-kafka-backups \ --server-side-encryption-configuration '{ "Rules": [ { "ApplyServerSideEncryptionByDefault": { "SSEAlgorithm": "AES256" } } ] }' ``` s3.tf ``` resource "aws_s3_bucket" "kafka_backups" { bucket = "my-kafka-backups" } resource "aws_s3_bucket_versioning" "kafka_backups" { bucket = aws_s3_bucket.kafka_backups.id versioning_configuration { status = "Enabled" } } resource "aws_s3_bucket_server_side_encryption_configuration" "kafka_backups" { bucket = aws_s3_bucket.kafka_backups.id rule { apply_server_side_encryption_by_default { sse_algorithm = "AES256" } } } resource "aws_s3_bucket_lifecycle_configuration" "kafka_backups" { bucket = aws_s3_bucket.kafka_backups.id rule { id = "transition-to-ia" status = "Enabled" transition { days = 30 storage_class = "STANDARD_IA" } transition { days = 90 storage_class = "GLACIER" } expiration { days = 365 } } } ``` kafka-backup-policy.json ``` { "Version": "2012-10-17", "Statement": [ { "Sid": "ListBucket", "Effect": "Allow", "Action": [ "s3:ListBucket", "s3:GetBucketLocation" ], "Resource": "arn:aws:s3:::my-kafka-backups" }, { "Sid": "ObjectOperations", "Effect": "Allow", "Action": [ "s3:GetObject", "s3:PutObject", "s3:DeleteObject" ], "Resource": "arn:aws:s3:::my-kafka-backups/*" } ] } ``` ``` # Create user aws iam create-user --user-name kafka-backup # Attach policy aws iam put-user-policy \ --user-name kafka-backup \ --policy-name kafka-backup-s3 \ --policy-document file://kafka-backup-policy.json # Create access keys aws iam create-access-key --user-name kafka-backup ``` trust-policy.json ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": { "Service": "ec2.amazonaws.com" }, "Action": "sts:AssumeRole" } ] } ``` ``` # Create role aws iam create-role \ --role-name kafka-backup-role \ --assume-role-policy-document file://trust-policy.json # Attach policy aws iam put-role-policy \ --role-name kafka-backup-role \ --policy-name kafka-backup-s3 \ --policy-document file://kafka-backup-policy.json ``` For Kubernetes on EKS, use IAM Roles for Service Accounts: ``` # Create OIDC provider (if not exists) eksctl utils associate-iam-oidc-provider \ --cluster my-cluster \ --approve # Create service account with IAM role eksctl create iamserviceaccount \ --name kafka-backup \ --namespace kafka-backup \ --cluster my-cluster \ --attach-policy-arn arn:aws:iam::123456789:policy/kafka-backup-s3 \ --approve ``` backup.yaml ``` storage: backend: s3 bucket: my-kafka-backups region: us-west-2 prefix: production/daily access_key: ${AWS_ACCESS_KEY_ID} secret_key: ${AWS_SECRET_ACCESS_KEY} ``` backup.yaml ``` storage: backend: s3 bucket: my-kafka-backups region: us-west-2 prefix: production/daily # No credentials needed - uses instance profile or IRSA ``` backup.yaml ``` storage: backend: s3 bucket: kafka-backups region: us-east-1 endpoint: https://minio.example.com:9000 access_key: ${MINIO_ACCESS_KEY} secret_key: ${MINIO_SECRET_KEY} ``` | Variable | Description | | --- | --- | | `AWS_ACCESS_KEY_ID` | AWS access key | | `AWS_SECRET_ACCESS_KEY` | AWS secret key | | `AWS_REGION` | AWS region | | `AWS_DEFAULT_REGION` | Alternative region variable | | `AWS_PROFILE` | AWS CLI profile name | Optimize costs with lifecycle policies: | Storage Class | Use Case | Cost | | --- | --- | --- | | `STANDARD` | Recent backups (< 30 days) | $$$ | | `STANDARD_IA` | Infrequent access (30-90 days) | $$ | | `GLACIER_IR` | Archive with fast retrieval | $ | | `GLACIER` | Long-term archive | $ | | `DEEP_ARCHIVE` | Compliance archives | $ | ``` aws s3api put-bucket-lifecycle-configuration \ --bucket my-kafka-backups \ --lifecycle-configuration '{ "Rules": [ { "ID": "archive-old-backups", "Status": "Enabled", "Filter": { "Prefix": "production/" }, "Transitions": [ { "Days": 30, "StorageClass": "STANDARD_IA" }, { "Days": 90, "StorageClass": "GLACIER" } ], "Expiration": { "Days": 365 } } ] }' ``` For disaster recovery: ``` # Enable versioning on destination bucket aws s3api put-bucket-versioning \ --bucket my-kafka-backups-dr \ --versioning-configuration Status=Enabled # Create replication configuration aws s3api put-bucket-replication \ --bucket my-kafka-backups \ --replication-configuration '{ "Role": "arn:aws:iam::123456789:role/s3-replication-role", "Rules": [ { "Status": "Enabled", "Priority": 1, "DeleteMarkerReplication": { "Status": "Enabled" }, "Filter": {}, "Destination": { "Bucket": "arn:aws:s3:::my-kafka-backups-dr" } } ] }' ``` ``` # Test AWS credentials aws sts get-caller-identity # Test bucket access aws s3 ls s3://my-kafka-backups/ # Test write access echo "test" | aws s3 cp - s3://my-kafka-backups/test.txt aws s3 rm s3://my-kafka-backups/test.txt ``` ``` # Run a test backup kafka-backup backup --config backup.yaml # List backups kafka-backup list --path s3://my-kafka-backups/production/daily ``` ``` # Check bucket policy aws s3api get-bucket-policy --bucket my-kafka-backups # Check IAM permissions aws iam simulate-principal-policy \ --policy-source-arn arn:aws:iam::123456789:user/kafka-backup \ --action-names s3:PutObject s3:GetObject \ --resource-arns arn:aws:s3:::my-kafka-backups/* ``` - Use multipart uploads for large segments - Check network bandwidth - Consider using S3 Transfer Acceleration - Enable Intelligent-Tiering for automatic class transitions - Use lifecycle policies aggressively - Monitor with S3 Storage Lens 1. **Enable versioning** - Protect against accidental deletes 2. **Enable encryption** - Use SSE-S3 or SSE-KMS 3. **Use IAM roles** - Avoid static credentials 4. **Enable access logging** - Audit bucket access 5. **Block public access** - Ensure bucket is private 6. **Enable MFA delete** - For critical backups ``` # Block all public access aws s3api put-public-access-block \ --bucket my-kafka-backups \ --public-access-block-configuration \ "BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true" ``` - [Backup to S3 Guide](https://kafkabackup.com/guides/backup-to-s3.md) - Step-by-step backup tutorial - [Configuration Reference](https://kafkabackup.com/reference/config-yaml.md) - All storage options - [Performance Tuning](https://kafkabackup.com/guides/performance-tuning.md) - Optimize throughput --- title: Azure Blob Storage Setup description: Configure Azure Blob Storage as backend for OSO Kafka Backup source_url: html: https://kafkabackup.com/deployment/cloud-setup/azure-blob md: https://kafkabackup.com/deployment/cloud-setup/azure-blob.md --- # Azure Blob Storage Setup Configure Azure Blob Storage for Kafka backups. - Azure subscription - Azure CLI installed - Permissions to create storage accounts ``` # Set variables RESOURCE_GROUP="kafka-backup-rg" STORAGE_ACCOUNT="kafkabackups$(date +%s)" # Must be globally unique LOCATION="westus2" # Create resource group az group create --name $RESOURCE_GROUP --location $LOCATION # Create storage account az storage account create \ --name $STORAGE_ACCOUNT \ --resource-group $RESOURCE_GROUP \ --location $LOCATION \ --sku Standard_LRS \ --kind StorageV2 \ --min-tls-version TLS1_2 \ --allow-blob-public-access false # Create container az storage container create \ --name kafka-backups \ --account-name $STORAGE_ACCOUNT \ --auth-mode login ``` azure-storage.tf ``` resource "azurerm_resource_group" "kafka_backup" { name = "kafka-backup-rg" location = "West US 2" } resource "azurerm_storage_account" "kafka_backup" { name = "kafkabackups${random_string.suffix.result}" resource_group_name = azurerm_resource_group.kafka_backup.name location = azurerm_resource_group.kafka_backup.location account_tier = "Standard" account_replication_type = "LRS" min_tls_version = "TLS1_2" blob_properties { versioning_enabled = true delete_retention_policy { days = 30 } } } resource "azurerm_storage_container" "kafka_backups" { name = "kafka-backups" storage_account_name = azurerm_storage_account.kafka_backup.name container_access_type = "private" } resource "random_string" "suffix" { length = 8 special = false upper = false } ``` ``` # Get storage account key az storage account keys list \ --account-name $STORAGE_ACCOUNT \ --resource-group $RESOURCE_GROUP \ --query "[0].value" -o tsv ``` ``` # Get connection string az storage account show-connection-string \ --name $STORAGE_ACCOUNT \ --resource-group $RESOURCE_GROUP \ --query connectionString -o tsv ``` ``` # Enable system-assigned managed identity on VM az vm identity assign \ --resource-group $RESOURCE_GROUP \ --name my-vm # Get principal ID PRINCIPAL_ID=$(az vm show \ --resource-group $RESOURCE_GROUP \ --name my-vm \ --query identity.principalId -o tsv) # Assign Storage Blob Data Contributor role az role assignment create \ --role "Storage Blob Data Contributor" \ --assignee $PRINCIPAL_ID \ --scope "/subscriptions//resourceGroups/$RESOURCE_GROUP/providers/Microsoft.Storage/storageAccounts/$STORAGE_ACCOUNT" ``` ``` # Create service principal az ad sp create-for-rbac \ --name kafka-backup-sp \ --role "Storage Blob Data Contributor" \ --scopes "/subscriptions//resourceGroups/$RESOURCE_GROUP/providers/Microsoft.Storage/storageAccounts/$STORAGE_ACCOUNT" ``` OSO Kafka Backup v0.2.1+ supports multiple Azure authentication methods with automatic detection. | Method | Use Case | Required Fields | | --- | --- | --- | | Account Key | Development/testing | `account_key` | | SAS Token | Time-limited access | `sas_token` | | Service Principal | CI/CD pipelines | `client_id`, `tenant_id`, `client_secret` | | Workload Identity | AKS with managed identity | `use_workload_identity: true` (or auto-detected) | | DefaultAzureCredential | Fallback chain | No auth fields (uses Azure SDK default) | backup.yaml ``` storage: backend: azure account_name: kafkabackups123456 container_name: kafka-backups account_key: ${AZURE_STORAGE_KEY} prefix: production/daily ``` backup.yaml ``` storage: backend: azure account_name: kafkabackups123456 container_name: kafka-backups sas_token: ${AZURE_SAS_TOKEN} prefix: production/daily ``` backup.yaml ``` storage: backend: azure account_name: kafkabackups123456 container_name: kafka-backups client_id: ${AZURE_CLIENT_ID} tenant_id: ${AZURE_TENANT_ID} client_secret: ${AZURE_CLIENT_SECRET} prefix: production/daily ``` backup.yaml ``` storage: backend: azure account_name: kafkabackups123456 container_name: kafka-backups use_workload_identity: true # Or auto-detected via AZURE_FEDERATED_TOKEN_FILE prefix: production/daily # client_id / tenant_id are optional overrides of the webhook-injected values (v0.22.0+) ``` This works for the CLI running as a plain Kubernetes Job — no operator needed. The AKS Workload Identity webhook injects `AZURE_CLIENT_ID`, `AZURE_TENANT_ID`, `AZURE_AUTHORITY_HOST` and `AZURE_FEDERATED_TOKEN_FILE` into a pod when: 1. its **ServiceAccount** is annotated `azure.workload.identity/client-id: `, and 2. the **pod** carries the label `azure.workload.identity/use: "true"`. On the Azure side the managed identity (or app registration) needs a **federated credential** with issuer = the cluster's OIDC issuer (`az aks show … --query oidcIssuerProfile.issuerUrl`), subject = `system:serviceaccount::`, audience `api://AzureADTokenExchange`, and the **Storage Blob Data Contributor** role on the container. Run with `RUST_LOG=debug` to see `Azure authentication: Workload Identity (client_id=…, tenant_id=…, token_file=…)`; a `not set` token file means the label or annotation is missing. Credential precedence: `sas_token` → `account_key` → service principal → Workload Identity → DefaultAzureCredential chain. backup.yaml ``` storage: backend: azure account_name: kafkabackups123456 container_name: kafka-backups # No credentials - uses Azure SDK's DefaultAzureCredential chain prefix: production/daily ``` For Azure Government, China, or private endpoints: backup.yaml ``` storage: backend: azure account_name: kafkabackups123456 container_name: kafka-backups endpoint: https://kafkabackups123456.blob.core.usgovcloudapi.net account_key: ${AZURE_STORAGE_KEY} prefix: production/daily ``` | Variable | Description | | --- | --- | | `AZURE_STORAGE_ACCOUNT` | Storage account name | | `AZURE_STORAGE_KEY` | Storage account key | | `AZURE_SAS_TOKEN` | SAS token for time-limited access | | `AZURE_CLIENT_ID` | Service principal or managed identity client ID | | `AZURE_CLIENT_SECRET` | Service principal secret | | `AZURE_TENANT_ID` | Azure AD tenant ID | | `AZURE_FEDERATED_TOKEN_FILE` | Auto-injected by Workload Identity webhook | | `AZURE_AUTHORITY_HOST` | Azure AD authority (auto-injected) | ``` # Generate SAS token with 1-year expiry END_DATE=$(date -u -d "+1 year" '+%Y-%m-%dT%H:%MZ') az storage container generate-sas \ --account-name $STORAGE_ACCOUNT \ --name kafka-backups \ --permissions rwdl \ --expiry $END_DATE \ --auth-mode login \ --as-user \ -o tsv ``` For AKS, use workload identity: ``` # Enable workload identity on cluster az aks update \ --resource-group $RESOURCE_GROUP \ --name my-aks-cluster \ --enable-oidc-issuer \ --enable-workload-identity # Create managed identity az identity create \ --name kafka-backup-identity \ --resource-group $RESOURCE_GROUP # Get identity client ID CLIENT_ID=$(az identity show \ --name kafka-backup-identity \ --resource-group $RESOURCE_GROUP \ --query clientId -o tsv) # Assign role az role assignment create \ --role "Storage Blob Data Contributor" \ --assignee $CLIENT_ID \ --scope "/subscriptions//resourceGroups/$RESOURCE_GROUP/providers/Microsoft.Storage/storageAccounts/$STORAGE_ACCOUNT" # Create federated credential az identity federated-credential create \ --name kafka-backup-federated \ --identity-name kafka-backup-identity \ --resource-group $RESOURCE_GROUP \ --issuer $(az aks show --name my-aks-cluster --resource-group $RESOURCE_GROUP --query oidcIssuerProfile.issuerUrl -o tsv) \ --subject system:serviceaccount:kafka-backup:kafka-backup \ --audience api://AzureADTokenExchange ``` Kubernetes ServiceAccount: ``` apiVersion: v1 kind: ServiceAccount metadata: name: kafka-backup namespace: kafka-backup annotations: azure.workload.identity/client-id: labels: azure.workload.identity/use: "true" ``` Use Azure Key Vault with the CSI Secrets Store Driver to securely manage credentials for Kafka and storage access. - AKS cluster with Workload Identity enabled - Azure Key Vault created - Azure CSI Secrets Store Driver installed ``` # Add the Helm repo helm repo add csi-secrets-store-provider-azure \ https://azure.github.io/secrets-store-csi-driver-provider-azure/charts helm repo update # Install the Azure provider helm install csi-secrets-store-provider-azure \ csi-secrets-store-provider-azure/csi-secrets-store-provider-azure \ --namespace kube-system ``` ``` # Get the managed identity principal ID PRINCIPAL_ID=$(az identity show \ --name kafka-backup-identity \ --resource-group $RESOURCE_GROUP \ --query principalId -o tsv) # Grant secret access to Key Vault az keyvault set-policy \ --name $KEY_VAULT_NAME \ --object-id $PRINCIPAL_ID \ --secret-permissions get list ``` ``` # Store Kafka credentials az keyvault secret set \ --vault-name $KEY_VAULT_NAME \ --name kafka-sasl-username \ --value "" az keyvault secret set \ --vault-name $KEY_VAULT_NAME \ --name kafka-sasl-password \ --value "" # Store storage account key (if not using Workload Identity for storage) az keyvault secret set \ --vault-name $KEY_VAULT_NAME \ --name azure-storage-account-key \ --value "" ``` The SecretProviderClass syncs secrets from Azure Key Vault to Kubernetes secrets: secretproviderclass.yaml ``` apiVersion: secrets-store.csi.x-k8s.io/v1 kind: SecretProviderClass metadata: name: kafka-backup-secrets namespace: kafka-backup spec: provider: azure parameters: usePodIdentity: "false" clientID: keyvaultName: tenantId: objects: | array: - | objectName: kafka-sasl-username objectType: secret - | objectName: kafka-sasl-password objectType: secret - | objectName: azure-storage-account-key objectType: secret secretObjects: - secretName: kafka-backup-secrets type: Opaque data: - objectName: kafka-sasl-username key: KAFKA_SASL_USERNAME - objectName: kafka-sasl-password key: KAFKA_SASL_PASSWORD - objectName: azure-storage-account-key key: AZURE_STORAGE_KEY ``` Mount the CSI volume and reference the synced Kubernetes secret: pod-with-secrets.yaml ``` apiVersion: v1 kind: Pod metadata: name: kafka-backup namespace: kafka-backup labels: azure.workload.identity/use: "true" spec: serviceAccountName: kafka-backup containers: - name: kafka-backup image: ghcr.io/osodevops/kafka-backup:latest env: - name: KAFKA_SASL_USERNAME valueFrom: secretKeyRef: name: kafka-backup-secrets key: KAFKA_SASL_USERNAME - name: KAFKA_SASL_PASSWORD valueFrom: secretKeyRef: name: kafka-backup-secrets key: KAFKA_SASL_PASSWORD - name: AZURE_STORAGE_KEY valueFrom: secretKeyRef: name: kafka-backup-secrets key: AZURE_STORAGE_KEY volumeMounts: - name: secrets-store mountPath: /mnt/secrets-store readOnly: true volumes: - name: secrets-store csi: driver: secrets-store.csi.k8s.io readOnly: true volumeAttributes: secretProviderClass: kafka-backup-secrets ``` When using the Kafka Backup Operator Helm chart, configure Key Vault integration: values-azure-keyvault.yaml ``` deployment: tenantId: workloadIdentityClientId: serviceAccountName: kafka-backup syncSecrets: keyVaultName: envSecrets: kafka-sasl-username: KAFKA_SASL_USERNAME kafka-sasl-password: KAFKA_SASL_PASSWORD azure-storage-account-key: AZURE_STORAGE_KEY ``` This configuration: - Uses Workload Identity for authentication (no stored credentials) - Syncs secrets from Azure Key Vault to Kubernetes secrets - Maps Key Vault secret names to environment variable names ``` # Create lifecycle policy JSON cat > lifecycle-policy.json << 'EOF' { "rules": [ { "enabled": true, "name": "archive-old-backups", "type": "Lifecycle", "definition": { "actions": { "baseBlob": { "tierToCool": { "daysAfterModificationGreaterThan": 30 }, "tierToArchive": { "daysAfterModificationGreaterThan": 90 }, "delete": { "daysAfterModificationGreaterThan": 365 } } }, "filters": { "blobTypes": ["blockBlob"], "prefixMatch": ["kafka-backups/production/"] } } } ] } EOF # Apply policy az storage account management-policy create \ --account-name $STORAGE_ACCOUNT \ --resource-group $RESOURCE_GROUP \ --policy @lifecycle-policy.json ``` | Tier | Use Case | Access Time | Cost | | --- | --- | --- | --- | | Hot | Frequent access | Instant | $$$ | | Cool | Infrequent (30+ days) | Instant | $$ | | Archive | Rare access (180+ days) | Hours | $ | For disaster recovery: ``` # Create storage account with geo-redundancy az storage account create \ --name $STORAGE_ACCOUNT \ --resource-group $RESOURCE_GROUP \ --location $LOCATION \ --sku Standard_GRS \ --kind StorageV2 ``` Redundancy options: | SKU | Description | | --- | --- | | `Standard_LRS` | Locally redundant | | `Standard_ZRS` | Zone redundant | | `Standard_GRS` | Geo-redundant | | `Standard_RAGRS` | Read-access geo-redundant | ``` # Enable infrastructure encryption az storage account update \ --name $STORAGE_ACCOUNT \ --resource-group $RESOURCE_GROUP \ --require-infrastructure-encryption ``` ``` # Restrict to specific VNet az storage account network-rule add \ --account-name $STORAGE_ACCOUNT \ --resource-group $RESOURCE_GROUP \ --vnet-name my-vnet \ --subnet my-subnet # Set default action to deny az storage account update \ --name $STORAGE_ACCOUNT \ --resource-group $RESOURCE_GROUP \ --default-action Deny ``` ``` az storage blob service-properties delete-policy update \ --account-name $STORAGE_ACCOUNT \ --enable true \ --days-retained 30 ``` ``` # Test with Azure CLI az storage blob list \ --account-name $STORAGE_ACCOUNT \ --container-name kafka-backups \ --auth-mode login # Test write echo "test" | az storage blob upload \ --account-name $STORAGE_ACCOUNT \ --container-name kafka-backups \ --name test.txt \ --data @- \ --auth-mode login # Clean up az storage blob delete \ --account-name $STORAGE_ACCOUNT \ --container-name kafka-backups \ --name test.txt \ --auth-mode login ``` ``` # Run backup kafka-backup backup --config backup.yaml # List backups kafka-backup list --path azure://$STORAGE_ACCOUNT/kafka-backups/production/daily ``` ``` # Check credentials az storage account keys list --account-name $STORAGE_ACCOUNT # Test with connection string az storage container list --connection-string "$AZURE_STORAGE_CONNECTION_STRING" ``` ``` # Check firewall rules az storage account network-rule list \ --account-name $STORAGE_ACCOUNT \ --resource-group $RESOURCE_GROUP # Test connectivity curl -v https://$STORAGE_ACCOUNT.blob.core.windows.net/ ``` - [Configuration Reference](https://kafkabackup.com/reference/config-yaml.md) - All storage options - [Backup Guide](https://kafkabackup.com/guides/backup-to-s3.md) - Backup walkthrough - [Security Setup](https://kafkabackup.com/guides/security-setup.md) - Security configuration --- title: Google Cloud Storage Setup description: Configure Google Cloud Storage as backend for OSO Kafka Backup source_url: html: https://kafkabackup.com/deployment/cloud-setup/gcs md: https://kafkabackup.com/deployment/cloud-setup/gcs.md --- # Google Cloud Storage Setup Configure Google Cloud Storage (GCS) for Kafka backups. - Google Cloud project - gcloud CLI installed and configured - Permissions to create buckets and service accounts ``` # Set variables PROJECT_ID="my-project" BUCKET_NAME="my-kafka-backups" REGION="us-west1" # Create bucket gcloud storage buckets create gs://$BUCKET_NAME \ --project=$PROJECT_ID \ --location=$REGION \ --uniform-bucket-level-access # Enable versioning gcloud storage buckets update gs://$BUCKET_NAME --versioning # Set lifecycle policy cat > lifecycle.json << 'EOF' { "lifecycle": { "rule": [ { "action": {"type": "SetStorageClass", "storageClass": "NEARLINE"}, "condition": {"age": 30} }, { "action": {"type": "SetStorageClass", "storageClass": "COLDLINE"}, "condition": {"age": 90} }, { "action": {"type": "Delete"}, "condition": {"age": 365} } ] } } EOF gcloud storage buckets update gs://$BUCKET_NAME --lifecycle-file=lifecycle.json ``` gcs.tf ``` resource "google_storage_bucket" "kafka_backups" { name = "my-kafka-backups" location = "US-WEST1" project = var.project_id uniform_bucket_level_access = true versioning { enabled = true } lifecycle_rule { condition { age = 30 } action { type = "SetStorageClass" storage_class = "NEARLINE" } } lifecycle_rule { condition { age = 90 } action { type = "SetStorageClass" storage_class = "COLDLINE" } } lifecycle_rule { condition { age = 365 } action { type = "Delete" } } } ``` ``` # Create service account gcloud iam service-accounts create kafka-backup \ --display-name="Kafka Backup Service Account" \ --project=$PROJECT_ID # Get service account email SA_EMAIL="kafka-backup@${PROJECT_ID}.iam.gserviceaccount.com" # Grant bucket access gcloud storage buckets add-iam-policy-binding gs://$BUCKET_NAME \ --member="serviceAccount:$SA_EMAIL" \ --role="roles/storage.objectAdmin" # Create key file gcloud iam service-accounts keys create kafka-backup-key.json \ --iam-account=$SA_EMAIL ``` ``` # Enable workload identity on cluster gcloud container clusters update my-cluster \ --zone=us-west1-a \ --workload-pool=${PROJECT_ID}.svc.id.goog # Create Kubernetes service account kubectl create serviceaccount kafka-backup -n kafka-backup # Bind Kubernetes SA to Google SA gcloud iam service-accounts add-iam-policy-binding $SA_EMAIL \ --role="roles/iam.workloadIdentityUser" \ --member="serviceAccount:${PROJECT_ID}.svc.id.goog[kafka-backup/kafka-backup]" # Annotate Kubernetes SA kubectl annotate serviceaccount kafka-backup \ -n kafka-backup \ iam.gke.io/gcp-service-account=$SA_EMAIL ``` On GCE VMs or Cloud Run: ``` # No explicit configuration needed # Uses instance metadata service automatically ``` backup.yaml ``` storage: backend: gcs bucket: my-kafka-backups prefix: production/daily service_account_json: /path/to/kafka-backup-key.json ``` backup.yaml ``` storage: backend: gcs bucket: my-kafka-backups prefix: production/daily # Uses GOOGLE_APPLICATION_CREDENTIALS environment variable ``` backup.yaml ``` storage: backend: gcs bucket: my-kafka-backups prefix: production/daily # No credentials needed - uses workload identity ``` | Variable | Description | | --- | --- | | `GOOGLE_APPLICATION_CREDENTIALS` | Path to service account JSON key | | `GOOGLE_CLOUD_PROJECT` | Default project ID | | `CLOUDSDK_CORE_PROJECT` | Alternative project variable | | Role | Description | | --- | --- | | `roles/storage.objectViewer` | Read backups | | `roles/storage.objectCreator` | Create backups | | `roles/storage.objectAdmin` | Full access (recommended) | | `roles/storage.admin` | Bucket management | Minimum required permissions: ``` - storage.objects.create - storage.objects.delete - storage.objects.get - storage.objects.list ``` | Class | Use Case | Minimum Storage | Retrieval Cost | | --- | --- | --- | --- | | `STANDARD` | Frequent access | None | Free | | `NEARLINE` | Monthly access | 30 days | $ | | `COLDLINE` | Quarterly access | 90 days | $$ | | `ARCHIVE` | Yearly access | 365 days | $$$ | ``` gcloud storage buckets update gs://$BUCKET_NAME \ --default-storage-class=NEARLINE ``` For high availability: ``` # Create multi-region bucket gcloud storage buckets create gs://$BUCKET_NAME \ --location=US \ --uniform-bucket-level-access # Create dual-region bucket gcloud storage buckets create gs://$BUCKET_NAME \ --location=NAM4 \ --uniform-bucket-level-access ``` Location options: | Type | Examples | Use Case | | --- | --- | --- | | Region | `us-west1` | Single region | | Dual-region | `NAM4` (Iowa + SC) | HA within continent | | Multi-region | `US`, `EU`, `ASIA` | Global access | ``` # View current lifecycle gcloud storage buckets describe gs://$BUCKET_NAME --format="json(lifecycle)" # Update lifecycle cat > lifecycle.json << 'EOF' { "lifecycle": { "rule": [ { "action": {"type": "SetStorageClass", "storageClass": "NEARLINE"}, "condition": {"age": 30, "matchesPrefix": ["production/"]} }, { "action": {"type": "Delete"}, "condition": {"age": 365} }, { "action": {"type": "Delete"}, "condition": {"numNewerVersions": 3} } ] } } EOF gcloud storage buckets update gs://$BUCKET_NAME --lifecycle-file=lifecycle.json ``` ``` gcloud storage buckets update gs://$BUCKET_NAME \ --uniform-bucket-level-access ``` ``` gcloud storage buckets update gs://$BUCKET_NAME --versioning ``` ``` # Create key ring gcloud kms keyrings create kafka-backup-ring \ --location=us-west1 \ --project=$PROJECT_ID # Create key gcloud kms keys create kafka-backup-key \ --keyring=kafka-backup-ring \ --location=us-west1 \ --purpose=encryption \ --project=$PROJECT_ID # Set bucket encryption gcloud storage buckets update gs://$BUCKET_NAME \ --default-encryption-key=projects/$PROJECT_ID/locations/us-west1/keyRings/kafka-backup-ring/cryptoKeys/kafka-backup-key ``` ``` # Create logging bucket gcloud storage buckets create gs://${BUCKET_NAME}-logs \ --location=$REGION # Enable logging gcloud storage buckets update gs://$BUCKET_NAME \ --log-bucket=gs://${BUCKET_NAME}-logs ``` ``` # Test with gcloud gcloud storage ls gs://$BUCKET_NAME/ # Test write echo "test" | gcloud storage cp - gs://$BUCKET_NAME/test.txt gcloud storage rm gs://$BUCKET_NAME/test.txt # Test with gsutil gsutil ls gs://$BUCKET_NAME/ ``` ``` # Check active account gcloud auth list # Test with service account gcloud auth activate-service-account --key-file=kafka-backup-key.json gcloud storage ls gs://$BUCKET_NAME/ ``` ``` # Set credentials export GOOGLE_APPLICATION_CREDENTIALS="/path/to/kafka-backup-key.json" # Run backup kafka-backup backup --config backup.yaml # List backups kafka-backup list --path gs://$BUCKET_NAME/production/daily ``` deployment.yaml ``` apiVersion: batch/v1 kind: CronJob metadata: name: kafka-backup namespace: kafka-backup spec: schedule: "0 2 * * *" jobTemplate: spec: template: spec: serviceAccountName: kafka-backup # For workload identity containers: - name: kafka-backup image: ghcr.io/osodevops/kafka-backup:latest args: ["backup", "--config", "/config/backup.yaml"] volumeMounts: - name: config mountPath: /config # Only needed if not using workload identity: # - name: gcp-sa # mountPath: /var/secrets/google # env: # - name: GOOGLE_APPLICATION_CREDENTIALS # value: /var/secrets/google/key.json volumes: - name: config configMap: name: kafka-backup-config # - name: gcp-sa # secret: # secretName: gcp-sa-key ``` ``` # Check IAM policy gcloud storage buckets get-iam-policy gs://$BUCKET_NAME # Check service account permissions gcloud projects get-iam-policy $PROJECT_ID \ --filter="bindings.members:$SA_EMAIL" \ --format="table(bindings.role)" ``` ``` # Verify workload identity binding gcloud iam service-accounts get-iam-policy $SA_EMAIL # Check pod service account annotation kubectl get sa kafka-backup -n kafka-backup -o yaml ``` - Use regional buckets close to your Kafka cluster - Enable parallel composite uploads - Check network bandwidth - [Configuration Reference](https://kafkabackup.com/reference/config-yaml.md) - All storage options - [Backup Guide](https://kafkabackup.com/guides/backup-to-s3.md) - Backup walkthrough - [Performance Tuning](https://kafkabackup.com/guides/performance-tuning.md) - Optimize throughput --- title: Backup to S3 description: Step-by-step guide to backing up Kafka topics to Amazon S3 source_url: html: https://kafkabackup.com/guides/backup-to-s3 md: https://kafkabackup.com/guides/backup-to-s3.md --- # Backup to S3 This guide walks through setting up and running Kafka backups to Amazon S3. - OSO Kafka Backup installed - Kafka cluster accessible - AWS S3 bucket created ([AWS S3 Setup](https://kafkabackup.com/deployment/cloud-setup/aws-s3.md)) - AWS credentials configured ``` export AWS_ACCESS_KEY_ID="AKIA..." export AWS_SECRET_ACCESS_KEY="..." export AWS_REGION="us-west-2" ``` ``` # ~/.aws/credentials [default] aws_access_key_id = AKIA... aws_secret_access_key = ... # ~/.aws/config [default] region = us-west-2 ``` No configuration needed when running on EC2 with instance profile or EKS with IRSA. Create `s3-backup.yaml`: s3-backup.yaml ``` mode: backup backup_id: "production-daily" source: bootstrap_servers: - kafka-0.kafka.svc:9092 - kafka-1.kafka.svc:9092 - kafka-2.kafka.svc:9092 topics: include: - orders - payments - "events-*" exclude: - "__consumer_offsets" - "_schemas" storage: backend: s3 bucket: my-kafka-backups region: us-west-2 prefix: production/daily backup: compression: zstd compression_level: 3 checkpoint_interval_secs: 30 include_offset_headers: true source_cluster_id: "prod-us-west-2" ``` ``` kafka-backup backup --config s3-backup.yaml ``` With verbose logging: ``` kafka-backup -v backup --config s3-backup.yaml ``` ``` [INFO] Starting backup: production-daily [INFO] Connecting to Kafka cluster... [INFO] Connected to cluster: prod-us-west-2 [INFO] Storage: s3://my-kafka-backups/production/daily [INFO] Topics to backup: - orders (6 partitions) - payments (3 partitions) - events-clickstream (12 partitions) [INFO] Backing up topic: orders Partition 0: 150,234 records Partition 1: 148,892 records ... [INFO] Backup completed successfully [INFO] Summary: Topics: 3 Records: 1,234,567 Compressed Size: 128 MB Duration: 4m 23s S3 Objects: 45 ``` ``` kafka-backup list --path s3://my-kafka-backups/production/daily ``` ``` kafka-backup describe \ --path s3://my-kafka-backups/production/daily \ --backup-id production-daily ``` ``` # Quick validation kafka-backup validate \ --path s3://my-kafka-backups/production/daily \ --backup-id production-daily # Deep validation kafka-backup validate \ --path s3://my-kafka-backups/production/daily \ --backup-id production-daily \ --deep ``` ``` # List backup objects aws s3 ls s3://my-kafka-backups/production/daily/production-daily/ --recursive # Check manifest aws s3 cp s3://my-kafka-backups/production/daily/production-daily/manifest.json - | jq . ``` For subsequent backups, use a different backup\_id or timestamp: ``` backup_id: "production-daily-$(date +%Y%m%d)" ``` Or use continuous backup mode: ``` backup: continuous: true # Stream changes continuously start_offset: latest # Start from current position ``` > [!NOTE] > > [!NOTE] > > Offset tracking in continuous mode > > [!NOTE] > > Continuous backups track progress in a local SQLite database stored at `$TMPDIR/{backup_id}-offsets.db` (typically `/tmp`). This is periodically synced to your remote storage for durability. > > > > You can customize the path via `offset_storage.db_path`: > > > > ``` > > offset_storage: > > db_path: /data/offsets.db # Custom path (e.g. for K8s volume mounts) > > ``` > > > > **Kubernetes note:** If your pods use `readOnlyRootFilesystem: true`, you must mount `/tmp` as an `emptyDir` volume. See [Kubernetes Deployment](https://kafkabackup.com/deployment/kubernetes.md) for details. ``` # Daily at 2 AM 0 2 * * * /usr/local/bin/kafka-backup backup --config /etc/kafka-backup/s3-backup.yaml >> /var/log/kafka-backup/backup.log 2>&1 ``` ``` apiVersion: batch/v1 kind: CronJob metadata: name: kafka-backup-s3 spec: schedule: "0 2 * * *" jobTemplate: spec: template: spec: containers: - name: kafka-backup image: ghcr.io/osodevops/kafka-backup:latest args: ["backup", "--config", "/config/s3-backup.yaml"] env: - name: AWS_REGION value: us-west-2 # Use IRSA or mount credentials ``` Automatically transition old backups to cheaper storage: ``` aws s3api put-bucket-lifecycle-configuration \ --bucket my-kafka-backups \ --lifecycle-configuration file://lifecycle.json ``` Protect against accidental overwrites: ``` aws s3api put-bucket-versioning \ --bucket my-kafka-backups \ --versioning-configuration Status=Enabled ``` ``` aws s3api put-bucket-encryption \ --bucket my-kafka-backups \ --server-side-encryption-configuration '{ "Rules": [{"ApplyServerSideEncryptionByDefault": {"SSEAlgorithm": "AES256"}}] }' ``` ``` # Test credentials aws sts get-caller-identity # Test bucket access aws s3 ls s3://my-kafka-backups/ ``` - Check network bandwidth to S3 - Use a bucket in the same region as your Kafka cluster - Consider using S3 Transfer Acceleration - Increase segment size for fewer uploads - Check security groups and network ACLs - Verify VPC endpoint configuration (if using) - Check for proxy settings - [Point-in-Time Recovery](https://kafkabackup.com/guides/restore-pitr.md) - Restore from S3 backup - [Performance Tuning](https://kafkabackup.com/guides/performance-tuning.md) - Optimize backup speed - [Kubernetes Operator](https://kafkabackup.com/operator.md) - Automate S3 backups --- title: Incremental Backups description: Configure resumable incremental backups that only capture new data on each run source_url: html: https://kafkabackup.com/guides/incremental-backups md: https://kafkabackup.com/guides/incremental-backups.md --- # Incremental Backups By default, one-shot and snapshot backups start from the configured `start_offset` (typically `earliest`) on every run, producing a **full backup** each time. Starting from **v0.13.5**, you can make these backups **incremental** by adding `offset_storage` to your configuration. Each run then resumes from where the previous one stopped, backing up only new messages. OSO Kafka Backup tracks progress using a local SQLite database (the **offset store**). After each partition is backed up, the last consumed offset is saved. On the next run: 1. The offset store is loaded (from local disk or remote storage) 2. Each partition resumes from `last_saved_offset + 1` 3. Only new messages are fetched and written to storage 4. The manifest is **merged** with the existing one — segments from all runs are preserved 5. The updated offset store is synced back to remote storage This means the backup grows incrementally over time, and each run is fast because it only processes new data. Add the `offset_storage` section to your backup config: incremental-backup.yaml ``` mode: backup backup_id: "production-incremental" source: bootstrap_servers: - broker-1.kafka.svc:9092 topics: include: - "orders" - "payments" storage: backend: s3 bucket: kafka-backups region: us-west-2 prefix: incremental backup: compression: zstd stop_at_current_offsets: true # Exit after catching up include_offset_headers: true # This section enables incremental backups offset_storage: db_path: /data/offsets.db # Local path for SQLite database sync_interval_secs: 30 # Sync to remote storage every 30s ``` | Field | Default | Description | | --- | --- | --- | | `offset_storage.db_path` | `$TMPDIR/{backup_id}-offsets.db` | Local path for the SQLite offset database | | `offset_storage.sync_interval_secs` | `30` | How often the local DB is synced to remote storage | The offset database is also synced to remote storage at `{backup_id}/offsets.db`, so it survives pod restarts and machine changes. On startup, the engine checks remote storage for an existing offset database and loads it if the local one is empty. | Mode | Config | Incremental? | Exits? | Best for | | --- | --- | --- | --- | --- | | **Full one-shot** | `continuous: false` (default) | No — always starts from `start_offset` | Yes | Initial full backup, one-off snapshots | | **Incremental one-shot** | `continuous: false` + `offset_storage` | Yes — resumes from last checkpoint | Yes | Scheduled cron backups, hourly/daily DR | | **Snapshot** | `stop_at_current_offsets: true` | No (without `offset_storage`) | Yes | Consistent point-in-time snapshots | | **Incremental snapshot** | `stop_at_current_offsets: true` + `offset_storage` | Yes — resumes from last checkpoint | Yes | Scheduled incremental DR snapshots | | **Continuous** | `continuous: true` | Yes (automatic) | No | Streaming replication, zero-loss | A common pattern is running an incremental backup every hour via cron: ``` # Crontab entry 0 * * * * kafka-backup backup --config /etc/kafka-backup/incremental.yaml ``` **First run** (e.g. Monday 00:00): ``` [INFO] No saved offsets found, starting from earliest [INFO] Backed up 1,200,000 records across 3 topics [INFO] Offset database synced to s3://kafka-backups/incremental/production-incremental/offsets.db ``` **Second run** (Monday 01:00): ``` [INFO] Loaded offset database from remote storage [INFO] Resuming from saved offsets [INFO] orders:0 — starting at offset 450001 (saved: 450000) [INFO] Backed up 45,000 new records across 3 topics ``` **Third run** (Monday 02:00): ``` [INFO] Loaded offset database from remote storage [INFO] Resuming from saved offsets [INFO] Backed up 38,000 new records across 3 topics ``` Each subsequent run takes seconds instead of minutes, and the manifest accumulates all segments from every run. After running an incremental backup, you can verify the manifest contains segments from multiple runs: ``` kafka-backup describe \ --path s3://kafka-backups/incremental \ --backup-id production-incremental \ --format json ``` Look for multiple segments per partition with non-overlapping offset ranges — each segment corresponds to a single backup run's output. **If the offset database is lost locally:** The engine loads it from remote storage on startup. As long as the remote copy exists, no data is lost. **If both local and remote offset databases are lost:** The backup falls back to `start_offset` (default: `earliest`). This produces a full backup, but duplicate segments are deduplicated during manifest merging — existing segments with the same key or start offset are preserved, and the new run fills in any gaps. **If a backup fails mid-run:** Offsets are saved incrementally during the run (not only at the end). The next run picks up from the last checkpointed offset, so at most a few seconds of work is repeated. For Kubernetes deployments, enable incremental backups using the `checkpoint` section in the `KafkaBackup` CRD: ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: incremental-hourly spec: schedule: "0 0 * * * * *" stopAtCurrentOffsets: true checkpoint: enabled: true intervalSecs: 30 # ... rest of config ``` See [Scheduled Backups](https://kafkabackup.com/operator/guides/scheduled-backups.md) for complete examples. --- title: Retention and Erasure description: Bounded retention for backup sets and re-applying GDPR erasure requests on restore source_url: html: https://kafkabackup.com/guides/retention-and-erasure md: https://kafkabackup.com/guides/retention-and-erasure.md --- # Retention and Erasure A backup taken before a record was deleted still contains that record, and a restore reinstates it. For personal data this is the question every data protection officer asks: _what happens to an erasure request once the data is in a backup?_ This guide gives the answer kafka-backup supports today, in the order regulators expect it: bounded retention first, erasure re-applied on restore second, and removal from the archive itself as the enterprise layer. Both the UK ICO's right-to-erasure guidance and the EDPB's 2025/26 coordinated enforcement report on Article 17 accept that personal data may remain in a **passive, time-bounded backup** that is _put beyond use_, provided that: 1. the backup is not used to make decisions about, or otherwise affect, the individual; 2. it is deleted on its normal rotation — retention is bounded and documented; 3. the erasure is **re-applied if the backup is ever restored**, so the data never re-enters an active system. kafka-backup implements the mechanics for all three. What you document in your DPIA is the retention window, the restore procedure that re-applies erasure, and the evidence each restore produces. Backup sets grow: incremental and continuous backups append segments under one stable `backup_id` and rewrite `manifest.json` on every run. Two safe ways to bound them, both available since kafka-backup **0.21.0**: **Prune on a schedule** — plan-only by default, `--execute` to delete: ``` kafka-backup prune --config backup.yaml --older-than 30d # plan kafka-backup prune --config backup.yaml --older-than 30d --execute # delete ``` **Retention in the backup config** — applied at the end of every run: ``` backup: retention: max_age: 30d # delete segments older than this keep_segments: 1 # never remove a partition's newest segment ``` Either way the manifest is rewritten **before** any object is deleted, the removed offset range is recorded as a _pruned range_ (so `validate` stays green and `describe` shows what was removed and when), and a later run can never resurrect a pruned segment. The resume position of an incremental set is protected: segments the next run still needs are never pruned. > [!CAUTION] > > [!NOTE] > > Do not use bucket lifecycle rules on incremental sets > > [!NOTE] > > An age-based S3/Azure/GCS lifecycle rule deletes segments the manifest still references and never expires the manifest itself (it is rewritten every run). The result is a backup that reports healthy and fails at restore. Lifecycle rules are only safe when every run has its own `backup_id` (for example `backup_id: "daily-${DATE}"`). See the [storage guide](https://github.com/osodevops/kafka-backup/blob/main/docs/storage_guide.md#retention-use-prune-not-bucket-lifecycle-rules). A runnable walk-through is in the demos repository: [`cli/retention-prune`](https://github.com/osodevops/kafka-backup-demos/tree/main/cli/retention-prune). The restore engine has a per-record filter hook (`Keep | Drop | Tombstone`, kafka-backup 0.21.0+). The Enterprise binary uses it for **restore-time key suppression**: an erasure register — one erased record key per line — is re-applied while restoring, so records for those keys are produced as tombstones (compacted topics) or skipped (plain topics). The restore output reports the list's SHA-256 and the counts, which is the audit evidence a DPO needs: ``` Records suppressed: 2 (0 dropped, 2 tombstoned; keys_file=/etc/kafka-backup/erasure/suppressed-keys.txt, entries=1, sha256=d4f0…a9fa) ``` If the feature is not licensed, or the list cannot be read, the restore is **refused** — it never silently runs without the list. Configuration, keys-file format and the tombstone-versus-drop rule are on the [Enterprise erasure page](https://kafkabackup.com/enterprise/erasure.md); the demo [`cli/gdpr-erasure`](https://github.com/osodevops/kafka-backup-demos/tree/main/cli/gdpr-erasure) runs the whole flow under the built-in 14-day trial. **Without the Enterprise feature**, the equivalent manual procedure is: keep an erasure register (key + timestamp); after any restore and before consumers are unfrozen, re-produce tombstones for every key erased since the backup was taken. Compaction then removes them again. In-place redaction of an existing backup set (rewriting the affected segments and re-signing the evidence) and crypto-shredding via per-set data-encryption keys are the next erasure layers. They are designed so that a backup set is declared **immutable** (cyber-recovery, WORM) _or_ **erasable** — never both. Track progress on the [enterprise erasure issue](https://github.com/osodevops/kafka-backup-enterprise/issues/16). - **Retention window** for backups (`prune --older-than` / `retention.max_age`), and that lifecycle rules are not used on incremental sets. - **Restore procedure**: restores of topics containing personal data run with the erasure register applied (`enterprise.erasure.suppression`); the register is version-controlled and its SHA-256 is recorded with each restore ticket. - **Evidence**: the restore output (and `validate-restore --format json`, which carries the same summary under `enterprise.erasure_suppression`). - **Access control**: backups live in object storage with its own IAM; the backup is passive and never read by application code. - [Incremental Backups](https://kafkabackup.com/guides/incremental-backups.md) — how the manifest and offset store grow across runs. - [Validation and Compliance Evidence](https://kafkabackup.com/guides/validation-compliance.md) — signed evidence for backup and restore runs. - [Enterprise erasure](https://kafkabackup.com/enterprise/erasure.md) — the suppression feature. --- title: Point-in-Time Recovery (PITR) description: Restore Kafka data to a specific point in time using PITR source_url: html: https://kafkabackup.com/guides/restore-pitr md: https://kafkabackup.com/guides/restore-pitr.md --- # Point-in-Time Recovery (PITR) Restore Kafka data to any specific moment within your backup window with millisecond precision. Point-in-Time Recovery allows you to restore data from a backup to a specific timestamp, rather than restoring the entire backup. This is useful for: - **Disaster recovery**: Restore to just before a data corruption event - **Debugging**: Reproduce issues with data from a specific time - **Compliance**: Retrieve data as it existed at a particular moment - **Testing**: Create test environments with data from specific periods - A backup (offset headers are on by default; PITR itself does not need them — only header-based consumer offset recovery does) - Backup storage accessible - Target Kafka cluster available When you back up Kafka topics, each message includes its original timestamp. PITR uses these timestamps to filter which messages to restore. ``` Backup Timeline: ├── Nov 1, 00:00 ─ Backup started │ ├── Message 1 (timestamp: Nov 1, 00:00:01) │ ├── Message 2 (timestamp: Nov 1, 00:00:02) │ ├── ... │ ├── Message N (timestamp: Nov 15, 23:59:59) └── Nov 15, 23:59 ─ Backup ended PITR Window: Nov 10, 10:00 → Nov 10, 14:00 Result: Only messages with timestamps in this 4-hour window are restored ``` The window is applied to the **record timestamp stored by Kafka** — the value in each record's timestamp field. Whether that is the producer's clock (`CreateTime`, the default) or the broker's (`LogAppendTime`) depends on the topic's `message.timestamp.type`; kafka-backup does not distinguish the two. After a replay, mirror, or re-publish, that timestamp reflects when the record was (re)produced, not the business event time inside the payload. Filtering on a field **inside the payload** (an event-time attribute) is not supported by the OSS engine. If your event time lives in the payload, restore the widest window that covers the incident and let idempotent consumers filter by event id or event time; a payload-time filter built on the enterprise record-filter hook is on the roadmap. First, check the time range available in your backup: ``` kafka-backup describe \ --path s3://my-kafka-backups/production \ --backup-id production-backup-001 ``` Output: ``` Backup: production-backup-001 ═══════════════════════════════════════ Time Range: Earliest Message: 2024-11-01T00:00:00Z Latest Message: 2024-11-15T23:59:59Z Topics: orders 6 partitions 523,456 records payments 3 partitions 234,567 records ``` PITR uses Unix timestamps in **milliseconds**. Convert your target times: ``` # Convert ISO timestamp to Unix milliseconds date -d "2024-11-10T10:00:00Z" +%s000 # Output: 1731236400000 date -d "2024-11-10T14:00:00Z" +%s000 # Output: 1731250800000 ``` ``` from datetime import datetime start = datetime(2024, 11, 10, 10, 0, 0) end = datetime(2024, 11, 10, 14, 0, 0) print(f"Start: {int(start.timestamp() * 1000)}") print(f"End: {int(end.timestamp() * 1000)}") ``` | Window | Start Formula | Example | | --- | --- | --- | | Last hour | `now - 3600000` | Debugging recent issues | | Last 24 hours | `now - 86400000` | Daily incident recovery | | Specific day | Midnight to midnight | Compliance reporting | pitr-restore.yaml ``` mode: restore backup_id: "production-backup-001" target: bootstrap_servers: - kafka-dr-0.example.com:9092 - kafka-dr-1.example.com:9092 storage: backend: s3 bucket: my-kafka-backups region: us-west-2 prefix: production restore: # Point-in-Time Recovery window time_window_start: 1731236400000 # Nov 10, 2024 10:00:00 UTC time_window_end: 1731250800000 # Nov 10, 2024 14:00:00 UTC # Auto-create topics if they don't exist (v0.3.0+) create_topics: true default_replication_factor: 3 # Optional: Only restore specific topics # topics: # - orders # - payments # Optional: Remap topic names topic_mapping: orders: orders_pitr_restore payments: payments_pitr_restore # Add x-original-* / x-source-partition headers to restored records (default: false) include_original_offset_header: true # Or, for records identical to the source: strip_offset_headers: true (v0.19.0+) # Consumer offset handling consumer_group_strategy: skip ``` > [!TIP] > > [!NOTE] > > Auto Topic Creation > > [!NOTE] > > When using `topic_mapping` to restore to new topic names, enable `create_topics: true` to automatically create the target topics. Set `default_replication_factor` to match your cluster's requirements (typically 3 for production). Always validate PITR configuration first: ``` kafka-backup validate-restore --config pitr-restore.yaml ``` Output: ``` Restore Validation Report ═══════════════════════════════════════ Status: ✓ VALID PITR Time Window: Start: 2024-11-10T10:00:00Z (1731236400000) End: 2024-11-10T14:00:00Z (1731250800000) Duration: 4 hours Data to Restore: Segments matching: 23 of 156 Records in window: ~45,678 Estimated size: 12.3 MB Topics: orders → orders_pitr_restore (est. 23,456 records) payments → payments_pitr_restore (est. 22,222 records) ``` ``` kafka-backup restore --config pitr-restore.yaml ``` Output: ``` [INFO] Starting PITR restore from: production-backup-001 [INFO] Time window: 2024-11-10T10:00:00Z to 2024-11-10T14:00:00Z [INFO] Filtering records by timestamp... [INFO] Restoring topic: orders → orders_pitr_restore Records in window: 23,456 Restoring partition 0: 3,892 records Restoring partition 1: 3,901 records ... [INFO] PITR restore completed [INFO] Summary: Records Restored: 45,678 Skipped (outside window): 477,889 Duration: 2m 15s ``` ``` # Check topic was created kafka-topics --bootstrap-server kafka-dr:9092 --describe --topic orders_pitr_restore # Sample messages kafka-console-consumer \ --bootstrap-server kafka-dr:9092 \ --topic orders_pitr_restore \ --from-beginning \ --max-messages 5 \ --property print.timestamp=true ``` If you know an incident occurred at 14:32:15 UTC: ``` restore: # Restore up to 1 minute before the incident time_window_end: 1731251535000 # 14:32:15 - 60000 = 14:31:15 ``` ``` restore: time_window_start: 1731236400000 time_window_end: 1731250800000 source_partitions: - 0 - 1 - 2 # Only restore partitions 0, 1, 2 ``` ``` restore: time_window_start: 1731236400000 time_window_end: 1731250800000 partition_mapping: 0: 0 1: 0 # Merge partition 1 into partition 0 2: 1 # Move partition 2 to partition 1 ``` ``` restore: # From Nov 10 to latest available time_window_start: 1731236400000 # time_window_end: omitted = restore to end of backup ``` After a PITR restore, consumer groups need their offsets adjusted. Since only a subset of messages is restored, offsets will differ. ``` kafka-consumer-groups \ --bootstrap-server kafka-dr:9092 \ --group my-consumer-group \ --topic orders_pitr_restore \ --reset-offsets \ --to-earliest \ --execute ``` Restored messages carry their original offset in the `x-original-offset` header (little-endian `i64`) — archived by default at backup time, and re-added on restore with `include_original_offset_header: true`. Consumers can read this header to track position. Use the three-phase restore for automatic offset handling: ``` restore: time_window_start: 1731236400000 time_window_end: 1731250800000 consumer_group_strategy: header-based reset_consumer_offsets: true consumer_groups: - my-consumer-group ``` ``` kafka-backup three-phase-restore --config pitr-restore.yaml ``` PITR may be slower than full restore because: 1. All segments must be scanned to find matching records 2. Filtering adds CPU overhead 3. Smaller batches may be written - Use narrower time windows when possible - Increase `max_concurrent_partitions` for parallelism - Use faster storage (SSD) for backup data ``` Warning: No records found in specified time window ``` **Causes:** - Time window is outside backup range - Timestamps are in wrong format (seconds vs milliseconds) - Topic has no data in that period **Solution:** Check backup time range with `describe` command. Messages may have unexpected timestamps if: - Producers set custom timestamps - Messages were replicated with different timestamps - Clock skew between producers 1. **Always validate first** - Use `validate-restore` before executing 2. **Restore to new topics** - Use `topic_mapping` to avoid overwriting 3. **Document recovery points** - Note important timestamps for compliance 4. **Test regularly** - Practice PITR restores before you need them 5. **Monitor backup timestamps** - Ensure backups capture expected time ranges - [Offset Management](https://kafkabackup.com/guides/offset-management.md) - Handle consumer offsets after restore - [Three-Phase Restore](https://kafkabackup.com/reference/cli-reference.md#three-phase-restore) - Automated restore with offset reset - [Disaster Recovery Use Case](https://kafkabackup.com/use-cases/disaster-recovery.md) - DR planning with PITR --- title: Offset Management description: Manage consumer group offsets during backup and restore operations source_url: html: https://kafkabackup.com/guides/offset-management md: https://kafkabackup.com/guides/offset-management.md --- # Offset Management Managing consumer group offsets is critical when restoring Kafka data. This guide covers offset concepts and management strategies. Kafka offsets are sequential IDs assigned to each message in a partition: ``` Partition 0: ┌────┬────┬────┬────┬────┬────┬────┐ │ 0 │ 1 │ 2 │ 3 │ 4 │ 5 │ 6 │ ← Offsets └────┴────┴────┴────┴────┴────┴────┘ ↑ ↑ Earliest Latest (High Watermark) ``` Consumer groups track their position (committed offset) in each partition: ``` Consumer Group: order-processor orders-0: committed offset 5 (next to read: 6) orders-1: committed offset 3 (next to read: 4) orders-2: committed offset 7 (next to read: 8) ``` When data is restored, original offsets don't match the new offsets: ``` Original Cluster: Restored Cluster: ┌────┬────┬────┬────┐ ┌────┬────┬────┬────┐ │ 0 │ 1 │ 2 │ 3 │ │ 0 │ 1 │ 2 │ 3 │ └────┴────┴────┴────┘ └────┴────┴────┴────┘ Consumer at offset 2 Consumer still thinks offset 2 but data changed! ``` Don't modify offsets during restore. Handle manually afterward. ``` restore: consumer_group_strategy: skip ``` Best for: - Development/testing environments - When consumers will reprocess all data anyway - Manual offset management workflows Store original offset in message headers. Consumers read headers to track position. The backup already archives `x-original-offset` / `x-original-timestamp` by default (`backup.include_offset_headers: true`); the header-based strategy adds the restore-side set too (`include_original_offset_header` defaults to `false` but is implied by the strategy): ``` restore: consumer_group_strategy: header-based include_original_offset_header: true ``` Each restored message includes (binary little-endian values, not strings): ``` Headers: x-original-offset: 12345 (i64 LE, 8 bytes) x-original-timestamp: 1701234567890 (i64 LE, 8 bytes) x-source-partition: 0 (i32 LE, 4 bytes) ``` To restore _without_ any of these headers, set `restore.strip_offset_headers: true` (v0.19.0+) — see [Offset-tracking headers](https://kafkabackup.com/reference/config-yaml.md#offset-tracking-headers). Reset consumer offsets as part of the restore process. ``` restore: consumer_group_strategy: header-based reset_consumer_offsets: true consumer_groups: - order-processor - payment-service - analytics-consumer ``` After backup, view the offset mapping: ``` kafka-backup show-offset-mapping \ --path /data/backups \ --backup-id production-backup-001 ``` Output: ``` Offset Mapping: production-backup-001 ══════════════════════════════════════════════════════════════ Topic: orders ───────────────────────────────────────────────────────────── Partition First Offset Last Offset Record Count ───────────────────────────────────────────────────────────── 0 0 150233 150234 1 0 148891 148892 2 0 152456 152457 Topic: payments ───────────────────────────────────────────────────────────── Partition First Offset Last Offset Record Count ───────────────────────────────────────────────────────────── 0 0 78234 78235 1 0 76543 76544 ``` Export as CSV for analysis: ``` kafka-backup show-offset-mapping \ --path /data/backups \ --backup-id production-backup-001 \ --format csv > offset-mapping.csv ``` ``` # Reset to earliest kafka-consumer-groups \ --bootstrap-server kafka:9092 \ --group order-processor \ --topic orders \ --reset-offsets \ --to-earliest \ --execute # Reset to specific offset kafka-consumer-groups \ --bootstrap-server kafka:9092 \ --group order-processor \ --topic orders:0 \ --reset-offsets \ --to-offset 150233 \ --execute # Reset to timestamp kafka-consumer-groups \ --bootstrap-server kafka:9092 \ --group order-processor \ --topic orders \ --reset-offsets \ --to-datetime 2024-12-01T10:00:00.000 \ --execute ``` ``` kafka-backup offset-reset plan \ --path /data/backups \ --backup-id production-backup-001 \ --groups order-processor,payment-service \ --bootstrap-servers kafka:9092 ``` Output: ``` Offset Reset Plan ══════════════════════════════════════════════════════════════ Consumer Group: order-processor ───────────────────────────────────────────────────────────── Topic Partition Current Target Action ───────────────────────────────────────────────────────────── orders 0 - 150233 SET orders 1 - 148891 SET orders 2 - 152456 SET Consumer Group: payment-service ───────────────────────────────────────────────────────────── payments 0 - 78234 SET payments 1 - 76543 SET Total partitions to reset: 5 ``` ``` kafka-backup offset-reset execute \ --path /data/backups \ --backup-id production-backup-001 \ --groups order-processor \ --bootstrap-servers kafka:9092 ``` For review before execution: ``` kafka-backup offset-reset script \ --path /data/backups \ --backup-id production-backup-001 \ --groups order-processor \ --bootstrap-servers kafka:9092 \ --output reset-offsets.sh ``` Generated script: ``` #!/bin/bash # Offset reset script for backup: production-backup-001 # Generated: 2024-12-03T10:00:00Z # Stop consumers before running this script! kafka-consumer-groups --bootstrap-server kafka:9092 \ --group order-processor \ --topic orders:0 \ --reset-offsets --to-offset 150233 --execute kafka-consumer-groups --bootstrap-server kafka:9092 \ --group order-processor \ --topic orders:1 \ --reset-offsets --to-offset 148891 --execute # ... more commands ``` For large numbers of partitions, use bulk reset: ``` kafka-backup offset-reset-bulk \ --path /data/backups \ --backup-id production-backup-001 \ --groups order-processor,payment-service,analytics \ --bootstrap-servers kafka:9092 \ --max-concurrent 100 ``` Benefits: - ~50x faster than sequential reset - Per-partition retry with backoff - Detailed progress reporting - Performance metrics (p50/p99 latency) ``` kafka-backup offset-rollback snapshot \ --path /data/offset-snapshots \ --groups order-processor,payment-service \ --bootstrap-servers kafka:9092 \ --description "Before restore operation" ``` ``` kafka-backup offset-rollback list \ --path /data/offset-snapshots ``` Output: ``` Offset Snapshots ══════════════════════════════════════════════════════════════ ID Created Groups Description ────────────────────────────────────────────────────────────── snapshot-20241203-100000 2024-12-03T10:00:00Z 2 Before restore snapshot-20241203-090000 2024-12-03T09:00:00Z 2 Before migration ``` ``` kafka-backup offset-rollback show \ --path /data/offset-snapshots \ --snapshot-id snapshot-20241203-100000 ``` If something goes wrong: ``` kafka-backup offset-rollback rollback \ --path /data/offset-snapshots \ --snapshot-id snapshot-20241203-100000 \ --bootstrap-servers kafka:9092 ``` ``` kafka-backup offset-rollback verify \ --path /data/offset-snapshots \ --snapshot-id snapshot-20241203-100000 \ --bootstrap-servers kafka:9092 ``` For complete automated recovery: three-phase-restore.yaml ``` mode: restore backup_id: "production-backup-001" target: bootstrap_servers: - kafka-dr:9092 storage: backend: s3 bucket: my-kafka-backups region: us-west-2 restore: consumer_group_strategy: header-based reset_consumer_offsets: true consumer_groups: - order-processor - payment-service - analytics-consumer ``` ``` kafka-backup three-phase-restore --config three-phase-restore.yaml ``` Phases: 1. **Phase 1**: Collect offset headers from backup 2. **Phase 2**: Restore data to target cluster 3. **Phase 3**: Reset consumer group offsets 1. **Stop consumers** - Prevent offset commits during restore 2. **Snapshot current offsets** - Enable rollback if needed 3. **Document consumer groups** - Know which groups need reset 1. **Use header-based strategy** - Preserve original offset information 2. **Enable dry-run first** - Validate configuration 3. **Monitor progress** - Watch for errors 1. **Verify offset mapping** - Check offsets are set correctly 2. **Start consumers gradually** - Monitor for reprocessing issues 3. **Keep snapshots** - Retain for potential rollback ``` Error: Consumer group 'my-group' not found ``` The group may not exist if: - Consumers never started - Group was deleted - Using incorrect group name Solution: Create the group by starting a consumer, or use `--reset-offsets --to-earliest --dry-run` first. ``` Error: Offset 150000 is out of range for partition 0 (current: 0-100) ``` Target offset doesn't exist in restored topic. This can happen with PITR restores. Solution: Reset to `--to-earliest` or `--to-latest` instead of specific offset. ``` Error: Consumer group has active members ``` Stop all consumers before resetting offsets: ``` # Check active members kafka-consumer-groups \ --bootstrap-server kafka:9092 \ --group order-processor \ --describe --members ``` - [Three-Phase Restore](https://kafkabackup.com/reference/cli-reference.md#three-phase-restore) - Automated restore - [Performance Tuning](https://kafkabackup.com/guides/performance-tuning.md) - Optimize operations - [Disaster Recovery](https://kafkabackup.com/use-cases/disaster-recovery.md) - DR planning --- title: Performance Tuning description: Optimize OSO Kafka Backup throughput and resource usage source_url: html: https://kafkabackup.com/guides/performance-tuning md: https://kafkabackup.com/guides/performance-tuning.md --- # Performance Tuning Optimize backup and restore operations for maximum throughput and efficiency. OSO Kafka Backup is built in Rust for high performance: - **Throughput**: 100+ MB/s per partition - **Memory efficient**: Streaming processing, minimal buffering - **CPU efficient**: Native compilation, zero-copy where possible ``` kafka-backup -v backup --config backup.yaml ``` Output includes: ``` [INFO] Backup completed [INFO] Summary: Records: 2,456,789 Duration: 4m 23s Throughput: 9,346 records/sec Throughput: 45.2 MB/sec (uncompressed) Compression ratio: 4.9x ``` ``` kafka-backup status \ --path /data/backups \ --backup-id production-backup ``` Compression significantly impacts both storage costs and performance. | Algorithm | Speed | Ratio | Use Case | | --- | --- | --- | --- | | `none` | Fastest | 1x | Pre-compressed data | | `lz4` | Very Fast | 2-3x | Speed priority | | `zstd` (level 1-3) | Fast | 3-4x | Balanced (default) | | `zstd` (level 4-9) | Moderate | 4-6x | Storage priority | | `zstd` (level 10+) | Slow | 5-7x | Archive | ``` backup: compression: zstd compression_level: 3 # Balanced performance ``` Larger segments mean fewer storage operations but more memory usage: ``` backup: segment_max_bytes: 268435456 # 256 MB (larger segments) segment_max_interval_ms: 120000 # Roll every 2 minutes ``` | Segment Size | Pros | Cons | | --- | --- | --- | | 64 MB | Lower memory, frequent checkpoints | More S3 PUTs | | 128 MB | Balanced (default) | \- | | 256 MB | Fewer storage operations | Higher memory | | 512 MB | Minimal storage ops | High memory, delayed checkpoints | Control concurrent partition processing: ``` backup: max_concurrent_partitions: 8 # Default: 8 ``` Guidelines: - 2-4 partitions: Single-core systems - 4-8 partitions: Multi-core systems - 8-16 partitions: High-performance systems Control the delay between consumer poll attempts in continuous backup mode. Lower values reduce lag but increase CPU usage: ``` backup: poll_interval_ms: 100 # Default: 100ms ``` | Interval | Lag Impact | CPU Impact | Use Case | | --- | --- | --- | --- | | 50ms | Very low lag | Higher CPU | Real-time replication | | 100ms | Low lag (default) | Moderate | Balanced | | 500ms | Moderate lag | Low CPU | Resource-constrained | | 1000ms | Higher lag | Minimal CPU | Low-priority backups | Balance durability vs. performance: ``` backup: checkpoint_interval_secs: 60 # Less frequent = faster sync_interval_secs: 120 # Storage sync interval ``` ``` restore: max_concurrent_partitions: 8 # Increase for faster restore ``` ``` restore: produce_batch_size: 5000 # Records per produce batch (default: 1000) ``` Larger batches = higher throughput but more memory. Avoid overwhelming target cluster: ``` restore: rate_limit_records_per_sec: 100000 # Cap at 100k records/sec rate_limit_bytes_per_sec: 50000000 # Cap at 50 MB/sec ``` Set to `null` for unlimited. For local/NFS storage: - Use SSDs for backup storage - Use XFS or ext4 filesystems - Mount with `noatime` option - Ensure adequate I/O bandwidth ``` # Check I/O performance fio --name=test --rw=write --bs=128k --size=1G --runtime=30 ``` Optimize S3 performance: 1. **Use regional buckets** - Same region as Kafka cluster 2. **Enable Transfer Acceleration** - For cross-region backups 3. **Use larger segments** - Reduce PUT requests 4. **Use VPC endpoints** - Reduce latency ``` storage: backend: s3 bucket: my-kafka-backups region: us-west-2 # Same as Kafka cluster # endpoint: https://s3-accelerate.amazonaws.com # For acceleration ``` ``` storage: backend: azure container: kafka-backups # Use Premium block blob for high performance ``` ``` storage: backend: gcs bucket: my-kafka-backups # Use regional bucket in same region as Kafka ``` TCP keepalive and nodelay settings are critical for maintaining stable connections, especially with cloud-hosted Kafka services. ``` source: connection: tcp_keepalive: true # Prevent idle disconnections keepalive_time_secs: 60 # First probe after 60s idle keepalive_interval_secs: 20 # Probe every 20s thereafter tcp_nodelay: true # Reduce latency (disable Nagle) ``` | Setting | Impact | Use Case | | --- | --- | --- | | `tcp_keepalive: true` | Prevents connection drops | **Required** for Confluent Cloud | | `keepalive_time_secs: 30` | More aggressive keepalive | Unstable networks | | `keepalive_time_secs: 120` | Less overhead | Stable networks | | `tcp_nodelay: true` | Lower latency | Real-time backups | > [!WARNING] > > [!NOTE] > > Confluent Cloud > > [!NOTE] > > Confluent Cloud terminates idle connections after ~5 minutes. Always use TCP keepalive when connecting to Confluent Cloud to avoid "Broken pipe" errors. The backup process uses Kafka consumer settings internally. These are optimized by default: - `fetch.min.bytes`: 1 MB - `fetch.max.wait.ms`: 500 ms - `max.partition.fetch.bytes`: 10 MB Restore operations use optimized producer settings: - `batch.size`: 1 MB - `linger.ms`: 5 ms - `buffer.memory`: 64 MB ``` Memory ≈ (concurrent_partitions × segment_size) + overhead Example: 8 partitions × 128 MB = 1 GB segment buffers + 256 MB overhead = ~1.25 GB total ``` ``` backup: max_concurrent_partitions: 4 # Reduce parallelism segment_max_bytes: 67108864 # 64 MB segments ``` ``` resources: requests: memory: 1Gi cpu: 500m limits: memory: 2Gi cpu: 2000m ``` Compression is the main CPU consumer: | Compression | CPU Usage | | --- | --- | | none | Minimal | | lz4 | Low | | zstd level 1-3 | Moderate | | zstd level 4-9 | High | | zstd level 10+ | Very High | Increase parallelism to use more cores: ``` backup: max_concurrent_partitions: 8 # Use 8 cores ``` ``` Bandwidth = throughput × (1 + compression overhead) Example: 50 MB/s uncompressed data 4x compression ratio = ~12.5 MB/s to storage + Kafka read bandwidth ``` 1. **Enable compression** - Reduces storage bandwidth 2. **Use local storage** - Eliminate network for storage 3. **Same-region resources** - Kafka, backup, storage in same region ``` # Time a backup time kafka-backup backup --config backup.yaml # With metrics kafka-backup -v backup --config backup.yaml 2>&1 | tee backup.log grep -E "(Throughput|Duration|Records)" backup.log ``` ``` # Validate without executing (check estimated records) kafka-backup validate-restore --config restore.yaml # Time actual restore time kafka-backup restore --config restore.yaml ``` - Kafka cluster has available bandwidth - Storage has sufficient space and I/O - Network path is optimized (same region/VPC) - Compression level matches use case - Monitor Kafka consumer lag - Monitor storage write latency - Monitor memory usage - Monitor CPU usage - Verify backup integrity - Check compression ratio - Review throughput metrics - Optimize for next backup **Symptoms**: Low throughput, long duration **Causes and Solutions**: | Cause | Solution | | --- | --- | | High compression level | Reduce `compression_level` | | Low parallelism | Increase `max_concurrent_partitions` | | High poll interval | Reduce `poll_interval_ms` (default: 100) | | Slow storage | Use faster storage, increase segment size | | Network bottleneck | Use same-region storage | | Kafka throttling | Check broker load | **Symptoms**: Low restore throughput **Causes and Solutions**: | Cause | Solution | | --- | --- | | Target cluster overloaded | Add rate limiting, reduce parallelism | | Small batch size | Increase `produce_batch_size` | | Low parallelism | Increase `max_concurrent_partitions` | | Decompression overhead | Pre-decompress in pipeline | **Symptoms**: OOM errors, memory pressure **Solutions**: 1. Reduce `max_concurrent_partitions` 2. Reduce `segment_max_bytes` 3. Increase container/process memory limits ``` backup: compression: lz4 # Fast compression max_concurrent_partitions: 16 # High parallelism poll_interval_ms: 50 # Aggressive polling segment_max_bytes: 268435456 # 256 MB segments checkpoint_interval_secs: 120 # Less frequent checkpoints ``` ``` backup: compression: zstd compression_level: 3 max_concurrent_partitions: 8 poll_interval_ms: 100 segment_max_bytes: 134217728 # 128 MB checkpoint_interval_secs: 30 ``` ``` backup: compression: zstd compression_level: 1 # Minimal compression max_concurrent_partitions: 2 # Low parallelism poll_interval_ms: 500 # Less frequent polling segment_max_bytes: 67108864 # 64 MB segments checkpoint_interval_secs: 60 ``` - [Deployment Guide](https://kafkabackup.com/deployment.md) - Production deployment - [Configuration Reference](https://kafkabackup.com/reference/config-yaml.md) - All options - [Metrics Reference](https://kafkabackup.com/reference/metrics.md) - Monitor performance --- title: Security Setup description: Configure TLS, SASL, and encryption for OSO Kafka Backup source_url: html: https://kafkabackup.com/guides/security-setup md: https://kafkabackup.com/guides/security-setup.md --- # Security Setup Secure your Kafka backup operations with TLS encryption and SASL authentication. OSO Kafka Backup supports multiple security configurations: | Feature | Description | | --- | --- | | **TLS/SSL** | Encrypt data in transit | | **SASL** | Authenticate to Kafka | | **mTLS** | Mutual TLS authentication | | **Storage Encryption** | Encrypt data at rest | Encrypt communication without authentication: ``` source: bootstrap_servers: - kafka:9093 security: security_protocol: SSL ssl_ca_location: /certs/ca.crt ``` Mutual TLS for encryption and authentication: ``` source: bootstrap_servers: - kafka:9093 security: security_protocol: SSL ssl_ca_location: /certs/ca.crt ssl_certificate_location: /certs/client.crt ssl_key_location: /certs/client.key ssl_key_password: ${SSL_KEY_PASSWORD} # If key is encrypted ``` | File | Purpose | Format | | --- | --- | --- | | `ca.crt` | CA certificate to verify broker | PEM | | `client.crt` | Client certificate | PEM | | `client.key` | Client private key | PEM (PKCS#8) | Using OpenSSL: ``` # Generate CA openssl genrsa -out ca.key 4096 openssl req -new -x509 -days 365 -key ca.key -out ca.crt \ -subj "/CN=Kafka-CA" # Generate client key and CSR openssl genrsa -out client.key 4096 openssl req -new -key client.key -out client.csr \ -subj "/CN=kafka-backup" # Sign client certificate openssl x509 -req -days 365 -in client.csr -CA ca.crt -CAkey ca.key \ -CAcreateserial -out client.crt ``` Simple username/password authentication: ``` source: bootstrap_servers: - kafka:9092 security: security_protocol: SASL_PLAINTEXT # Or SASL_SSL sasl_mechanism: PLAIN sasl_username: backup-user sasl_password: ${KAFKA_PASSWORD} ``` SCRAM authentication using SHA-256 (RFC 5802). Uses PBKDF2 key derivation and mutual server signature verification to prevent MITM attacks: ``` source: bootstrap_servers: - kafka:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA256 sasl_username: backup-user sasl_password: ${KAFKA_PASSWORD} ssl_ca_location: /certs/ca.crt ``` SCRAM authentication using SHA-512. Same protocol as SCRAM-SHA-256 but with a stronger hash. Recommended for production: ``` source: bootstrap_servers: - kafka:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA512 sasl_username: backup-user sasl_password: ${KAFKA_PASSWORD} ssl_ca_location: /certs/ca.crt ``` Kerberos authentication via the GSSAPI mechanism (RFC 4752). Typical in enterprise environments that already run an MIT Kerberos KDC. > [!NOTE] > > [!NOTE] > > Opt-in feature > > [!NOTE] > > GSSAPI is gated behind the `gssapi` cargo feature. **Release binaries and the default Docker image do not include it.** You need to either build from source with `--features gssapi` or build a Docker image with `--build-arg FEATURES=gssapi`. See [Build & runtime prerequisites](#gssapi-prerequisites) below. kerberos-backup.yaml ``` source: bootstrap_servers: - kafka.prod.corp:9098 security: security_protocol: SASL_PLAINTEXT # or SASL_SSL if the broker advertises both sasl_mechanism: GSSAPI sasl_kerberos_service_name: kafka # must match the broker's service principal sasl_keytab_path: /etc/kafka-backup/client.keytab sasl_krb5_config_path: /etc/krb5.conf ``` The three GSSAPI-specific fields: | Field | Required | Description | | --- | --- | --- | | `sasl_kerberos_service_name` | Yes | The service name portion of the broker's service principal (e.g. `kafka` for `kafka/broker.host@REALM`). | | `sasl_keytab_path` | Yes | Absolute path to a keytab containing the client principal's long-term key. Read-only to the process. | | `sasl_krb5_config_path` | No | Override the default `/etc/krb5.conf`. Useful when packaging kafka-backup inside a container that doesn't share the host's krb5 config. | GSSAPI dynamically links against the system MIT Kerberos library — the same `libgssapi_krb5.so.2` your Kafka clients and brokers already load. **At build time** (the machine running `cargo build`): | Platform | Install | | --- | --- | | macOS | `brew install krb5` and `export PKG_CONFIG_PATH="$(brew --prefix krb5)/lib/pkgconfig:$PKG_CONFIG_PATH"` before building. Apple's bundled Heimdal does **not** expose the symbols `libgssapi` needs. | | Debian/Ubuntu | `apt-get install libkrb5-dev` | | Fedora/RHEL | `dnf install krb5-devel` | Then: ``` cargo build --release --features gssapi -p kafka-backup-cli ``` **At runtime** (the machine the binary runs on): | Platform | Package | | --- | --- | | Debian/Ubuntu | `libkrb5-3` (almost always already present if any Kerberos client is installed) | | RHEL/Fedora | `krb5-libs` | | Alpine | `krb5-libs` | If the host can already SASL-authenticate to Kafka with any Kerberos tool, `libgssapi_krb5.so.2` is already there. A keytab is a binary file holding the client principal and its long-term Kerberos key. Create and scope one per kafka-backup identity: ``` # On a machine with kadmin access to the KDC kadmin -p admin/admin > addprinc -randkey kafka-backup/host.corp@CORP.EXAMPLE > ktadd -k /tmp/kafka-backup.keytab kafka-backup/host.corp@CORP.EXAMPLE ``` Move the keytab to the backup host and lock it down: ``` install -m 0400 -o kafka-backup /tmp/kafka-backup.keytab \ /etc/kafka-backup/client.keytab ``` Grant the principal the Kafka ACLs listed under [ACL Requirements](#acl-requirements); Kerberos identities appear on the Kafka side as `User:kafka-backup/host.corp@CORP.EXAMPLE`. Kerberos service-principal matching is strict. The hostname in your `bootstrap_servers` must match the FQDN in the broker's service principal. If the broker runs as `kafka/broker.prod.corp@CORP.EXAMPLE`, your config must reach the broker via `broker.prod.corp` — not `localhost`, not a load-balancer DNS name, not a container alias. Mismatches surface as `KRB5KDC_ERR_S_PRINCIPAL_UNKNOWN`. Your `krb5.conf` needs the realm's KDC and its hostname resolution. Minimal example: /etc/krb5.conf ``` [libdefaults] default_realm = CORP.EXAMPLE rdns = false # disable reverse DNS — spn matching must be honest [realms] CORP.EXAMPLE = { kdc = kdc.corp.example:88 } [domain_realm] .corp.example = CORP.EXAMPLE corp.example = CORP.EXAMPLE ``` If the broker advertises a session lifetime (`connections.max.reauth.ms`), kafka-backup schedules a Kerberos re-handshake at ~80 % of that window (with ±5 s jitter and a 30 s floor) **before** the broker forcibly drops the connection. This happens transparently — no operator action required. You don't need to set anything. The scheduler is on the `SaslMechanismPlugin` trait that both built-in mechanisms and GSSAPI share. For dev or production hosts without crates.io access: ``` # On an internet-connected machine, from the kafka-backup repo root cargo vendor > .cargo/config-vendor.toml git add vendor .cargo/config-vendor.toml git commit -m "chore: vendor cargo deps for airgapped builds" # Transfer the repo to the airgapped host, pre-install libkrb5-dev from your # internal apt/rpm mirror, then build with: cargo build --release --features gssapi -p kafka-backup-cli --offline ``` The vendored `libgssapi`, `libgssapi-sys`, and all transitive crates live inside the repo. No network access is needed after `cargo vendor`. The only remaining system prerequisite is the krb5 development headers, which every enterprise Linux distribution ships. The default `oso/kafka-backup:` image does not include GSSAPI. Build a variant yourself: ``` FROM rust:1.82-slim AS builder RUN apt-get update && apt-get install -y --no-install-recommends \ libkrb5-dev pkg-config build-essential && \ rm -rf /var/lib/apt/lists/* WORKDIR /src COPY . . RUN cargo build --release --features gssapi -p kafka-backup-cli FROM debian:bookworm-slim RUN apt-get update && apt-get install -y --no-install-recommends \ libkrb5-3 libgssapi-krb5-2 ca-certificates && \ rm -rf /var/lib/apt/lists/* COPY --from=builder /src/target/release/kafka-backup /usr/local/bin/ ENTRYPOINT ["/usr/local/bin/kafka-backup"] ``` Then mount the keytab and `krb5.conf` at runtime: ``` docker run --rm \ -v /etc/krb5.conf:/etc/krb5.conf:ro \ -v /etc/kafka-backup/client.keytab:/etc/kafka-backup/client.keytab:ro \ -v $(pwd)/config:/config:ro \ my-registry/kafka-backup:gssapi \ backup --config /config/kerberos-backup.yaml ``` kafka-backup exposes a `SaslMechanismPlugin` Rust trait so downstream crates can add mechanisms (OAUTHBEARER, MSK IAM, custom) **without forking core**. GSSAPI is the first in-tree consumer; OAUTHBEARER and MSK IAM are expected community additions. If you are an operator, this section is informational — you do not need to interact with the trait directly. You configure a mechanism via `sasl_mechanism:` in YAML; the plugin is wired automatically. If you are building a downstream mechanism, see `examples/custom_sasl_plugin.rs` in the source repository for a minimal OAUTHBEARER reference implementation, and the [PRD at `docs/PRD-sasl-mechanism-plugin.md`](https://github.com/osodevops/kafka-backup/blob/main/docs/PRD-sasl-mechanism-plugin.md) for the trait contract (handshake dispatch, KIP-368 re-auth scheduler, RFC 7628 error parsing). | Protocol | Encryption | Authentication | | --- | --- | --- | | `PLAINTEXT` | No | No | | `SSL` | Yes (TLS) | Optional (mTLS) | | `SASL_PLAINTEXT` | No | Yes (SASL) | | `SASL_SSL` | Yes (TLS) | Yes (SASL) | | Mechanism | `sasl_mechanism:` value | Built-in | Feature gate | | --- | --- | --- | --- | | SASL/PLAIN | `PLAIN` | Yes | — | | SASL/SCRAM-SHA-256 | `SCRAM-SHA256` | Yes | — | | SASL/SCRAM-SHA-512 | `SCRAM-SHA512` | Yes | — | | SASL/GSSAPI (Kerberos) | `GSSAPI` | Opt-in | `--features gssapi` at build time | A CLI built without `--features gssapi` that loads a config with `sasl_mechanism: GSSAPI` fails fast at startup with a clear error pointing at the missing feature flag. kafka-credentials.yaml ``` apiVersion: v1 kind: Secret metadata: name: kafka-credentials namespace: kafka-backup type: Opaque stringData: username: backup-user password: your-secure-password ``` kafka-tls.yaml ``` apiVersion: v1 kind: Secret metadata: name: kafka-tls namespace: kafka-backup type: Opaque data: ca.crt: client.crt: client.key: ``` ``` spec: containers: - name: kafka-backup env: - name: KAFKA_USERNAME valueFrom: secretKeyRef: name: kafka-credentials key: username - name: KAFKA_PASSWORD valueFrom: secretKeyRef: name: kafka-credentials key: password volumeMounts: - name: tls mountPath: /certs readOnly: true volumes: - name: tls secret: secretName: kafka-tls ``` The operator handles secrets automatically: ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: secure-backup spec: kafkaCluster: bootstrapServers: - kafka:9093 securityProtocol: SASL_SSL tlsSecret: name: kafka-tls caKey: ca.crt saslSecret: name: kafka-credentials mechanism: SCRAM-SHA256 usernameKey: username passwordKey: password ``` Strimzi stores the cluster CA certificate and per-user client certificates in separate secrets. Use `caSecret` to reference the CA independently: ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: strimzi-backup spec: kafkaCluster: bootstrapServers: - my-cluster-kafka-bootstrap:9093 securityProtocol: SSL caSecret: name: my-cluster-cluster-ca-cert # Strimzi cluster CA caKey: ca.crt tlsSecret: name: my-kafka-user # Strimzi KafkaUser secret certKey: user.crt keyKey: user.key ``` This preserves Strimzi's automatic certificate rotation without requiring a manually combined secret. ``` storage: backend: s3 bucket: my-kafka-backups region: us-west-2 # S3 encrypts at rest automatically with SSE-S3 # Or configure SSE-KMS in bucket settings ``` Enable in bucket: ``` aws s3api put-bucket-encryption \ --bucket my-kafka-backups \ --server-side-encryption-configuration '{ "Rules": [{ "ApplyServerSideEncryptionByDefault": { "SSEAlgorithm": "aws:kms", "KMSMasterKeyID": "arn:aws:kms:us-west-2:123456789:key/12345" } }] }' ``` Azure encrypts all data at rest by default. For customer-managed keys: ``` az storage account update \ --name mystorageaccount \ --resource-group mygroup \ --encryption-key-source Microsoft.Keyvault \ --encryption-key-vault https://myvault.vault.azure.net \ --encryption-key-name mykey ``` GCS encrypts all data at rest by default. For customer-managed keys: ``` gcloud storage buckets update gs://my-bucket \ --default-encryption-key=projects/PROJECT_ID/locations/LOCATION/keyRings/KEY_RING/cryptoKeys/KEY ``` ``` # Allow reading from all topics kafka-acls --bootstrap-server kafka:9092 \ --add --allow-principal User:backup-user \ --operation Read --operation Describe \ --topic '*' # Allow reading consumer group offsets kafka-acls --bootstrap-server kafka:9092 \ --add --allow-principal User:backup-user \ --operation Read --operation Describe \ --group '*' # Optional: Cluster describe for metadata kafka-acls --bootstrap-server kafka:9092 \ --add --allow-principal User:backup-user \ --operation Describe \ --cluster ``` ``` # Allow writing to topics kafka-acls --bootstrap-server kafka:9092 \ --add --allow-principal User:restore-user \ --operation Write --operation Create --operation Describe \ --topic '*' # For offset reset kafka-acls --bootstrap-server kafka:9092 \ --add --allow-principal User:restore-user \ --operation Read --operation Write \ --group '*' ``` Secure credentials using environment variables: ``` export KAFKA_PASSWORD="your-password" export SSL_KEY_PASSWORD="key-password" export AWS_ACCESS_KEY_ID="AKIA..." export AWS_SECRET_ACCESS_KEY="..." ``` Reference in config: ``` source: security: sasl_password: ${KAFKA_PASSWORD} ssl_key_password: ${SSL_KEY_PASSWORD} storage: backend: s3 access_key: ${AWS_ACCESS_KEY_ID} secret_key: ${AWS_SECRET_ACCESS_KEY} ``` 1. **Never commit credentials** - Use secrets management 2. **Rotate credentials** - Regular rotation policy 3. **Least privilege** - Minimal required permissions 4. **Audit access** - Log and monitor credential usage 1. **Automate renewal** - Use cert-manager or similar 2. **Monitor expiry** - Alert before certificates expire 3. **Secure storage** - Protect private keys 4. **Use short-lived certs** - 90 days or less 1. **Use TLS** - Always encrypt in transit 2. **Private networks** - Keep Kafka on private subnets 3. **Firewall rules** - Restrict access to backup service 4. **VPC endpoints** - Use private connectivity to cloud storage ``` Error: SSL handshake failed ``` **Causes:** - Certificate mismatch - Expired certificate - Wrong CA certificate **Solution:** ``` # Verify certificate openssl s_client -connect kafka:9093 -CAfile ca.crt # Check expiry openssl x509 -in client.crt -noout -dates ``` ``` Error: SASL authentication failed ``` **Causes:** - Wrong username/password - User not created in Kafka - Wrong SASL mechanism **Solution:** ``` # Test with kafka-console-consumer kafka-console-consumer \ --bootstrap-server kafka:9092 \ --consumer.config client.properties \ --topic test ``` ``` Error: Not authorized to access topic ``` **Solution:** Check and update ACLs: ``` kafka-acls --bootstrap-server kafka:9092 \ --list --principal User:backup-user ``` ``` fatal error: 'gssapi.h' file not found failed to generate gssapi bindings ``` **Cause:** krb5 development headers are missing on the build host. Dynamic link needs the header files at compile time even though the `.so` ships in the runtime OS package. **Solution:** Install the development package for your distribution: ``` # Debian/Ubuntu sudo apt-get install libkrb5-dev # Fedora/RHEL sudo dnf install krb5-devel # macOS (Apple Heimdal does not work — must be MIT krb5) brew install krb5 export PKG_CONFIG_PATH="$(brew --prefix krb5)/lib/pkgconfig:$PKG_CONFIG_PATH" ``` ``` Error: acquire_cred failed: KRB5KDC_ERR_S_PRINCIPAL_UNKNOWN ``` **Cause:** the hostname you're connecting to doesn't match the broker's service principal. If the broker runs as `kafka/broker.prod.corp@CORP.EXAMPLE` and you connect via `localhost:9098` or a load balancer DNS name, the KDC refuses to issue a ticket for the service. **Solution:** connect via the exact FQDN that appears in the broker's service principal. Verify: ``` kvno kafka/broker.prod.corp@CORP.EXAMPLE ``` If you're testing against a local Docker fixture, add a `/etc/hosts` entry pointing that FQDN at `127.0.0.1`. ``` Error: Authentication failed due to invalid credentials with SASL mechanism GSSAPI ``` **Cause:** MIT Kerberos preferred a stale ticket in the OS default credential cache over a fresh TGT from your keytab. Most common on macOS, where the default cache is `API:` and persists across logins — and when the broker's service key has rotated since that cache was populated, AP-REQ decryption fails. **Solution:** kafka-backup isolates its credential cache automatically when a keytab is configured (`KRB5CCNAME=MEMORY:` per plugin instance). If you still hit this, confirm: - The keytab you configured is fresh (`klist -kte /path/to/keytab` — check the KVNO column matches what the KDC knows about your principal). - Nothing on the host is `kinit`\-ing the same principal into the default ccache with a stale key. - `rdns = false` is set in your `krb5.conf` — reverse-DNS fuzz on SPN matching is a common silent failure. ``` Error: sasl_mechanism: GSSAPI is configured but this binary was built without the `gssapi` feature ``` **Cause:** you're running the default (non-gssapi) CLI against a Kerberized broker. **Solution:** rebuild with the feature (`cargo build --release --features gssapi -p kafka-backup-cli`) or use a Docker image built with `--build-arg FEATURES=gssapi`. See [Build & runtime prerequisites](#gssapi-prerequisites). secure-backup.yaml ``` mode: backup backup_id: "secure-production-backup" source: bootstrap_servers: - kafka-0.kafka.svc:9093 - kafka-1.kafka.svc:9093 - kafka-2.kafka.svc:9093 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA512 sasl_username: backup-service sasl_password: ${KAFKA_PASSWORD} ssl_ca_location: /certs/ca.crt ssl_certificate_location: /certs/client.crt ssl_key_location: /certs/client.key topics: include: - "*" exclude: - "__consumer_offsets" storage: backend: s3 bucket: secure-kafka-backups region: us-west-2 # Uses IAM role - no static credentials backup: compression: zstd compression_level: 3 include_offset_headers: true ``` For enterprises running an MIT Kerberos KDC. Assumes the CLI was built with `--features gssapi` and `libkrb5-3` is installed on the runtime host. kerberos-backup.yaml ``` mode: backup backup_id: "kerberos-production-backup" source: bootstrap_servers: - broker-0.prod.corp:9098 - broker-1.prod.corp:9098 - broker-2.prod.corp:9098 security: security_protocol: SASL_PLAINTEXT # SASL_SSL if broker advertises TLS too sasl_mechanism: GSSAPI sasl_kerberos_service_name: kafka sasl_keytab_path: /etc/kafka-backup/client.keytab sasl_krb5_config_path: /etc/krb5.conf topics: include: - "*" exclude: - "__consumer_offsets" storage: backend: s3 bucket: secure-kafka-backups region: us-west-2 backup: compression: zstd compression_level: 3 include_offset_headers: true ``` - [Deployment Guide](https://kafkabackup.com/deployment.md) - Production deployment - [Kubernetes Operator](https://kafkabackup.com/operator.md) - Secure K8s setup - [Enterprise CSFLE Metadata Backup](https://kafkabackup.com/enterprise/encryption.md) - Preserve Confluent KEKs, DEKs, and schema encryption rules --- title: Monitoring Setup description: Scrape OSO Kafka Backup with Prometheus and visualize it in Grafana source_url: html: https://kafkabackup.com/guides/monitoring-setup md: https://kafkabackup.com/guides/monitoring-setup.md --- # Monitoring Setup OSO Kafka Backup serves Prometheus metrics while a backup or restore is running. This guide creates a small local Prometheus and Grafana stack and then shows the Kubernetes setup. Add a `metrics` block to the backup or restore configuration: ``` mode: backup backup_id: production-backup source: bootstrap_servers: ["kafka:9092"] topics: include: ["orders", "events"] storage: backend: filesystem path: /var/lib/kafka-backup/data backup: stop_at_current_offsets: true metrics: enabled: true port: 8080 bind_address: "0.0.0.0" path: /metrics update_interval_ms: 500 keep_alive_seconds: 60 max_partition_labels: 100 ``` `keep_alive_seconds` lets Prometheus collect final values after a one-shot operation completes. Set it to at least twice your scrape interval. ``` kafka-backup backup --config backup.yaml curl --fail http://127.0.0.1:8080/metrics ``` Download [compose.yaml](https://kafkabackup.com/assets/files/compose-389258835834b535a22e9836774a974c.md) and [prometheus.yml](https://kafkabackup.com/assets/files/prometheus-16af523dd3f4dcc7a79fb5a5a18f518e.md), or create the following files. First, `compose.yaml`: ``` services: prometheus: image: prom/prometheus:v2.54.1 command: - --config.file=/etc/prometheus/prometheus.yml ports: - "9090:9090" volumes: - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro - prometheus-data:/prometheus extra_hosts: - host.docker.internal:host-gateway grafana: image: grafana/grafana:11.2.0 environment: GF_SECURITY_ADMIN_USER: admin GF_SECURITY_ADMIN_PASSWORD: admin GF_USERS_ALLOW_SIGN_UP: "false" ports: - "3000:3000" volumes: - grafana-data:/var/lib/grafana depends_on: - prometheus volumes: prometheus-data: {} grafana-data: {} ``` Create `prometheus.yml` beside it: ``` global: scrape_interval: 30s scrape_configs: - job_name: kafka-backup metrics_path: /metrics static_configs: - targets: ["host.docker.internal:8080"] ``` Start the stack: ``` docker compose up -d docker compose ps ``` - Prometheus: [http://localhost:9090](http://localhost:9090) - Grafana: [http://localhost:3000](http://localhost:3000), using `admin` / `admin` On Linux, `host-gateway` maps `host.docker.internal` to the host. Docker Desktop provides the same hostname on macOS and Windows. In Grafana, add `http://prometheus:9090` as a Prometheus data source. Finite backup progress: ``` 100 * ( 1 - kafka_backup_snapshot_records_remaining / clamp_min(kafka_backup_snapshot_records_target, 1) ) ``` Aggregate lag: ``` kafka_backup_lag_records_sum ``` Backup throughput: ``` rate(kafka_backup_records_total[5m]) ``` Storage write latency: ``` histogram_quantile( 0.99, sum by (le, backend, operation) ( rate(kafka_backup_storage_write_latency_seconds_bucket[5m]) ) ) ``` Errors: ``` sum by (backup_id, error_type) ( increase(kafka_backup_errors_total[5m]) ) ``` See the [metrics reference](https://kafkabackup.com/reference/metrics.md) for the complete list and label sets. The two operators expose metrics differently: - The [OSO Kafka Backup Operator](https://kafkabackup.com/operator/metrics.md) runs the core library in its own process and exposes controller/completion metrics from one Service. - The [Strimzi Backup Operator](https://kafkabackup.com/strimzi-operator/metrics.md) creates separate Job pods. Enable both its ServiceMonitor and Job PodMonitor to collect controller health and live backup progress. For short-lived Jobs, a PodMonitor is preferable to a ServiceMonitor because it discovers the pods directly. kafka-backup-alerts.yaml ``` groups: - name: kafka-backup rules: - alert: KafkaBackupStalled expr: >- increase(kafka_backup_records_total[10m]) == 0 and kafka_backup_lag_records_sum > 0 for: 15m labels: severity: critical annotations: summary: Kafka backup is not making progress - alert: KafkaBackupErrors expr: increase(kafka_backup_errors_total[5m]) > 0 for: 1m labels: severity: warning annotations: summary: Kafka backup emitted errors ``` Metrics from a one-shot process are ephemeral. Use Job/CR status or an external batch-success metric for durable last-success alerting. --- title: Backup Validation & Compliance Evidence description: Step-by-step guide to validating Kafka backups and generating signed compliance evidence reports source_url: html: https://kafkabackup.com/guides/validation-compliance md: https://kafkabackup.com/guides/validation-compliance.md --- # Backup Validation & Compliance Evidence Automatically validate that your Kafka backups can be restored correctly and generate cryptographically signed evidence reports for auditors. The validation suite runs checks against a restored Kafka cluster and compares the results against the original backup manifest. It produces: - **JSON evidence reports** — machine-readable, deterministic, suitable for automation - **PDF evidence reports** — auditor-ready, branded, suitable for direct submission - **Detached signatures** — ECDSA-P256-SHA256 cryptographic proof of report integrity - **Compliance mappings** — automatic mapping to SOX ITGC, CMMC RE.3.139, and GDPR Article 32 - OSO Kafka Backup installed (v0.11.0+) - An existing backup in object storage or filesystem - A Kafka cluster with the backup data restored (the "target" cluster) - Optional: OpenSSL for generating signing keys > [!NOTE] > > [!NOTE] > > info > > [!NOTE] > > The validation tool does **not** perform the restore itself. Run `kafka-backup restore` first, then validate the result. This separation ensures the validation is an independent check. Create `validation.yaml`: validation.yaml ``` # Backup to validate against backup_id: "production-daily-001" # Where the backup is stored storage: backend: s3 bucket: my-kafka-backups region: us-west-2 prefix: production/daily # The restored Kafka cluster to validate target: bootstrap_servers: - restored-kafka:9092 # Which checks to run checks: message_count: enabled: true mode: exact # exact | sample offset_range: enabled: true consumer_group_offsets: enabled: false # Enable if consumer groups were restored # Evidence report settings evidence: formats: - json - pdf storage: prefix: "evidence-reports/" retention_days: 2555 # ~7 years (SOX requirement) ``` ``` $ kafka-backup validation run --config validation.yaml ``` You'll see output like: ``` === Validation Results === Overall: PASSED Checks: 2/2 passed, 0 failed, 0 skipped Duration: 18ms [PASSED] MessageCountCheck — 3 topics; 1000 messages expected, 1000 restored; 0 discrepancies [PASSED] OffsetRangeCheck — 9 partitions checked; 9 passed; 0 issues JSON evidence uploaded: evidence-reports/validation-9275b4aa/2026/04/validation-9275b4aa.json PDF evidence uploaded: evidence-reports/validation-9275b4aa/2026/04/validation-9275b4aa.pdf ``` The command exits with code 0 on success, code 1 if any check fails. When an auditor requests a specific point-in-time validation: ``` $ kafka-backup validation run \ --config validation.yaml \ --pitr 1711929600000 \ --triggered-by "KPMG Q1 2026 audit" ``` The `--triggered-by` string is recorded in the evidence report, providing a clear chain of custody. The JSON evidence report contains: - **Backup metadata** — ID, source cluster, topics, partitions, record counts - **Validation results** — per-check pass/fail with machine-readable data - **Integrity information** — SHA-256 checksums, signature algorithm - **Compliance mappings** — which checks satisfy which regulatory controls ``` # List available evidence reports $ kafka-backup validation evidence-list --path s3://my-kafka-backups # Download a specific report $ kafka-backup validation evidence-get \ --path s3://my-kafka-backups \ --report-id validation-9275b4aa \ --format json \ --output evidence-report.json ``` Include `pdf` in the formats list: ``` evidence: formats: - json - pdf ``` The PDF contains: - **Page 1** — Cover page with overall result (PASSED/FAILED), report ID, timestamp - **Page 2** — Validation check results table - **Page 3** — Integrity details and compliance framework mappings (SOX, CMMC, GDPR) See the [Evidence Signing Guide](https://kafkabackup.com/guides/evidence-signing.md) for detailed key management instructions. Quick setup: ``` # Generate an ECDSA-P256 key pair $ openssl ecparam -genkey -name prime256v1 -noout | \ openssl pkcs8 -topk8 -nocrypt -out signing-key.pem $ openssl ec -in signing-key.pem -pubout -out signing-key-pub.pem ``` Add to your config: validation.yaml ``` evidence: signing: enabled: true private_key_path: "/etc/kafka-backup/signing-key.pem" ``` The signed report produces a `.sig` file alongside the JSON and PDF. ``` $ kafka-backup validation evidence-verify \ --report evidence-report.json \ --signature evidence-report.sig \ --public-key signing-key-pub.pem ``` ``` Report ID: validation-9275b4aa-2aeb-4910-a3a6-9e4aa1dc016a Algorithm: ECDSA-P256-SHA256 Report SHA-256: 2482bbdfa113146e39a4884767002554... SHA-256 checksum: VALID ECDSA signature: VALID Evidence report integrity: VERIFIED ``` Get alerted when validation passes or fails: validation.yaml ``` notifications: slack: webhook_url: "https://hooks.slack.com/services/T00/B00/xxxxx" pagerduty: integration_key: "your-pagerduty-integration-key" severity: critical # Triggers on failure only ``` Slack receives a Block Kit message with the result, check summary, and a link to the evidence report. PagerDuty receives an Events API v2 trigger on failure and auto-resolves on the next success. Compares per-partition record counts between the backup manifest and the restored cluster. Fails if any partition has more discrepancies than `fail_threshold`. ``` checks: message_count: enabled: true mode: exact # exact: all partitions | sample: random subset sample_percentage: 100 topics: [] # Empty = all topics in the backup fail_threshold: 0 # 0 = fail on any discrepancy ``` Verifies that the high watermark and low watermark for each partition in the restored cluster match the backup manifest's segment offset ranges. ``` checks: offset_range: enabled: true verify_high_watermark: true verify_low_watermark: true ``` Verifies that consumer group offsets are present and valid in the restored cluster. ``` checks: consumer_group_offsets: enabled: true verify_all_groups: true # false = only check groups listed below groups: [] # Empty + verify_all_groups = all groups ``` Call your own validation endpoint. The tool POSTs a JSON payload with the backup ID and restored cluster details, and expects a pass/fail response. ``` checks: custom_webhooks: - name: application-health-check url: "https://internal.example.com/kafka-validation-hook" timeout_seconds: 120 expected_status_code: 200 fail_on_timeout: true ``` validation-full.yaml ``` backup_id: "production-daily-001" storage: backend: s3 bucket: my-kafka-backups region: us-west-2 prefix: production/daily target: bootstrap_servers: - restored-kafka-0:9092 - restored-kafka-1:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA-512 sasl_username: backup-validator sasl_password: "${KAFKA_PASSWORD}" checks: message_count: enabled: true mode: exact fail_threshold: 0 offset_range: enabled: true consumer_group_offsets: enabled: true verify_all_groups: true custom_webhooks: - name: order-service-check url: "https://internal.example.com/validation/orders" timeout_seconds: 120 evidence: formats: [json, pdf] signing: enabled: true private_key_path: "/etc/kafka-backup/signing-key.pem" storage: prefix: "evidence-reports/" retention_days: 2555 notifications: slack: webhook_url: "https://hooks.slack.com/services/T00/B00/xxxxx" pagerduty: integration_key: "abc123def456" severity: critical triggered_by: "weekly-cron-job" ``` - [Evidence Signing Guide](https://kafkabackup.com/guides/evidence-signing.md) — deep-dive on key management and PKI integration - [Validation Config Reference](https://kafkabackup.com/reference/validation-config.md) — complete option reference - [Evidence Report Schema](https://kafkabackup.com/reference/evidence-report-schema.md) — JSON schema documentation - [SOX Compliance Example](https://kafkabackup.com/examples/compliance-evidence-sox.md) — end-to-end SOX scenario - [GDPR Compliance Example](https://kafkabackup.com/examples/compliance-evidence-gdpr.md) — GDPR Article 32 scenario - [Compliance Evidence Use Cases](https://kafkabackup.com/use-cases/compliance-evidence.md) — why and when to use this --- title: Evidence Report Signing description: Generate, sign, and verify cryptographic evidence reports with ECDSA-P256-SHA256 source_url: html: https://kafkabackup.com/guides/evidence-signing md: https://kafkabackup.com/guides/evidence-signing.md --- # Evidence Report Signing Cryptographically sign evidence reports to prove they have not been tampered with after generation. Auditors can independently verify the signature using the public key. Unsigned evidence is an assertion. Signed evidence is proof. | Without signing | With signing | | --- | --- | | Auditor must trust the operator produced the report | Auditor can verify the report independently | | Reports can be modified after generation | Any modification invalidates the signature | | No chain of custody | Signing key binds report to a specific identity | | May not satisfy SOX tamper-evidence requirements | SHA-256 checksums accepted by SOX auditors | ``` 1. Run validation checks → collect results 2. Serialize to canonical JSON (sorted keys, no whitespace) 3. Compute SHA-256 of canonical JSON → report_sha256 4. Sign the canonical JSON with ECDSA-P256 private key → signature 5. Store: report.json + report.pdf + report.sig ``` The `.sig` file is a detached signature — the JSON report and signature are stored separately. This allows auditors to hash the JSON file themselves and verify it matches. ``` # Generate an ECDSA P-256 private key in PKCS#8 format $ openssl ecparam -genkey -name prime256v1 -noout | \ openssl pkcs8 -topk8 -nocrypt -out signing-key.pem # Extract the public key $ openssl ec -in signing-key.pem -pubout -out signing-key-pub.pem ``` > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > The private key must be in **PKCS#8** format (begins with `-----BEGIN PRIVATE KEY-----`). If your key begins with `-----BEGIN EC PRIVATE KEY-----`, convert it: > > > > ``` > > $ openssl pkcs8 -topk8 -nocrypt -in ec-key.pem -out pkcs8-key.pem > > ``` | Environment | Recommendation | | --- | --- | | Development | Local file with restricted permissions (`chmod 600`) | | Docker/VM | Mount as a read-only volume | | Kubernetes | Kubernetes Secret or external secret operator (Vault, AWS Secrets Manager) | | CI/CD | Injected as a pipeline secret, never committed to git | Add to your `validation.yaml`: validation.yaml ``` evidence: signing: enabled: true private_key_path: "/etc/kafka-backup/signing-key.pem" public_key_path: "/etc/kafka-backup/signing-key-pub.pem" # Optional ``` When signing is enabled, the stored JSON report uses canonical (compact) serialization to ensure the stored bytes exactly match what was signed. ``` $ kafka-backup validation evidence-verify \ --report evidence-report.json \ --signature evidence-report.sig \ --public-key signing-key-pub.pem ``` Output: ``` Report ID: validation-9275b4aa-2aeb-4910-a3a6-9e4aa1dc016a Algorithm: ECDSA-P256-SHA256 Report SHA-256: 2482bbdfa113146e39a4884767002554a622ecc573f16f74e183a532de94590a SHA-256 checksum: VALID ECDSA signature: VALID Evidence report integrity: VERIFIED ``` Auditors can verify without the kafka-backup CLI: ``` # 1. Compute SHA-256 of the JSON report $ sha256sum evidence-report.json 2482bbdfa113146e39a4884767002554... evidence-report.json # 2. Compare against the SHA-256 in the .sig file $ cat evidence-report.sig -----BEGIN KAFKA BACKUP EVIDENCE SIGNATURE----- Algorithm: ECDSA-P256-SHA256 Report-ID: validation-9275b4aa Report-SHA256: 2482bbdfa113146e39a4884767002554... Signature: MEUCIQDk... -----END KAFKA BACKUP EVIDENCE SIGNATURE----- # 3. If SHA-256 matches, the report has not been modified ``` The `.sig` file uses a simple text format: ``` -----BEGIN KAFKA BACKUP EVIDENCE SIGNATURE----- Algorithm: ECDSA-P256-SHA256 Report-ID: weekly-compliance-check-20260406-020000 Report-SHA256: e3b0c44298fc1c149afb4c8996fb924... Signature: MEUCIQDk8nGq... (base64-encoded ECDSA signature) -----END KAFKA BACKUP EVIDENCE SIGNATURE----- ``` | Field | Description | | --- | --- | | `Algorithm` | Always `ECDSA-P256-SHA256` | | `Report-ID` | Unique validation run identifier | | `Report-SHA256` | SHA-256 hex digest of the canonical JSON report | | `Signature` | Base64-encoded ECDSA signature over the canonical JSON bytes | If your cluster uses cert-manager, you can use a Certificate resource to provision signing keys: ``` apiVersion: cert-manager.io/v1 kind: Certificate metadata: name: kafka-backup-signing spec: secretName: kafka-backup-signing-key issuerRef: name: internal-ca kind: ClusterIssuer privateKey: algorithm: ECDSA size: 256 usages: - digital signature ``` Mount the secret into the kafka-backup pod and reference the key path in your validation config. ``` # Store the signing key in Vault $ vault kv put secret/kafka-backup/signing private_key=@signing-key.pem # Retrieve at runtime $ vault kv get -field=private_key secret/kafka-backup/signing > /tmp/signing-key.pem ``` The private key is not in PKCS#8 format. Convert it: ``` $ openssl pkcs8 -topk8 -nocrypt -in your-key.pem -out pkcs8-key.pem ``` - Ensure the JSON report file has not been reformatted (the bytes must match exactly what was signed) - Verify you are using the correct public key that corresponds to the private key used for signing The report file was modified after signing. This could mean: - The file was pretty-printed or reformatted - The file was opened and saved by an editor that changed whitespace - The file was genuinely tampered with - [Backup Validation Guide](https://kafkabackup.com/guides/validation-compliance.md) — complete validation walkthrough - [Validation Config Reference](https://kafkabackup.com/reference/validation-config.md) — all configuration options - [SOX Compliance Example](https://kafkabackup.com/examples/compliance-evidence-sox.md) — end-to-end signed evidence for SOX --- title: MSK KRaft Migration Runbook description: Step-by-step production runbook for migrating AWS MSK from ZooKeeper to KRaft with kafka-backup Enterprise. source_url: html: https://kafkabackup.com/guides/msk-kraft-migration-runbook md: https://kafkabackup.com/guides/msk-kraft-migration-runbook.md --- # MSK KRaft Migration Runbook This runbook walks through a complete production migration from an AWS MSK ZooKeeper cluster to a KRaft cluster using kafka-backup Enterprise. Before starting, ensure you have: - **Source cluster**: MSK Provisioned cluster in ZooKeeper mode (Kafka 2.8+) - **Target cluster**: MSK Provisioned cluster in KRaft mode (Kafka 3.7+), pre-provisioned with the same or larger broker count - **S3 buckets**: Two buckets in the same region as your clusters — one for migration segments, one for evidence - **IAM permissions**: The migration runner needs access to both MSK clusters and both S3 buckets (see [Step 2](#step-2-generate-the-migration-plan)) - **kafka-backup Enterprise**: Installed on a host that can reach both clusters and S3 ([Installation Guide](https://kafkabackup.com/enterprise/installation.md)) - **License**: Enterprise license with `migrations:msk-kraft` feature, or the 14-day free trial (activates automatically) - **Network**: The migration runner must have network access to both clusters' bootstrap servers > [!TIP] > > [!NOTE] > > Pre-flight with free tools > > [!NOTE] > > You can validate all prerequisites without a license. `plan` generates the IAM policies you need, and `precheck` verifies network connectivity, S3 access, and cluster compatibility — all free. Create a YAML configuration file that describes your source, target, and migration parameters. migration.yaml ``` enterprise: msk_kraft_migration: source: cluster_arn: arn:aws:kafka:us-east-1:123456789012:cluster/prod-zk-cluster/abc-def-123 auth: mode: iam target: cluster_arn: arn:aws:kafka:us-east-1:123456789012:cluster/prod-kraft-cluster/ghi-jkl-456 auth: mode: iam backup: s3_bucket: prod-migration-segments s3_prefix: zk-to-kraft/ evidence: s3_bucket: prod-migration-evidence s3_prefix: migrations/ retention: 7y cutover: drain_timeout: 30m drain_max_partition_lag: 100 drain_stable_window: 30s max_producer_freeze: 60s producer_freeze_webhook: https://internal-api.example.com/kafka/freeze validation: count_tolerance: 1 spot_check_records_per_partition: 3 acl: on_drift: merge ``` See the [Configuration Reference](https://kafkabackup.com/enterprise/msk-kraft-config-reference.md) for every field and its defaults. **IAM** (most common): ``` auth: mode: iam ``` **SCRAM-SHA-512**: ``` auth: mode: scram-sha-512 username: ${KAFKA_USERNAME} password: ${KAFKA_PASSWORD} ``` **mTLS**: ``` auth: mode: mtls keystore: /path/to/keystore.jks keystore_password: ${KEYSTORE_PASS} truststore: /path/to/truststore.jks truststore_password: ${TRUSTSTORE_PASS} ``` ``` kafka-backup migrate msk-kraft plan \ --config migration.yaml \ --format all \ --out-dir ./migration-plan ``` This generates six artifacts: | File | Purpose | | --- | --- | | `plan.json` | Machine-readable migration plan (topic list, partition counts, estimated data volume) | | `runbook.md` | Auto-generated step-by-step runbook customized to your clusters | | `aws-cli.sh` | AWS CLI commands for any infrastructure setup needed | | `iam-policy-templated.json` | IAM policy template with placeholder ARNs | | `iam-policy-concrete.json` | IAM policy with your actual cluster and bucket ARNs | | `cost-estimate.json` | Estimated S3 storage and data transfer costs | > [!WARNING] > > [!NOTE] > > Review the IAM policy > > [!NOTE] > > Attach `iam-policy-concrete.json` to the IAM role running the migration. Without these permissions, `execute` will fail with S3 or MSK access errors. Example plan summary from the May 8, 2026 AWS release qualification: ``` source=kbe-msk-kraft-e2e-source-zk metadata=ZOOKEEPER kafka=3.7.x brokers=3 target=kbe-msk-kraft-e2e-target-kraft-r4 metadata=KRAFT kafka=3.9.x.kraft brokers=3 topics_to_migrate=11 partitions_to_seed=68 estimated_seed_bytes=123831296 skipped_internal=["__amazon_msk_canary","__consumer_offsets"] precheck_findings=["W04","W07","W10","W03","W08"] sample topics: e2e.events partitions=9 rf=3 e2e.orders partitions=6 rf=3 e2e.compact partitions=3 rf=3 cleanup.policy=compact streams-orders-app-repartition partitions=6 rf=3 connect-offsets partitions=25 rf=3 cleanup.policy=compact ``` ``` kafka-backup migrate msk-kraft precheck --config migration.yaml ``` Precheck performs read-only analysis of both clusters and reports: - **Blockers** (B-codes): Must be resolved before migration can proceed - **Warnings** (W-codes): Proceed with awareness - **Info** (I-codes): Informational, no action needed Common blockers and their fixes: | Code | Issue | Fix | | --- | --- | --- | | B03 | Source is not ZK mode | Verify source ARN points to a ZooKeeper-mode cluster | | B04 | Target is not KRaft mode | Provision target in KRaft mode | | B07/B08 | S3 bucket not reachable | Create buckets or fix IAM permissions | | B09/B10 | Kafka brokers not reachable | Check security groups and bootstrap servers | | B11/B12 | Target message size too small | Increase target `message.max.bytes` | See the [Precheck Codes Reference](https://kafkabackup.com/enterprise/msk-kraft-precheck-codes.md) for all codes with detailed remediation. Example precheck output with no blockers: ``` W04 warn: could not verify target message-size floor (target broker DescribeConfigs returned no message.max.bytes or replica.fetch.max.bytes (dynamic-config only on this broker)) — ensure target `message.max.bytes` and `replica.fetch.max.bytes` ≥ largest source topic's effective max.message.bytes W03 info: KMS key ARN set on backup channel — CMK access is not verified by this precheck phase; ensure the caller has kms:Encrypt/Decrypt/GenerateDataKey I01 info: target is IAM-auth — ACLs will be emitted as access-map.json for customer IaC to translate to IAM policies (tool does not apply IAM) ``` ``` kafka-backup migrate msk-kraft execute \ --config migration.yaml \ --journal-dir ./journal ``` Execute runs through these phases automatically: 1. **Precheck** — re-verifies cluster compatibility 2. **Topology Copy** — creates topics and ACLs on target 3. **Seed** — bulk-copies all existing data through S3 4. **Tail** — continuously replicates new records until lag is within tolerance 5. **Drain Ready** — halts and waits for you to proceed The command blocks until `drain_ready` and then exits with code 0. This is your signal that the data is caught up and you can proceed to cutover. Example journal excerpt at drain-ready: ``` 2026-05-08T07:25:53.232357Z topology_copy -> seed 2026-05-08T07:30:06.232018Z seed -> tail 2026-05-08T07:31:23.055520Z tail -> drain_ready drain ready: max_partition_lag=0 records_replayed=1222 bytes_replayed=85708 ``` > [!NOTE] > > [!NOTE] > > Expected timeline > > [!NOTE] > > Seed phase duration depends on data volume. For 100GB of data, expect ~30 minutes for seed. Tail converges quickly for steady-state workloads. The entire execute phase typically completes in under an hour for clusters under 500GB. In another terminal, check migration status: ``` kafka-backup migrate msk-kraft status \ --config migration.yaml \ --migration-id \ --journal-dir ./journal ``` When execute completes with `drain_ready`, coordinate with your application teams: 1. **Schedule a maintenance window** — the producer freeze is typically under 60 seconds, but applications should be prepared 2. **Prepare client configs** — have new bootstrap servers ready to deploy (K8s ConfigMap, SSM Parameter, Consul KV, etc.) **Webhook (recommended)**: Configure `cutover.producer_freeze_webhook` in your config. The tool sends a POST request to freeze producers and a second POST to unfreeze them. **Manual TTY**: If no webhook is configured, the tool prompts in the terminal. You manually confirm when producers are frozen. > [!WARNING] > > [!NOTE] > > Non-interactive cutover needs a webhook > > [!NOTE] > > In CI, SSH automation, Kubernetes jobs, or any non-TTY shell, configure `cutover.producer_freeze_webhook`. Without a webhook the tool refuses to guess that producers are frozen. ``` kafka-backup migrate msk-kraft cutover \ --config migration.yaml \ --migration-id \ --journal-dir ./journal ``` Cutover performs: 1. Freezes producers (webhook or manual) 2. Publishes sentinel records to every partition 3. Drains final records from source to target 4. Snapshots all consumer group offsets from source 5. Translates offsets using the offset map (source → target) 6. Commits translated offsets on target 7. Verifies target log-start offsets have not advanced past the copied data 8. Logs `READY_FOR_CLIENT_SWITCH` Example cutover-ready output: ``` 2026-05-08T07:32:13.480416Z cutover -> awaiting_client_switch READY_FOR_CLIENT_SWITCH: groups_translated=3 offsets_committed=37 warnings=0 ``` After cutover completes, update your applications to point to the new KRaft cluster: ``` # Example: update Kubernetes ConfigMap kubectl patch configmap kafka-config -p '{"data":{"bootstrap.servers":"b-1.kraft.abc123.kafka.us-east-1.amazonaws.com:9098,b-2.kraft.abc123.kafka.us-east-1.amazonaws.com:9098"}}' # Roll deployments kubectl rollout restart deployment/order-service deployment/analytics-service ``` Consumers resume from translated target offsets so message continuity is preserved across the switch. Spot-check a consumer group to confirm it reads from the expected position: ``` kafka-consumer-groups.sh \ --bootstrap-server \ --group \ --describe ``` Once all clients are running against the target: ``` kafka-backup migrate msk-kraft cutover-ack \ --config migration.yaml \ --migration-id \ --journal-dir ./journal ``` This moves the migration to the `validating` state. ``` kafka-backup migrate msk-kraft finalize \ --config migration.yaml \ --migration-id \ --journal-dir ./journal ``` Finalize runs the 5-check validation suite: 1. **Topic parity** — partition counts match 2. **Counts & offsets** — record counts within tolerance and target offset-floor safety holds 3. **Spot-check records** — sampled records are byte-equal 4. **Sentinel presence** — cutover markers landed 5. **Consumer group reconciliation** — translated offsets committed correctly On success, the Ed25519-signed evidence bundle is uploaded to S3. Example AWS release rehearsal validation summary: ``` overall=PASSED topic_parity=PASSED detail="11 topic(s) match on partition count" counts_and_offsets=PASSED detail="68 partition(s) within count_tolerance=1" offset_floor_violations=0 spot_check_records=PASSED detail="204 samples compared, all matched" sentinel_presence=PASSED detail="sentinel observed on 68 partition(s)" consumer_group_reconciliation=PASSED detail="3 group(s) reconciled across 37 (group,topic,partition) triples" ``` Download and inspect the evidence: ``` aws s3 cp \ s3://prod-migration-evidence/migrations//evidence.json \ ./evidence.json # Verify the signature (optional) cat evidence.json | jq '.bundle_json' -r | sha256sum # Check validation outcome cat evidence.json | jq -r '.bundle_json' | jq -r '.validation.overall' # Check individual validation outcomes cat evidence.json | jq -r '.bundle_json' | jq '{ topic_parity: .validation.topic_parity.outcome, counts_and_offsets: .validation.counts_and_offsets.outcome, offset_floor_violations: .validation.counts_and_offsets.data.offset_floor_violations, spot_check_records: .validation.spot_check_records.outcome, sentinel_presence: .validation.sentinel_presence.outcome, consumer_group_reconciliation: .validation.consumer_group_reconciliation.outcome }' ``` The evidence bundle contains the complete migration journal, cluster snapshots, topology diff, ACL plan, seed/tail statistics, cutover report, and validation results. Share it with your compliance team. > [!WARNING] > > [!NOTE] > > Rollback is only available before cutover completes > > [!NOTE] > > Once cutover commits translated offsets to the target, rollback is no longer available. The source cluster remains untouched throughout — if you need to abort post-cutover, point applications back to the source cluster manually. To rollback a migration before cutover: ``` kafka-backup migrate msk-kraft rollback \ --config migration.yaml \ --migration-id \ --journal-dir ./journal ``` Rollback: - Unfreezes producers (if frozen) - Marks the migration as `rolled_back` in the journal - Uploads a rollback report to the evidence bucket - Does **not** delete topics/data on the target (manual cleanup) If the migration fails mid-execution (network error, timeout, crash): ``` kafka-backup migrate msk-kraft resume \ --config migration.yaml \ --migration-id \ --journal-dir ./journal ``` Resume reads the journal to find the last successful state and re-enters execution from there. The offset-map sidecar and tail checkpoints are persisted in S3, so no data is re-transferred. If the resume fingerprint doesn't match (config changed between runs), use `--force-restart` to override — but note this may cause duplicate records on the target in the seed phase. Rough estimates for a 3-broker MSK cluster: | Data volume | Seed phase | Tail convergence | Cutover window | Total | | --- | --- | --- | --- | --- | | 10 GB | ~5 min | ~2 min | < 30s | ~10 min | | 100 GB | ~30 min | ~5 min | < 30s | ~45 min | | 1 TB | ~4 hrs | ~15 min | < 60s | ~5 hrs | | 10 TB | ~36 hrs | ~30 min | < 60s | ~40 hrs | Factors that affect timing: - **Network bandwidth** between runner, MSK, and S3 - **Partition count** — more partitions = more parallelism in seed - **Message size** — large messages consume bandwidth faster - **Producer throughput** during tail — high-throughput topics take longer to converge - [Configuration Reference](https://kafkabackup.com/enterprise/msk-kraft-config-reference.md) — tune cutover, validation, and seed parameters - [Monitoring Guide](https://kafkabackup.com/guides/msk-kraft-monitoring.md) — what to watch during migration - [Troubleshooting](https://kafkabackup.com/troubleshooting/msk-kraft-migration.md) — common errors and fixes - [Precheck Codes](https://kafkabackup.com/enterprise/msk-kraft-precheck-codes.md) — all precheck findings explained --- title: MSK KRaft Migration Monitoring description: Operator monitoring guide for kafka-backup Enterprise MSK ZooKeeper to KRaft migrations — what to watch, when to alert, and how to interpret the journal. source_url: html: https://kafkabackup.com/guides/msk-kraft-monitoring md: https://kafkabackup.com/guides/msk-kraft-monitoring.md --- # MSK KRaft Migration Monitoring During a migration, track these key signals: | Phase | Monitor | Healthy signal | | --- | --- | --- | | **Seed** | S3 write throughput, source cluster CPU | Steady throughput, CPU 2x expected duration | Investigate: check source CPU, S3 throughput, network | | Tail lag not decreasing for 10+ minutes | Check source producer throughput, runner resources | | Cutover producer freeze approaching timeout | Check webhook health, increase `max_producer_freeze` | | Validation check FAILED | Keep clients on the target only if the issue is understood; repair the affected topic/partition and retry `finalize` | | Migration in `failed` state | Check logs, resolve issue, `resume` | The journal (`journal.jsonl`) records every state transition: ``` cat ./journal//journal.jsonl | jq . ``` Each entry contains: - `from` / `to`: State transition - `at`: UTC timestamp - `reason`: Phase-specific summary (records processed, lag status, etc.) This excerpt is from the May 8, 2026 AWS release qualification with 11 migrated topics and 68 partitions: ``` 2026-05-08T07:25:52.456012Z null -> planned planned | resume_fp=6f45dcfd5e980225 2026-05-08T07:25:52.532740Z planned -> precheck 2026-05-08T07:25:52.593949Z precheck -> topology_copy 2026-05-08T07:25:53.232357Z topology_copy -> seed 2026-05-08T07:30:06.232018Z seed -> tail 2026-05-08T07:31:23.055520Z tail -> drain_ready drain ready: max_partition_lag=0 records_replayed=1222 bytes_replayed=85708 2026-05-08T07:32:13.480416Z cutover -> awaiting_client_switch READY_FOR_CLIENT_SWITCH: groups_translated=3 offsets_committed=37 warnings=0 2026-05-08T07:32:14.657229Z awaiting_client_switch -> validating operator confirmed client cutover ``` ``` {"from": "seed", "to": "failed", "at": "2026-04-24T12:00:00Z", "reason": "S3 PutObject AccessDenied"} {"from": "failed", "to": "seed", "at": "2026-04-24T12:30:00Z", "reason": "resumed from seed"} {"from": "seed", "to": "tail", "at": "2026-04-24T16:00:00Z", "reason": "seed complete"} ``` 1. Run `status` to see the current state and lag 2. Check RUST\_LOG=debug output for the last operation 3. If in `tail`, check per-partition lag — one lagging partition can hold up drain-ready 4. If in `seed`, check S3 write throughput and source cluster health 1. Read the journal's last entry for the failure reason 2. Check the full log output (redirect to file with `2>&1 | tee migration.log`) 3. Resolve the issue (permissions, network, cluster health) 4. Run `resume` — it re-enters from the last successful state - **Before cutover**: Run `rollback` to cleanly abort. Source is untouched. - **After cutover**: The source is untouched but offsets are committed on target. Point applications back to source if needed. - [Production Runbook](https://kafkabackup.com/guides/msk-kraft-migration-runbook.md) — step-by-step guide - [Troubleshooting](https://kafkabackup.com/troubleshooting/msk-kraft-migration.md) — detailed error reference - [CLI Reference](https://kafkabackup.com/enterprise/msk-kraft-cli-reference.md) — status command options --- title: CLI Reference description: Complete reference for all OSO Kafka Backup CLI commands, flags, and options source_url: html: https://kafkabackup.com/reference/cli-reference md: https://kafkabackup.com/reference/cli-reference.md --- # CLI Reference Complete reference documentation for all OSO Kafka Backup CLI commands. These options apply to all commands: ``` kafka-backup [OPTIONS] ``` | Option | Description | | --- | --- | | `-v, --verbose` | Enable verbose logging. Use `-v` for debug, `-vv` for trace | | `-h, --help` | Print help information | | `-V, --version` | Print version information | * * * Run a backup operation from a Kafka cluster to storage. ``` kafka-backup backup --config ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-c, --config` | PATH | Yes | Path to backup configuration YAML | ``` # Basic backup kafka-backup backup --config backup.yaml # With debug logging kafka-backup -v backup --config /etc/kafka-backup/production.yaml # With trace logging kafka-backup -vv backup --config backup.yaml ``` | Code | Meaning | | --- | --- | | 0 | Backup completed successfully (or graceful shutdown) | | 1 | Backup failed | The backup and restore commands handle `SIGTERM` and `SIGINT` (Ctrl+C) for graceful shutdown (v0.8.1+). On signal: 1. In-progress segments are flushed to storage 2. A final checkpoint is saved 3. The process exits with code 0 This ensures safe interruption and allows resuming from the last checkpoint. * * * Restore data from a backup to a Kafka cluster. ``` kafka-backup restore --config ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-c, --config` | PATH | Yes | Path to restore configuration YAML | ``` # Basic restore kafka-backup restore --config restore.yaml # Restore with verbose output kafka-backup -v restore --config dr-restore.yaml ``` * * * List available backups or show details of a specific backup. ``` kafka-backup list --path [--backup-id ] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-p, --path` | PATH | Yes | Path to storage location (local path or S3 URI) | | `-b, --backup-id` | STRING | No | Specific backup ID to show details for | ``` # List all backups in local storage kafka-backup list --path /var/lib/kafka-backup/data # List backups in S3 kafka-backup list --path s3://my-bucket/backups # Show details for a specific backup kafka-backup list --path /data --backup-id daily-backup-001 ``` ``` Available Backups: ───────────────────────────────────────────────────────────── daily-backup-001 Created: 2024-12-03T02:00:00Z Topics: 5 Records: 1,234,567 Size: 128 MB (compressed) ───────────────────────────────────────────────────────────── ``` * * * Show status and statistics of a backup job. Supports two modes: 1. **Static inspection**: Inspect stored backup artifacts using `--path` and `--backup-id` 2. **Live monitoring**: Monitor a running backup using `--config` (optionally with `--watch`) ``` # Static inspection kafka-backup status --path --backup-id [--db-path ] # Live monitoring (one-shot) kafka-backup status --config # Live monitoring (continuous watch) kafka-backup status --config --watch [--interval ] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-p, --path` | PATH | No\* | Path to storage location (for static inspection) | | `-b, --backup-id` | STRING | No\* | Backup ID to show status for (for static inspection) | | `--db-path` | PATH | No | Path to offset database (for tracking progress) | | `-c, --config` | PATH | No\* | Path to config file (for live monitoring) | | `--watch` | FLAG | No | Enable continuous watch mode (requires `--config`) | | `--interval` | INT | No | Refresh interval in seconds for watch mode (default: 2) | \*Either `--config` OR both `--path` and `--backup-id` are required. **Static inspection of stored backup:** ``` kafka-backup status --path /data --backup-id backup-001 kafka-backup status -p /data -b backup-001 --db-path /var/lib/kafka-backup/offsets.db ``` **Live monitoring of running backup (one-shot):** ``` kafka-backup status --config backup.yaml ``` **Live monitoring with continuous watch:** ``` # Default 2-second refresh kafka-backup status --config backup.yaml --watch # Custom 5-second refresh interval kafka-backup status --config backup.yaml --watch --interval 5 ``` When using `--config`, the status command connects to the running backup's metrics server (`/health` and `/metrics` endpoints) and displays: ``` ================================================================ OSO Kafka Backup - Live Status ================================================================ Backup ID: production-backup Uptime: 00:15:32 Status: RUNNING ================================================================ Progress |- Records: 1,234,567 |- Bytes: 256.0 MB (compressed) |- Throughput: 15234 rec/s | 3.2 MB/s |- Lag: 45,000 records (orders-0) ================================================================ Components |- kafka: [OK] ok |- storage: [OK] ok |- checkpointing: [OK] ok ================================================================ Compression: 3.2x ratio | Errors: 0 ================================================================ Last updated: 2025-01-30 14:32:15 | Refresh: 2s | Ctrl+C to exit ``` When using `--path` and `--backup-id`, shows: - Manifest information (created date, source cluster) - Topic/partition/segment counts - Record and size statistics (compressed and uncompressed) - Compression ratio - Offset tracking status per partition Metrics must be enabled in your backup config: ``` metrics: enabled: true port: 8080 bind_address: "0.0.0.0" ``` * * * Show detailed backup manifest information. ``` kafka-backup describe --config [--format ] kafka-backup describe --path --backup-id [--format ] ``` | Flag | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `-c, --config` | PATH | one of | \- | Backup configuration file — storage (including `prefix`) and `backup_id` are taken from it (v0.22.0+). Conflicts with `--path`/`--backup-id` | | `-p, --path` | PATH | one of | \- | Path to storage location (requires `--backup-id`) | | `-b, --backup-id` | STRING | with `--path` | \- | Backup ID to describe | | `-f, --format` | STRING | No | text | Output format: `text`, `json`, `yaml` | ``` # From the backup config (same storage and prefix as the backup itself) kafka-backup describe --config backup.yaml # Human-readable output kafka-backup describe --path /data --backup-id backup-001 # JSON for scripting kafka-backup describe -p /data -b backup-001 --format json # YAML output kafka-backup describe -p /data -b backup-001 -f yaml ``` - Backup ID and creation timestamp - Source cluster ID and brokers - Compression algorithm - Topic/partition/segment/record counts - Compressed and uncompressed sizes - Compression ratio - Time range of backed-up data - Per-topic and per-partition details * * * Validate backup integrity. ``` kafka-backup validate --config [--deep] kafka-backup validate --path --backup-id [--deep] ``` | Flag | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `-c, --config` | PATH | one of | \- | Backup configuration file — storage (including `prefix`) and `backup_id` are taken from it (v0.22.0+). Conflicts with `--path`/`--backup-id` | | `-p, --path` | PATH | one of | \- | Path to storage location (requires `--backup-id`) | | `-b, --backup-id` | STRING | with `--path` | \- | Backup ID to validate | | `--deep` | BOOL | No | false | Perform deep validation (read and verify each segment) | The report also lists **pruned ranges** (deleted by retention — informational) and **missing topics** (literal include topics skipped under `backup.on_missing_topic: warn` — informational); neither makes the backup invalid. ``` # From the backup config kafka-backup validate --config backup.yaml --deep # Shallow validation (quick - checks existence and metadata) kafka-backup validate --path /data --backup-id backup-001 # Deep validation (thorough - reads and parses all segments) kafka-backup validate -p /data -b backup-001 --deep ``` ``` Validation Report: backup-001 ════════════════════════════════════════════════════════════ Segments: Checked: 156 Valid: 156 Missing: 0 Corrupted: 0 Records Validated: 2,456,789 Issues Found: 0 Result: ✓ VALID ``` * * * Delete aged or oversized segments from a backup set — safely (v0.21.0+). Plan-only by default; `--execute` performs the deletion. The manifest is rewritten **before** any object is deleted and every removed offset range is recorded as a _pruned range_, so `validate` stays green and later runs never resurrect the data. Never use bucket lifecycle rules on incremental sets — see [Retention and Erasure](https://kafkabackup.com/guides/retention-and-erasure.md). ``` kafka-backup prune --config [CRITERIA] [--execute] kafka-backup prune --path --backup-id [CRITERIA] [--execute] ``` | Flag | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `-c, --config` | PATH | one of | \- | Backup configuration file (storage incl. `prefix`, `backup_id`) | | `-p, --path` / `-b, --backup-id` | PATH / STRING | one of | \- | Explicit storage location and backup ID | | `--older-than` | DURATION | one of | \- | Prune segments older than `30d`, `12h`, `1d12h`, `45m`, … | | `--before` | TIMESTAMP | one of | \- | Prune segments older than an RFC 3339 instant or epoch milliseconds | | `--max-total-bytes` | INT | No | \- | Keep pruning oldest-first until the set's compressed size fits | | `--keep-segments` | INT | No | 1 | Never prune a partition below this many newest segments | | `--execute` | BOOL | No | false | Actually delete (default is a dry-run plan) | | `--force` | BOOL | No | false | Proceed even when a backup run looks live (job status `running` with a fresh checkpoint) | | `-f, --format` | STRING | No | text | `text` or `json` | Only a contiguous oldest-first prefix of each partition is ever pruned (no holes), and segments the next incremental run still needs to resume from are protected. `offsets.db` and `consumer-groups-snapshot.json` are never touched. ``` # Plan: what would a 30-day retention remove? kafka-backup prune --config backup.yaml --older-than 30d # Apply it kafka-backup prune --config backup.yaml --older-than 30d --execute # Cap a set at 2 TiB compressed, keeping the newest 2 segments per partition kafka-backup prune --path s3://bucket/prefix --backup-id daily --max-total-bytes 2199023255552 --keep-segments 2 --execute ``` The same policy can run automatically at the end of every backup with `backup.retention` (`max_age`, `max_total_bytes`, `keep_segments`) — see the [configuration reference](https://kafkabackup.com/reference/config-yaml.md). * * * Validate a restore configuration without executing it. ``` kafka-backup validate-restore --config [--format ] ``` | Flag | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `-c, --config` | PATH | Yes | \- | Path to restore configuration file | | `-f, --format` | STRING | No | text | Output format: `text`, `json`, `yaml` | ``` # Validate restore config kafka-backup validate-restore --config restore.yaml # Get JSON report for CI/CD kafka-backup validate-restore -c restore.yaml --format json ``` - Restore status (VALID or INVALID) - Topics to restore (count and names) - Segments to process - Records and bytes to restore - Time range of restore - Consumer offset actions - Errors and warnings * * * Show offset mapping for a backup. ``` kafka-backup show-offset-mapping --path --backup-id [--format ] ``` | Flag | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `-p, --path` | PATH | Yes | \- | Path to storage location | | `-b, --backup-id` | STRING | Yes | \- | Backup ID to show offset mapping for | | `-f, --format` | STRING | No | text | Output: `text`, `json`, `yaml`, `csv` | ``` # Human-readable output kafka-backup show-offset-mapping --path /data --backup-id backup-001 # CSV for spreadsheets kafka-backup show-offset-mapping -p /data -b backup-001 --format csv # JSON for programmatic use kafka-backup show-offset-mapping -p /data -b backup-001 -f json ``` ``` Offset Mapping: backup-001 ───────────────────────────────────────────────────────────── Topic Partition Source Start Source End Records ───────────────────────────────────────────────────────────── orders 0 0 150233 150234 orders 1 0 148891 148892 orders 2 0 152456 152457 payments 0 0 78234 78235 ... To reset consumer groups, use: kafka-consumer-groups --bootstrap-server \ --group --topic orders:0 \ --reset-offsets --to-offset 150233 --execute ``` * * * Generate or execute consumer group offset reset plans. Generate an offset reset plan from backup's offset mapping. ``` kafka-backup offset-reset plan \ --path \ --backup-id \ --groups \ --bootstrap-servers \ [--format ] \ [--dry-run ] ``` | Flag | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `-p, --path` | PATH | Yes | \- | Path to storage location | | `-b, --backup-id` | STRING | Yes | \- | Backup ID with offset mapping | | `-g, --groups` | STRING | Yes | \- | Consumer groups (comma-separated) | | `--bootstrap-servers` | STRING | Yes | \- | Kafka bootstrap servers | | `-f, --format` | STRING | No | text | Output: `text`, `json`, `csv`, `shell-script` | | `--dry-run` | BOOL | No | true | Preview only, no changes | ``` kafka-backup offset-reset plan \ --path /data \ --backup-id backup-001 \ --groups my-group,another-group \ --bootstrap-servers localhost:9092 # Generate shell script kafka-backup offset-reset plan \ -p /data -b backup-001 \ -g consumer-group \ --bootstrap-servers broker-1:9092 \ --format shell-script ``` Execute an offset reset plan. ``` kafka-backup offset-reset execute \ --path \ --backup-id \ --groups \ --bootstrap-servers \ [--security-protocol ] ``` | Flag | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `-p, --path` | PATH | Yes | \- | Path to storage location | | `-b, --backup-id` | STRING | Yes | \- | Backup ID | | `-g, --groups` | STRING | Yes | \- | Consumer groups | | `--bootstrap-servers` | STRING | Yes | \- | Kafka bootstrap servers | | `--security-protocol` | STRING | No | PLAINTEXT | Security: `PLAINTEXT`, `SSL`, `SASL_SSL`, `SASL_PLAINTEXT` | | `--sasl-mechanism` | STRING | No | \- | SASL mechanism: `PLAIN`, `SCRAM-SHA256`, `SCRAM-SHA512`, `GSSAPI` | | `--sasl-username` | STRING | No | \- | SASL username (for PLAIN / SCRAM) | | `--sasl-password` | STRING | No | \- | SASL password (for PLAIN / SCRAM). Prefer `SASL_PASSWORD` env var. | | `--sasl-kerberos-service-name` | STRING | No | \- | Broker service-principal name (GSSAPI only, e.g. `kafka`) | | `--sasl-keytab` | PATH | No | \- | Path to Kerberos keytab (GSSAPI only) | | `--sasl-krb5-config` | PATH | No | \- | Override path to `krb5.conf` (GSSAPI only) | > [!NOTE] > > [!NOTE] > > SASL/GSSAPI requires a feature-gated build > > [!NOTE] > > `--sasl-mechanism GSSAPI` is rejected at startup unless the CLI was built with `--features gssapi`. See [Security Setup → SASL/GSSAPI](https://kafkabackup.com/guides/security-setup.md#sasl-gssapi). ``` kafka-backup offset-reset execute \ --path /data \ --backup-id backup-001 \ --groups my-group \ --bootstrap-servers localhost:9092 # With SSL kafka-backup offset-reset execute \ -p /data -b backup-001 \ -g consumer-group \ --bootstrap-servers broker-1:9092 \ --security-protocol SSL # With SASL/GSSAPI (requires --features gssapi at build time) kafka-backup offset-reset execute \ -p /data -b backup-001 \ -g consumer-group \ --bootstrap-servers broker.prod.corp:9098 \ --security-protocol SASL_PLAINTEXT \ --sasl-mechanism GSSAPI \ --sasl-kerberos-service-name kafka \ --sasl-keytab /etc/kafka-backup/client.keytab \ --sasl-krb5-config /etc/krb5.conf ``` Generate a shell script for manual offset reset. ``` kafka-backup offset-reset script \ --path \ --backup-id \ --groups \ --bootstrap-servers \ [--output ] ``` | Flag | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `-o, --output` | PATH | No | stdout | Output file path | * * * Execute bulk parallel offset reset (optimized for large consumer groups). ``` kafka-backup offset-reset-bulk \ --path \ --backup-id \ --groups \ --bootstrap-servers \ [OPTIONS] ``` | Flag | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `-p, --path` | PATH | Yes | \- | Path to storage location | | `-b, --backup-id` | STRING | Yes | \- | Backup ID | | `-g, --groups` | STRING | Yes | \- | Consumer groups (comma-separated) | | `--bootstrap-servers` | STRING | Yes | \- | Kafka bootstrap servers | | `--max-concurrent` | INT | No | 50 | Maximum concurrent requests | | `--max-retries` | INT | No | 3 | Maximum retry attempts | | `--security-protocol` | STRING | No | PLAINTEXT | Security: `PLAINTEXT`, `SSL`, `SASL_SSL`, `SASL_PLAINTEXT` | | `--sasl-mechanism` | STRING | No | \- | SASL mechanism: `PLAIN`, `SCRAM-SHA256`, `SCRAM-SHA512`, `GSSAPI` | | `--sasl-username` | STRING | No | \- | SASL username (for PLAIN / SCRAM) | | `--sasl-password` | STRING | No | \- | SASL password (for PLAIN / SCRAM). Prefer `SASL_PASSWORD` env var. | | `--sasl-kerberos-service-name` | STRING | No | \- | Broker service-principal name (GSSAPI only, e.g. `kafka`) | | `--sasl-keytab` | PATH | No | \- | Path to Kerberos keytab (GSSAPI only) | | `--sasl-krb5-config` | PATH | No | \- | Override path to `krb5.conf` (GSSAPI only) | | `-f, --format` | STRING | No | text | Output: `text`, `json` | > [!NOTE] > > [!NOTE] > > SASL/GSSAPI requires a feature-gated build > > [!NOTE] > > `--sasl-mechanism GSSAPI` is rejected at startup unless the CLI was built with `--features gssapi`. See [Security Setup → SASL/GSSAPI](https://kafkabackup.com/guides/security-setup.md#sasl-gssapi). ``` kafka-backup offset-reset-bulk \ --path /data \ --backup-id backup-001 \ --groups group1,group2 \ --bootstrap-servers broker-1:9092,broker-2:9092 \ --max-concurrent 100 # With metrics output kafka-backup offset-reset-bulk \ -p /data -b backup-001 \ -g my-group \ --bootstrap-servers localhost:9092 \ --format json # With SASL/GSSAPI (requires --features gssapi at build time) kafka-backup offset-reset-bulk \ -p /data -b backup-001 \ -g my-group \ --bootstrap-servers broker.prod.corp:9098 \ --security-protocol SASL_PLAINTEXT \ --sasl-mechanism GSSAPI \ --sasl-kerberos-service-name kafka \ --sasl-keytab /etc/kafka-backup/client.keytab \ --sasl-krb5-config /etc/krb5.conf \ --max-concurrent 100 ``` - ~50x faster than sequential offset reset - Per-partition retry with exponential backoff - Detailed metrics (p50/p99 latency, throughput) - Progress reporting * * * Snapshot and rollback consumer group offsets. Create a snapshot of current consumer group offsets. ``` kafka-backup offset-rollback snapshot \ --path \ --groups \ --bootstrap-servers \ [--description ] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-p, --path` | PATH | Yes | Path to store snapshots | | `-g, --groups` | STRING | Yes | Consumer groups | | `--bootstrap-servers` | STRING | Yes | Kafka bootstrap servers | | `-d, --description` | STRING | No | Snapshot description | | `--security-protocol` | STRING | No | Security: `PLAINTEXT`, `SSL`, `SASL_SSL`, `SASL_PLAINTEXT` | | `--sasl-mechanism` | STRING | No | SASL mechanism: `PLAIN`, `SCRAM-SHA256`, `SCRAM-SHA512`, `GSSAPI` | | `--sasl-username` | STRING | No | SASL username (for PLAIN / SCRAM) | | `--sasl-password` | STRING | No | SASL password (for PLAIN / SCRAM). Prefer `SASL_PASSWORD` env var. | | `--sasl-kerberos-service-name` | STRING | No | Broker service-principal name (GSSAPI only, e.g. `kafka`) | | `--sasl-keytab` | PATH | No | Path to Kerberos keytab (GSSAPI only) | | `--sasl-krb5-config` | PATH | No | Override path to `krb5.conf` (GSSAPI only) | | `-f, --format` | STRING | No | Output: `text`, `json` | > [!NOTE] > > [!NOTE] > > SASL/GSSAPI requires a feature-gated build > > [!NOTE] > > `--sasl-mechanism GSSAPI` is rejected at startup unless the CLI was built with `--features gssapi`. See [Security Setup → SASL/GSSAPI](https://kafkabackup.com/guides/security-setup.md#sasl-gssapi). ``` kafka-backup offset-rollback snapshot \ --path /data/snapshots \ --groups my-group \ --bootstrap-servers localhost:9092 \ --description "Before maintenance" # With SASL/GSSAPI (requires --features gssapi at build time) kafka-backup offset-rollback snapshot \ --path /data/snapshots \ --groups my-group \ --bootstrap-servers broker.prod.corp:9098 \ --security-protocol SASL_PLAINTEXT \ --sasl-mechanism GSSAPI \ --sasl-kerberos-service-name kafka \ --sasl-keytab /etc/kafka-backup/client.keytab \ --sasl-krb5-config /etc/krb5.conf \ --description "Before GSSAPI restore" ``` List available offset snapshots. ``` kafka-backup offset-rollback list --path ``` Show details of a specific snapshot. ``` kafka-backup offset-rollback show \ --path \ --snapshot-id ``` Rollback offsets to a previous snapshot. ``` kafka-backup offset-rollback rollback \ --path \ --snapshot-id \ --bootstrap-servers \ [--verify ] ``` | Flag | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `--verify` | BOOL | No | true | Verify offsets after rollback | Verify current offsets match a snapshot. ``` kafka-backup offset-rollback verify \ --path \ --snapshot-id \ --bootstrap-servers ``` Delete a snapshot. ``` kafka-backup offset-rollback delete \ --path \ --snapshot-id ``` * * * Run complete three-phase restore with automatic offset reset. ``` kafka-backup three-phase-restore --config ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-c, --config` | PATH | Yes | Path to configuration file | 1. **Phase 1**: Collect offset headers from source backup 2. **Phase 2**: Restore data to target cluster 3. **Phase 3**: Reset consumer group offsets ``` kafka-backup three-phase-restore --config dr-restore.yaml ``` ``` Three-Phase Restore Report ════════════════════════════════════════════════════════════ Backup ID: production-backup-001 Status: ✓ SUCCESS Phase 2 - Data Restore: Records Restored: 2,456,789 Topics: 4 Duration: 8m 23s Phase 3 - Offset Reset: Strategy: header-based Consumer Groups: 3 Partitions Reset: 33 Warnings: 0 ``` * * * Snapshot consumer group committed offsets for all topics covered by a backup. Queries **every broker** individually (KRaft-safe — avoids missing groups whose coordinator lives on a non-bootstrap broker), fetches committed offsets, filters to groups with offsets on backed-up topics, and writes `{backup_id}/consumer-groups-snapshot.json` to the configured storage backend. The snapshot can be loaded automatically at restore time with [`auto_consumer_groups: true`](https://kafkabackup.com/reference/config-yaml.md#restore-options-reference) in the restore configuration, or read manually to reconstruct consumer positions after a disaster recovery event. ``` kafka-backup snapshot-groups --config ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-c`, `--config` | PATH | Yes | Path to backup configuration file (must have `mode: backup`) | | `-v`, `--verbose` | flag | No | Enable verbose logging (`-v` debug, `-vv` trace) | ``` # Capture consumer group offsets after a one-shot backup kafka-backup backup --config backup.yaml kafka-backup snapshot-groups --config backup.yaml # Verify the snapshot was written aws s3 cp s3://my-bucket/my-backup/consumer-groups-snapshot.json - | jq . # Use the snapshot at restore time kafka-backup restore --config restore.yaml # restore.yaml must include: # restore: # auto_consumer_groups: true ``` The command writes `{backup_id}/consumer-groups-snapshot.json` to the configured storage backend with this structure: ``` { "snapshot_time": 1744123456789, "groups": [ { "group_id": "my-app-group", "offsets": { "orders": { "0": 1500, "1": 1423, "2": 1601 } } } ] } ``` - The config file must use `mode: backup` — the backup manifest is used to determine which topics are in scope for filtering. - To run this automatically after each backup cycle, set `consumer_group_snapshot: true` in the backup configuration instead. - In KRaft mode each broker is group coordinator for only a subset of consumer groups. This command queries all known brokers and deduplicates results, ensuring complete coverage without requiring a ZooKeeper-style single coordinator. * * * Run backup validation checks and generate compliance evidence reports. ``` kafka-backup validation ``` Execute validation checks against a restored cluster and generate evidence. ``` kafka-backup validation run --config [OPTIONS] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-c, --config` | PATH | Yes | Path to validation configuration YAML | | `--pitr` | INT | No | PITR timestamp override (epoch milliseconds) | | `--triggered-by` | STRING | No | Record who/what triggered this run | ``` # Basic validation run $ kafka-backup validation run --config validation.yaml # With PITR timestamp and audit trigger $ kafka-backup validation run \ --config validation.yaml \ --pitr 1711929600000 \ --triggered-by "KPMG Q1 2026 audit" ``` | Code | Meaning | | --- | --- | | 0 | All validation checks passed | | 1 | One or more checks failed | List evidence reports stored in object storage. ``` kafka-backup validation evidence-list --path [OPTIONS] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-p, --path` | STRING | Yes | Storage location (e.g., `s3://bucket/prefix`) | | `-l, --limit` | INT | No | Maximum reports to show (default: 50) | Download an evidence report from storage. ``` kafka-backup validation evidence-get --path --report-id --output [OPTIONS] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-p, --path` | STRING | Yes | Storage location | | `-r, --report-id` | STRING | Yes | Report ID to download | | `-f, --format` | STRING | No | Format: `json` (default) or `pdf` | | `-o, --output` | STRING | Yes | Output file path | Verify the cryptographic signature of an evidence report. ``` kafka-backup validation evidence-verify --report --signature [OPTIONS] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `-r, --report` | PATH | Yes | Path to JSON evidence report | | `-s, --signature` | PATH | Yes | Path to detached signature (.sig) file | | `--public-key` | PATH | No | Path to PEM-encoded public key for ECDSA verification | ``` # Verify SHA-256 checksum only $ kafka-backup validation evidence-verify \ --report evidence-report.json \ --signature evidence-report.sig # Verify SHA-256 + ECDSA signature $ kafka-backup validation evidence-verify \ --report evidence-report.json \ --signature evidence-report.sig \ --public-key signing-key-pub.pem ``` ``` Report ID: validation-9275b4aa-2aeb-4910-a3a6-9e4aa1dc016a Algorithm: ECDSA-P256-SHA256 Report SHA-256: 2482bbdfa113146e39a4884767002554... SHA-256 checksum: VALID ECDSA signature: VALID Evidence report integrity: VERIFIED ``` * * * Enterprise MSK ZooKeeper-to-KRaft migration commands. `plan` and `precheck` are free — no license required. ``` kafka-backup migrate msk-kraft ``` | Subcommand | License Required | Description | | --- | --- | --- | | `plan` | No | Generate migration plan, runbook, IAM policies, cost estimate | | `precheck` | No | Read-only cluster compatibility checks | | `execute` | Yes | Run migration through drain-ready | | `resume` | Yes | Resume an in-flight migration | | `status` | Yes | Show migration state and lag | | `cutover` | Yes | Freeze producers, translate offsets, prepare for client switch | | `cutover-ack` | Yes | Acknowledge client switch, trigger validation | | `rollback` | Yes | Abort migration (pre-cutover only) | | `finalize` | Yes | Run validation, produce signed evidence bundle | See the [MSK KRaft Migration CLI Reference](https://kafkabackup.com/enterprise/msk-kraft-cli-reference.md) for complete flags and options for each subcommand. * * * The CLI respects these environment variables: | Variable | Description | | --- | --- | | `KAFKA_BACKUP_CONFIG` | Default config file path | | `RUST_LOG` | Logging level (e.g., `info`, `debug`, `trace`) | | `AWS_ACCESS_KEY_ID` | AWS access key for S3 | | `AWS_SECRET_ACCESS_KEY` | AWS secret key for S3 | | `AWS_REGION` | AWS region for S3 | | `AZURE_STORAGE_ACCOUNT` | Azure storage account name | | `AZURE_STORAGE_KEY` | Azure storage account key | | `GOOGLE_APPLICATION_CREDENTIALS` | Path to GCP service account JSON | --- title: Configuration Reference description: Complete reference for OSO Kafka Backup YAML configuration files source_url: html: https://kafkabackup.com/reference/config-yaml md: https://kafkabackup.com/reference/config-yaml.md --- # Configuration Reference OSO Kafka Backup uses YAML configuration files for backup and restore operations. This reference documents all available options. ``` # Required: Operation mode mode: backup # or "restore" # Required: Unique backup identifier backup_id: "my-backup-001" # Source/Target Kafka cluster configuration source: # For backup mode target: # For restore mode bootstrap_servers: [] security: {} topics: {} # Storage backend configuration storage: backend: filesystem # or s3, azure, gcs # Backend-specific options... # Mode-specific options backup: {} # Backup options restore: {} # Restore options # Optional Prometheus endpoint metrics: {} ``` * * * **Required.** Operation mode. ``` mode: backup # Run a backup operation mode: restore # Run a restore operation ``` **Required.** Unique identifier for the backup. ``` backup_id: "production-daily-001" backup_id: "dr-backup-$(date +%Y%m%d)" ``` * * * Used as `source` in backup mode and `target` in restore mode. **Required.** List of Kafka broker addresses. ``` source: bootstrap_servers: - broker-1.kafka.svc:9092 - broker-2.kafka.svc:9092 - broker-3.kafka.svc:9092 ``` Optional security configuration for Kafka connection. ``` source: security: # Security protocol security_protocol: SASL_SSL # PLAINTEXT, SSL, SASL_PLAINTEXT, SASL_SSL # SASL configuration sasl_mechanism: SCRAM-SHA256 # PLAIN, SCRAM-SHA256, SCRAM-SHA512 sasl_username: backup-user sasl_password: ${KAFKA_PASSWORD} # Environment variable substitution # SSL/TLS configuration ssl_ca_location: /etc/kafka/ca.crt ssl_certificate_location: /etc/kafka/client.crt ssl_key_location: /etc/kafka/client.key ssl_key_password: ${SSL_KEY_PASSWORD} ``` | Protocol | Description | | --- | --- | | `PLAINTEXT` | No encryption, no authentication | | `SSL` | TLS encryption, optional mTLS | | `SASL_PLAINTEXT` | SASL authentication, no encryption | | `SASL_SSL` | SASL authentication with TLS encryption | | Mechanism | Description | Feature gate | | --- | --- | --- | | `PLAIN` | Simple username/password | Built-in | | `SCRAM-SHA256` | SCRAM-SHA-256 — Salted Challenge Response (SHA-256) | Built-in | | `SCRAM-SHA512` | SCRAM-SHA-512 — Salted Challenge Response (SHA-512) | Built-in | | `GSSAPI` | Kerberos authentication (RFC 4752). Uses a keytab for non-interactive credential acquisition. | Requires CLI built with `--features gssapi` | > [!NOTE] > > [!NOTE] > > SASL/GSSAPI opt-in > > [!NOTE] > > `GSSAPI` is gated behind the `gssapi` cargo feature. The default `oso/kafka-backup` release binary and Docker image do **not** include it. A binary without the feature that loads a `sasl_mechanism: GSSAPI` config fails fast at startup with a clear error. See the [Security Setup guide](https://kafkabackup.com/guides/security-setup.md#sasl-gssapi) for build instructions. Required only when `sasl_mechanism: GSSAPI`. Ignored for other mechanisms. ``` source: security: security_protocol: SASL_PLAINTEXT # or SASL_SSL sasl_mechanism: GSSAPI sasl_kerberos_service_name: kafka sasl_keytab_path: /etc/kafka-backup/client.keytab sasl_krb5_config_path: /etc/krb5.conf ``` | Field | Type | Required | Description | | --- | --- | --- | --- | | `sasl_kerberos_service_name` | string | Yes | The service-name portion of the broker's service principal. For `kafka/broker.host@REALM`, this is `kafka`. | | `sasl_keytab_path` | path | Yes | Absolute path to the keytab holding the client principal's long-term key. Must be readable by the kafka-backup process; lock to `0400` in production. | | `sasl_krb5_config_path` | path | No | Override the default `/etc/krb5.conf`. Useful inside containers that don't share the host's krb5 config. | Session re-authentication (KIP-368) is handled automatically — if the broker advertises a session lifetime, kafka-backup re-handshakes at ~80 % of the window (30 s floor, ±5 s jitter) before the broker forcibly disconnects. Optional TCP connection settings. These are particularly important for cloud-hosted Kafka services like Confluent Cloud that terminate idle connections. ``` source: connection: # Enable TCP keepalive to prevent idle connection termination tcp_keepalive: true # Default: true # Time in seconds before first keepalive probe is sent keepalive_time_secs: 60 # Default: 60 # Interval in seconds between keepalive probes keepalive_interval_secs: 20 # Default: 20 # Enable TCP_NODELAY (disable Nagle's algorithm) for lower latency tcp_nodelay: true # Default: true ``` | Option | Type | Default | Description | | --- | --- | --- | --- | | `tcp_keepalive` | bool | `true` | Enable TCP keepalive probes | | `keepalive_time_secs` | int | `60` | Seconds idle before first probe | | `keepalive_interval_secs` | int | `20` | Seconds between probes | | `tcp_nodelay` | bool | `true` | Disable Nagle's algorithm | > [!TIP] > > [!NOTE] > > Confluent Cloud > > [!NOTE] > > Confluent Cloud terminates idle TCP connections after ~5 minutes. The default keepalive settings (60s time, 20s interval) prevent this. If you're experiencing "Broken pipe" errors with Confluent Cloud, ensure TCP keepalive is enabled. Topic selection for backup or restore. ``` source: topics: # Include specific topics or patterns include: - orders # Exact topic name - payments # Another exact name - "events-*" # Wildcard pattern - "logs-2024-*" # Date-based pattern # Exclude topics (applied after include) exclude: - "__consumer_offsets" # Internal Kafka topic - "_schemas" # Schema Registry topic - "*-internal" # Pattern exclusion ``` * * * **Required.** Storage backend type. ``` storage: backend: filesystem # Local filesystem or mounted volume backend: s3 # Amazon S3 or S3-compatible storage backend: azure # Azure Blob Storage backend: gcs # Google Cloud Storage ``` ``` storage: backend: filesystem path: "/var/lib/kafka-backup/data" prefix: "cluster-prod" # Optional subdirectory ``` ``` storage: backend: s3 bucket: my-kafka-backups region: us-west-2 prefix: backups/production # Optional key prefix # Optional: Custom endpoint for MinIO, Ceph RGW, etc. A custom endpoint # implies path-style requests; an http:// endpoint implies allow_http (v0.22.0+) endpoint: https://minio.example.com:9000 # path_style: true # force path-style without a custom endpoint # allow_http: true # plain-HTTP endpoints (in-cluster stores) # Credentials (optional - uses AWS credential chain if not specified). # access_key_id / secret_access_key are accepted as aliases (v0.22.0+). access_key: ${AWS_ACCESS_KEY_ID} secret_key: ${AWS_SECRET_ACCESS_KEY} ``` ``` storage: backend: azure account_name: mystorageaccount container_name: kafka-backups prefix: backups/production # endpoint: https://mystorageaccount.blob.core.usgovcloudapi.net # sovereign clouds # Credentials — precedence: sas_token → account_key → service principal # (client_secret) → Workload Identity → DefaultAzureCredential chain account_key: ${AZURE_STORAGE_KEY} # sas_token: ${AZURE_SAS_TOKEN} # client_id: ${AZURE_CLIENT_ID} # tenant_id: ${AZURE_TENANT_ID} # client_secret: ${AZURE_CLIENT_SECRET} # use_workload_identity: true # AKS; auto-enabled when AZURE_FEDERATED_TOKEN_FILE is set ``` ``` storage: backend: gcs bucket: my-kafka-backups prefix: backups/production # Credentials (uses GOOGLE_APPLICATION_CREDENTIALS if not specified) service_account_path: /etc/gcp/service-account.json ``` * * * Options specific to backup mode. ``` backup: # Compression settings compression: zstd # Options: zstd, lz4, none compression_level: 3 # 1-22 for zstd (default: 3) # Starting offset start_offset: earliest # earliest, latest, or specific offset # Segment settings segment_max_bytes: 134217728 # 128 MB - roll segment after this size segment_max_interval_ms: 60000 # 60 sec - roll segment after this time segment_max_records: 2000000 # Roll segment after this many records (v0.16.0+, default: unlimited) # Fetch sizing (v0.16.0+) # Maximum bytes requested per Kafka Fetch request. When unset, fetches # request min(segment_max_bytes, 16 MB). fetch_max_bytes: 16777216 # Continuous backup mode continuous: false # true for streaming backup # Internal topics include_internal_topics: false # Include __consumer_offsets, etc. # Checkpointing on_missing_topic: fail # or warn — skip absent literal topics and record them (v0.22.0+) sync_interval_secs: 30 # Sync offsets.db to storage every 30 seconds (default) # Offset-tracking headers (default: true). Adds x-original-offset and # x-original-timestamp to every archived record; needed for header-based # consumer offset recovery. Set false for a verbatim archive. include_offset_headers: true # Source cluster identifier source_cluster_id: "prod-cluster-east" # Snapshot backup mode (v0.8.0+) # Captures high watermarks at start and exits when all partitions reach them. # Incompatible with continuous: true. stop_at_current_offsets: false # Performance tuning (v0.8.0+) max_concurrent_partitions: 8 # Parallel partition processing (default: 8) poll_interval_ms: 100 # Delay between poll attempts in ms (default: 100) # Consumer group snapshot (v0.12.0+) # After each backup cycle, queries every broker for consumer groups (KRaft-safe), # fetches their committed offsets, filters to groups with offsets on backed-up # topics, and writes {backup_id}/consumer-groups-snapshot.json to storage. # Use kafka-backup snapshot-groups --config backup.yaml to run on demand instead. consumer_group_snapshot: false ``` | Option | Type | Default | Description | | --- | --- | --- | --- | | `compression` | string | `zstd` | Compression algorithm | | `compression_level` | int | `3` | Compression level (1-22 for zstd) | | `start_offset` | string | `earliest` | Starting offset: `earliest`, `latest` | | `segment_max_bytes` | int | `134217728` | Max segment size in bytes | | `segment_max_interval_ms` | int | `60000` | Max segment duration in ms | | `segment_max_records` | int | unlimited | Roll segment after this many records (v0.16.0+) | | `fetch_max_bytes` | int | `min(segment_max_bytes, 16777216)` | Max bytes requested per Kafka Fetch request (v0.16.0+) | | `continuous` | bool | `false` | Enable continuous/streaming backup | | `checkpoint_interval_secs` | int | `5` | **Deprecated (v0.22.0) — no effect.** Offsets are checkpointed at the end of every cycle; a warning is logged if set | | `sync_interval_secs` | int | `30` | How often the manifest and offset store are synced to storage (`offset_storage.sync_interval_secs` overrides it for the offset store, v0.22.0+) | | `include_internal_topics` | bool | `false` | Include internal Kafka topics | | `internal_topics` | list | `[]` | Internal topics to include when `include_internal_topics` is set (e.g. `__consumer_offsets`) | | `on_missing_topic` | string | `fail` | Literal `topics.include` entry absent from the cluster: `fail` errors; `warn` logs, records the names in the manifest's `missing_topics`, exposes `kafka_backup_missing_topics`, and continues (v0.22.0+). Globs matching nothing are always skipped | | `include_offset_headers` | bool | `true` | Add `x-original-offset` / `x-original-timestamp` headers to every archived record — see [Offset-tracking headers](#offset-tracking-headers) | | `source_cluster_id` | string | \- | Recorded in the `x-source-cluster` header (only with `include_offset_headers`) | | `stop_at_current_offsets` | bool | `false` | Snapshot mode: stop after reaching current high watermark (v0.8.0+) | | `max_concurrent_partitions` | int | `8` | Maximum parallel partition processing (v0.8.0+) | | `poll_interval_ms` | int | `100` | Delay in ms between consumer poll attempts (v0.8.0+) | | `consumer_group_snapshot` | bool | `false` | Write consumer group offsets to storage after each cycle — KRaft-safe, queries all brokers (v0.12.0+) | `include_offset_headers` is **on by default**. Every record is archived with two extra headers appended after the record's own headers: | Header | Value | Added when | | --- | --- | --- | | `x-original-offset` | source offset, little-endian `i64` | always | | `x-original-timestamp` | source timestamp (epoch ms), little-endian `i64` | always | | `x-source-cluster` | `source_cluster_id`, UTF-8 | `source_cluster_id` is set | They are what makes `consumer_group_strategy: header-based` recovery possible after a restore. The trade-off is that an archived — and therefore a restored — record is not header-for-header identical to its source. The backup logs which headers it is adding at startup (`include_offset_headers=true (default): ...`, v0.19.0+). To get a verbatim copy, either set `backup.include_offset_headers: false` at backup time, or set `restore.strip_offset_headers: true` (v0.19.0+) at restore time — that works for archives that already carry the headers and does not affect offset mapping, because the source offset is stored natively in the segment. > [!NOTE] > > [!NOTE] > > Record fidelity > > [!NOTE] > > Keys, values, timestamps and headers are restored verbatim, including the difference between a **null** and an **empty** key, value or header value. Null header values were archived as empty by versions before v0.18.0 ([#155](https://github.com/osodevops/kafka-backup/issues/155)); archives taken with those versions cannot be repaired and must be re-taken where the distinction matters. Duplicate header keys on one record are currently collapsed to the last one ([#156](https://github.com/osodevops/kafka-backup/issues/156)). > [!NOTE] > > [!NOTE] > > Unknown keys are warned about (v0.16.0+) > > [!NOTE] > > Config keys that kafka-backup does not recognize are logged as warnings at startup, for example: > > > > ``` > > WARN Ignoring unknown config key `backup.fetch_max_bytez` — check for typos; see https://kafkabackup.com/reference/config-yaml > > ``` > > > > Older versions silently ignored unrecognized keys, so a typo — or an option newer than the running binary — simply did nothing. If an option does not seem to take effect, check the job logs for this warning and confirm your kafka-backup version supports the option. * * * Configuration for the local SQLite database used to track backup progress. When present, this enables **resumable incremental backups** in any mode — one-shot, snapshot, or continuous. If the process exits and restarts with the same `backup_id`, it picks up where it left off. For continuous mode (`continuous: true`), the offset store is created automatically even without this section. For **one-shot** or **snapshot** mode, add this section explicitly to enable incremental behavior. > [!NOTE] > > [!NOTE] > > New in v0.13.5 > > [!NOTE] > > Prior to v0.13.5, the offset store was only created in continuous mode. One-shot and snapshot backups always started from `start_offset` (default: `earliest`), producing a full backup every time. From v0.13.5, adding `offset_storage` to your config enables incremental behavior in all modes. ``` offset_storage: backend: sqlite # sqlite (default) or memory db_path: /data/offsets.db # Path to local SQLite database sync_interval_secs: 30 # How often to sync to remote storage ``` | Option | Type | Default | Description | | --- | --- | --- | --- | | `backend` | string | `sqlite` | Storage backend: `sqlite` or `memory` | | `db_path` | string | `./offsets.db` | Path to local SQLite database file when `offset_storage` is present | | `s3_key` | string | \- | Remote storage key for syncing | | `sync_interval_secs` | int | `30` | How often to sync local DB to remote storage | > [!TIP] > > [!NOTE] > > Default path > > [!NOTE] > > When `offset_storage` is omitted, continuous mode creates its database in the system temporary directory as `{backup_id}-offsets.db`. When the section is present and `db_path` is omitted, the field default is `./offsets.db`. > > > > For persistent local state across pod restarts, configure `db_path` to point to a mounted volume (e.g. `/data/offsets.db`). Note that offsets are also synced to remote storage, so local persistence is optional. * * * ``` metrics: enabled: true port: 8080 bind_address: "0.0.0.0" path: /metrics update_interval_ms: 500 keep_alive_seconds: 60 max_partition_labels: 100 ``` | Option | Type | Default | Description | | --- | --- | --- | --- | | `enabled` | bool | `true` | Start the Prometheus HTTP server | | `port` | int | `8080` | Listener port | | `bind_address` | string | `0.0.0.0` | Listener address | | `path` | string | `/metrics` | Metrics endpoint path | | `update_interval_ms` | int | `500` | Metrics recalculation interval | | `keep_alive_seconds` | int | `0` | Keep serving after a one-shot operation completes | | `max_partition_labels` | int | `100` | Maximum unique per-partition label sets; `0` is unlimited | For a short-lived Kubernetes Job, set `keep_alive_seconds` to at least twice the Prometheus scrape interval. See the [metrics reference](https://kafkabackup.com/reference/metrics.md) for cardinality behavior and metric names. * * * Options specific to restore mode. ``` restore: # Point-in-Time Recovery (PITR) time_window_start: 1701417600000 # Unix milliseconds (optional) time_window_end: 1701504000000 # Unix milliseconds (optional) # Auto-create topics if they don't exist (v0.3.0+) create_topics: true default_replication_factor: 3 # Replication factor for new topics # Partition filtering source_partitions: # Only restore specific partitions - 0 - 1 - 2 # Partition mapping (remap partitions during restore) partition_mapping: 0: 0 1: 2 # Source partition 1 -> target partition 2 # Topic remapping topic_mapping: orders: orders_restored # orders -> orders_restored payments: payments_dr # payments -> payments_dr # Consumer offset strategy consumer_group_strategy: skip # skip, header-based, timestamp-based, cluster-scan, manual # Dry run mode dry_run: false # Validate without executing # Add x-original-offset / x-original-timestamp / x-source-partition to every # restored record (default: false; implied by consumer_group_strategy: header-based) include_original_offset_header: true # Remove the headers the backup added (x-original-*, x-source-*) before # producing, for a header-for-header identical restore (v0.19.0+, default: false) strip_offset_headers: false # Rate limiting rate_limit_records_per_sec: null # null for unlimited rate_limit_bytes_per_sec: null # null for unlimited # Performance tuning max_concurrent_partitions: 4 # Parallel partition processing produce_batch_size: 1000 # Records per produce batch # Resumable restores checkpoint_state: null # Path to checkpoint file checkpoint_interval_secs: 60 # Checkpoint frequency # Offset mapping report offset_report: /tmp/offset-mapping.json # Save offset mapping # Consumer group offset reset reset_consumer_offsets: false # Reset offsets after restore consumer_groups: # Groups to reset - my-consumer-group - analytics-consumer # Strimzi / pre-populated topic support (v0.12.0+) # Purge target topics via DeleteRecords before restoring. # Advances log-start-offset to the current end — makes the topic appear empty # without deleting it (safe for Strimzi-managed KafkaTopic resources). purge_topics: false # Automatically load consumer-groups-snapshot.json from storage (v0.12.0+) # Reads {backup_id}/consumer-groups-snapshot.json and uses those group IDs # to drive offset reset after restore. Sets reset_consumer_offsets: true # automatically. Silently ignored when the snapshot file is absent. auto_consumer_groups: false # Producer reliability tuning (v0.12.0+) # -1 = all in-sync replicas (default, highest durability) # 1 = leader only (faster, reduced durability — use when ISR is healthy) # 0 = fire-and-forget (fastest, no delivery guarantee) produce_acks: -1 produce_timeout_ms: 30000 # Broker-side produce timeout in milliseconds ``` | Option | Type | Default | Description | | --- | --- | --- | --- | | `time_window_start` | int | \- | PITR start timestamp (Unix ms) | | `time_window_end` | int | \- | PITR end timestamp (Unix ms) | | `create_topics` | bool | `false` | Auto-create topics if they don't exist (v0.3.0+) | | `default_replication_factor` | int | `1` | Replication factor for auto-created topics | | `source_partitions` | list | \- | Partitions to restore | | `partition_mapping` | map | \- | Partition remapping | | `topic_mapping` | map | \- | Topic remapping | | `consumer_group_strategy` | string | `skip` | Offset handling strategy | | `dry_run` | bool | `false` | Validate without executing | | `include_original_offset_header` | bool | `false` | Add `x-original-offset` / `x-original-timestamp` / `x-source-partition` headers to every restored record (also implied by `consumer_group_strategy: header-based`) | | `strip_offset_headers` | bool | `false` | Remove the headers kafka-backup added at backup time (`x-original-*`, `x-source-*`) before producing, for a header-for-header identical restore — see [Offset-tracking headers](#offset-tracking-headers) (v0.19.0+) | | `checkpoint_state` | string | \- | Path to the restore checkpoint file (resumable restores) | | `checkpoint_interval_secs` | int | `60` | Restore checkpoint frequency | | `offset_report` | string | \- | Path to write the offset-mapping report | | `repartitioning` | map | \- | Per target topic: `{strategy: murmur2 | automatic, target_partitions: N}`; mutually exclusive with `partition_mapping` | | `rate_limit_records_per_sec` | int | \- | Rate limit (records/sec) | | `rate_limit_bytes_per_sec` | int | \- | Rate limit (bytes/sec) | | `max_concurrent_partitions` | int | `4` | Parallel partitions | | `produce_batch_size` | int | `1000` | Batch size | | `reset_consumer_offsets` | bool | `false` | Reset consumer offsets | | `consumer_groups` | list | \- | Consumer groups to reset | | `purge_topics` | bool | `false` | Delete all records from target topic partitions before restoring via `DeleteRecords` API — safe for Strimzi (v0.12.0+) | | `auto_consumer_groups` | bool | `false` | Load consumer group IDs from `consumer-groups-snapshot.json` and reset their offsets after restore (v0.12.0+) | | `produce_acks` | int | `-1` | Producer ack level: `-1` all ISR (default, safest), `1` leader only (faster), `0` fire-and-forget (v0.12.0+) | | `produce_timeout_ms` | int | `30000` | Broker-side produce timeout in milliseconds (v0.12.0+) | | Strategy | Description | | --- | --- | | `skip` | Don't modify consumer offsets | | `header-based` | Use offset mapping from backup headers | | `timestamp-based` | Reset to timestamp-based offsets | | `cluster-scan` | Scan target cluster for offset mapping | | `manual` | Generate script for manual reset | * * * ``` mode: backup backup_id: "daily-backup" source: bootstrap_servers: - kafka:9092 topics: include: - orders - payments storage: backend: filesystem path: "/data/backups" backup: compression: zstd ``` ``` mode: backup backup_id: "prod-backup-${BACKUP_DATE}" source: bootstrap_servers: - broker-1.prod.kafka:9092 - broker-2.prod.kafka:9092 - broker-3.prod.kafka:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA256 sasl_username: backup-service sasl_password: ${KAFKA_PASSWORD} ssl_ca_location: /etc/kafka/ca.crt # Connection settings (recommended for cloud Kafka services) connection: tcp_keepalive: true keepalive_time_secs: 60 keepalive_interval_secs: 20 tcp_nodelay: true topics: include: - "*" exclude: - "__consumer_offsets" - "_schemas" - "*-internal" storage: backend: s3 bucket: company-kafka-backups region: us-west-2 prefix: production/${CLUSTER_NAME} backup: compression: zstd compression_level: 5 checkpoint_interval_secs: 60 include_offset_headers: true source_cluster_id: "prod-us-west-2" metrics: enabled: true keep_alive_seconds: 60 max_partition_labels: 100 ``` ``` mode: restore backup_id: "prod-backup-20241201" target: bootstrap_servers: - dr-broker-1:9092 - dr-broker-2:9092 storage: backend: s3 bucket: company-kafka-backups region: us-west-2 prefix: production/prod-cluster restore: # Restore only data from Dec 1, 2024 10:00 to 14:00 UTC time_window_start: 1701424800000 time_window_end: 1701439200000 topic_mapping: orders: orders_restored consumer_group_strategy: header-based reset_consumer_offsets: true consumer_groups: - order-processor - analytics-service ``` ``` mode: restore backup_id: "prod-backup-latest" target: bootstrap_servers: - dr-broker-1.dr.kafka:9092 - dr-broker-2.dr.kafka:9092 - dr-broker-3.dr.kafka:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA256 sasl_username: restore-service sasl_password: ${DR_KAFKA_PASSWORD} ssl_ca_location: /etc/kafka/ca.crt storage: backend: s3 bucket: company-kafka-backups region: us-east-1 # DR region prefix: production/prod-cluster restore: dry_run: false max_concurrent_partitions: 8 produce_batch_size: 5000 consumer_group_strategy: header-based reset_consumer_offsets: true consumer_groups: - order-service - payment-service - notification-service - analytics-pipeline ``` * * * Configuration files support `${VAR_NAME}` substitution. All `${...}` patterns are expanded from the process environment before YAML parsing. This works in any config value — passwords, paths, bucket names, backup IDs, etc. ``` source: security: sasl_password: ${KAFKA_PASSWORD} storage: backend: s3 access_key: ${AWS_ACCESS_KEY_ID} secret_key: ${AWS_SECRET_ACCESS_KEY} ``` Set variables before running: ``` export KAFKA_PASSWORD="secret123" export AWS_ACCESS_KEY_ID="AKIA..." export AWS_SECRET_ACCESS_KEY="..." kafka-backup backup --config backup.yaml ``` **Behavior notes:** - Unset variables are replaced with an empty string and produce a warning in the log - A bare `$` without `{` is left unchanged (e.g., `$5` stays as `$5`) - Expansion happens at all config loading sites: backup, restore, three-phase restore, validate-restore, and status watch --- title: Storage Format description: OSO Kafka Backup storage layout and file format reference source_url: html: https://kafkabackup.com/reference/storage-format md: https://kafkabackup.com/reference/storage-format.md --- # Storage Format This document describes the storage layout and file formats used by OSO Kafka Backup. Backups are organized in a hierarchical directory structure: ``` / ├── / │ ├── manifest.json # Backup metadata │ ├── topics/ │ │ ├── / │ │ │ ├── partition=/ │ │ │ │ ├── segment-00000000000000000000.bin.zst # Segment files, named by first offset │ │ │ │ ├── segment-00000000000000104857.bin.zst │ │ │ │ ├── segment-00000000000000209714.bin.zst │ │ │ │ └── ... │ │ │ └── ... │ │ └── ... │ ├── offsets.db # SQLite offset store; only with continuous: true or offset_storage: │ └── consumer-groups-snapshot.json # Only with backup.consumer_group_snapshot: true ├── / │ └── ... ├── evidence-reports////.json | .pdf | .sig # validation evidence └── offset-snapshots//snapshot.json | metadata.json # offset-rollback snapshots ``` `` is the filesystem `path`, or `/` for S3, Azure and GCS (`prefix` is joined as `{prefix}/{key}`; filesystem and memory backends have no prefix). CLI `--path` URLs: `s3://bucket/prefix`, `az://account.blob.core.windows.net/container`, `gs://bucket`, `file:///path`. Only `s3://` keeps a path prefix — `az://` / `gs://` discard it, so use the YAML `prefix:` key for those ([#162](https://github.com/osodevops/kafka-backup/issues/162)). There is no `state/`, `checkpoints/` or `checkpoint.json`; backup progress lives in `offsets.db` (see [Progress state](#progress-state)). `manifest.json` is the index of the backup: which topics and partitions were captured and where every segment lives. It is written by the backup engine and **merged** on every save — segments are de-duplicated by `(key, start_offset)` with the existing entry winning, so re-runs and resumes never drop segments. ``` { "backup_id": "production-backup-001", "created_at": 1733220000000, "source_cluster_id": "prod-cluster-east", "source_brokers": [], "compression": "zstd", "topics": [ { "name": "orders", "original_partition_count": 3, "partitions": [ { "partition_id": 0, "segments": [ { "key": "production-backup-001/topics/orders/partition=0/segment-00000000000000000000.bin.zst", "start_offset": 0, "end_offset": 124999, "start_timestamp": 1733216400123, "end_timestamp": 1733219999871, "record_count": 125000, "uncompressed_size": 268435456, "compressed_size": 41943040 } ] }, { "partition_id": 1, "segments": [ "..." ], "gaps": [ { "start_offset": 90000, "end_offset": 90512, "reason": "offset_out_of_range", "detected_at": 1733218800500 } ] } ] } ] } ``` | Field | Type | Description | | --- | --- | --- | | `backup_id` | string | Backup identifier (also the key prefix) | | `created_at` | int | Creation time, epoch milliseconds | | `source_cluster_id` | string | null | `backup.source_cluster_id`, if set | | `source_brokers` | array | Reserved; currently always `[]` | | `compression` | string | Segment codec: `zstd`, `lz4` or `none` (level is not recorded) | | `topics[].name` | string | Topic name | | `topics[].original_partition_count` | int | Partition count on the source at backup time; drives `create_topics` on restore | | `topics[].partitions[].partition_id` | int | Source partition | | `topics[].partitions[].segments[].key` | string | **Full storage key** of the segment — this is how restore and `validate` locate it | | `…segments[].start_offset` / `end_offset` | int | First / last source offset in the segment | | `…segments[].start_timestamp` / `end_timestamp` | int | Earliest / latest record timestamp (epoch ms); used for PITR segment selection | | `…segments[].record_count` | int | Records in the segment | | `…segments[].uncompressed_size` / `compressed_size` | int | Bytes before / after compression | | `topics[].partitions[].gaps[]` | array | Present only when retention deleted offsets before they were fetched: `start_offset`, `end_offset`, `reason` (`offset_out_of_range`), `detected_at`. Reported by `validate` and `describe`. | Totals (records, segments) are computed from the manifest on read; there is no `statistics`, `time_range` or `version` object. Segment files hold the record data of one partition, in offset order. They are named `segment-.bin.` — `.zst` for zstd, `.lz4` for lz4, `.bin` for none — and live under `topics//partition=/`. ``` ┌─────────────────────────────────────────┐ │ Segment Header (32 bytes) │ ├─────────────────────────────────────────┤ │ Compressed record block │ │ (Record 1, Record 2, ... Record N) │ ├─────────────────────────────────────────┤ │ Segment Footer (8 bytes) │ └─────────────────────────────────────────┘ ``` All integers are **little-endian**. | Offset | Size | Field | Description | | --- | --- | --- | --- | | 0 | 4 | Magic | `KBAK` | | 4 | 1 | Version | Format version (currently `1`) | | 5 | 1 | Compression | `0` = none, `1` = zstd, `2` = lz4 | | 6 | 2 | Reserved | Zero | | 8 | 8 | Record Count | Number of records in the segment (u64) | | 16 | 8 | Start Offset | Offset of the first record (i64) | | 24 | 8 | End Offset | Offset of the last record (i64) | The record block between header and footer is compressed as one unit with the codec named in the header. There is no per-batch framing: once decompressed, it is a plain sequence of records. | Offset | Size | Field | Description | | --- | --- | --- | --- | | 0 | 4 | CRC32 | CRC32 of every preceding byte (header + compressed block) | | 4 | 4 | End Magic | `BKAE` | Each record in the decompressed block is length-prefixed. Lengths of `-1` mean **null**, which is different from `0` (empty) — Kafka distinguishes the two for keys, values and header values, and so does the segment. ``` ┌─────────────────────────────────────────┐ │ Total Length (u32) — bytes that follow │ ├─────────────────────────────────────────┤ │ Timestamp (i64, epoch ms) │ ├─────────────────────────────────────────┤ │ Offset (i64, source offset) │ ├─────────────────────────────────────────┤ │ Key Length (i32, -1 = null) │ │ Key bytes │ ├─────────────────────────────────────────┤ │ Value Length (i32, -1 = null) │ │ Value bytes │ ├─────────────────────────────────────────┤ │ Header Count (u16) │ │ Headers, in order: │ │ Key Length (u16), Key bytes (UTF-8) │ │ Value Length (i32, -1 = null), bytes │ └─────────────────────────────────────────┘ ``` Because the source offset is stored in every record, restore can build the source→target offset mapping without relying on any header. With `backup.include_offset_headers: true` (the default), the backup appends these headers **after** the record's own headers: | Header Key | Value | Added when | | --- | --- | --- | | `x-original-offset` | source offset, little-endian `i64` (8 bytes) | always | | `x-original-timestamp` | source timestamp (epoch ms), little-endian `i64` | always | | `x-source-cluster` | `backup.source_cluster_id`, UTF-8 | `source_cluster_id` is set | A restore with `include_original_offset_header: true` (or `consumer_group_strategy: header-based`) adds `x-original-offset`, `x-original-timestamp` and `x-source-partition` (little-endian `i32`) to the records it produces; `restore.strip_offset_headers: true` (v0.19.0+) removes all of these from archived records before producing. See [Offset-tracking headers](https://kafkabackup.com/reference/config-yaml.md#offset-tracking-headers). > [!NOTE] > > [!NOTE] > > Fidelity across versions > > [!NOTE] > > Versions before v0.18.0 wrote a null header value with length `0` instead of `-1` ([#155](https://github.com/osodevops/kafka-backup/issues/155)); such archives cannot be repaired. Duplicate header keys on one record are collapsed to the last one at backup time ([#156](https://github.com/osodevops/kafka-backup/issues/156)). Segments written before the binary format are a compressed JSON array of records with base64-encoded `key`, `value` and header `value` fields (`null` for a null value). Restore recognises them by the missing `KBAK` magic and still reads them. There is no `checkpoint.json`. Two separate mechanisms track progress: The backup engine keeps per-partition progress in a local SQLite database and uploads that file byte-for-byte to `/offsets.db`: ``` CREATE TABLE IF NOT EXISTS offsets ( backup_id TEXT NOT NULL, topic TEXT NOT NULL, partition INTEGER NOT NULL, last_offset INTEGER NOT NULL, checkpoint_ts INTEGER NOT NULL DEFAULT (strftime('%s', 'now') * 1000), PRIMARY KEY (backup_id, topic, partition) ); CREATE TABLE IF NOT EXISTS backup_jobs ( backup_id TEXT PRIMARY KEY, source_cluster_id TEXT, status TEXT NOT NULL DEFAULT 'running', created_at INTEGER NOT NULL DEFAULT (strftime('%s', 'now') * 1000), last_heartbeat INTEGER NOT NULL DEFAULT (strftime('%s', 'now') * 1000), last_checkpoint INTEGER ); ``` - The store exists only when `backup.continuous: true` **or** an `offset_storage:` section is configured. A one-shot backup without either has no `offsets.db` and cannot resume. - Local path: `offset_storage.db_path`, else `$TMPDIR/-offsets.db`. - Progress is written after every fetched batch and synced to storage at most every `backup.sync_interval_secs` (default 30 s), at the end of every cycle, and on completion (`backup_jobs.status = 'completed'`). - **Resume:** at startup the engine downloads `/offsets.db` if the local database has no rows, then starts each partition at `last_offset + 1`. > [!NOTE] > > [!NOTE] > > Options that are currently ignored > > [!NOTE] > > `backup.checkpoint_interval_secs`, `offset_storage.s3_key` and `offset_storage.sync_interval_secs` are accepted but have no effect — the key is always `/offsets.db` and the cadence is `backup.sync_interval_secs` ([#161](https://github.com/osodevops/kafka-backup/issues/161)). `restore.checkpoint_state` names a **local** JSON file (it is not written to object storage): ``` { "backup_id": "production-backup-001", "start_time": 1733220000000, "last_checkpoint_time": 1733220315000, "segments_completed": [ "production-backup-001/topics/orders/partition=0/segment-00000000000000000000.bin.zst" ], "segments_in_progress": [["production-backup-001/topics/orders/partition=1/segment-00000000000000000000.bin.zst", 12582912]], "records_restored": 125000, "bytes_restored": 268435456, "config_hash": "a3f1c9…" } ``` > [!WARNING] > > [!NOTE] > > caution > > [!NOTE] > > The file is only _loaded and updated_ if it already exists — a fresh restore does not create it, so resumable restores currently require a pre-existing checkpoint file ([#160](https://github.com/osodevops/kafka-backup/issues/160)). On S3, Azure Blob and GCS the same tree maps to object keys under the configured prefix: ``` s3://my-bucket/kafka/ ├── backup-001/manifest.json ├── backup-001/offsets.db ├── backup-001/consumer-groups-snapshot.json ├── backup-001/topics/orders/partition=0/segment-00000000000000000000.bin.zst ├── backup-001/topics/orders/partition=0/segment-00000000000000125000.bin.zst ├── backup-001/topics/orders/partition=1/segment-00000000000000000000.bin.zst └── backup-002/... ``` `kafka-backup list --path s3://my-bucket/kafka` discovers backups by listing `*/manifest.json`. kafka-backup writes every object with the bucket's default storage class; there is no `storage_class` option ([#163](https://github.com/osodevops/kafka-backup/issues/163)). Use a bucket lifecycle policy to transition older backups to infrequent-access or archive tiers, remembering that objects in `GLACIER` / `DEEP_ARCHIVE` must be restored before a `kafka-backup restore` can read them. | `backup.compression` | Extension | Notes | | --- | --- | --- | | `zstd` (default) | `.zst` | Level from `backup.compression_level` (default 3; only zstd uses it) | | `lz4` | `.lz4` | `lz4_flex` block format with a size prefix — not the standard LZ4 frame format | | `none` | _(none)_ | Raw record block | Compression covers only the record block between the 32-byte header and the 8-byte footer, so a `.zst` segment is **not** a standalone zstd file — `zstd -d` will not decompress it; use `kafka-backup validate --deep` or a restore to read it. There is no gzip or snappy segment codec (broker-side gzip/snappy batches are decoded on fetch). Every segment carries one checksum: a CRC32 of the header plus the compressed record block, stored in the footer before the `BKAE` end magic. It is verified on **every** read (restore included), not only during validation. There is no per-record or per-batch checksum. ``` # Manifest loads; every segment key exists; sizes match compressed_size; gaps reported kafka-backup validate --path /data --backup-id backup-001 # Additionally opens every segment: end magic + CRC32, decompresses, and checks # record_count / start_offset / end_offset against the manifest kafka-backup validate --path /data --backup-id backup-001 --deep ``` The report is printed to stdout (JSON with `--format json`); nothing is written back to storage. A missing or unreadable manifest, missing segment, or corrupt segment makes the command exit non-zero. - **Segment format version** is `1` (`KBAK` header byte 4). Readers reject any other version outright; there is no forward-compatibility mode. - **Manifest** has no version field. All fields added since the first release (`source_cluster_id`, `source_brokers`, `compression`, `original_partition_count`, `gaps`, `uncompressed_size`, `compressed_size`) default when absent, so old manifests load in new readers, and new manifests load in old readers as long as the base fields are present. - **Legacy JSON segments** (pre-binary-format) are still restorable — see [Legacy JSON segments](#legacy-json-segments). Compressed legacy segments could not be decompressed before v0.19.2. `validate --deep` does not understand them and reports them as corrupt. - **Null header values** are written correctly from v0.18.0; archives from earlier versions store them as empty and cannot be repaired ([#155](https://github.com/osodevops/kafka-backup/issues/155)). --- title: Error Codes description: OSO Kafka Backup error codes, messages, and troubleshooting source_url: html: https://kafkabackup.com/reference/error-codes md: https://kafkabackup.com/reference/error-codes.md --- # Error Codes This reference documents all error types in OSO Kafka Backup with causes and solutions. | Category | Description | | --- | --- | | **CONFIG** | Configuration file errors | | **KAFKA** | Kafka connection and protocol errors | | **STORAGE** | Storage backend errors | | **VALIDATION** | Data validation errors | | **IO** | File system I/O errors | | **OFFSET** | Consumer offset management errors | * * * **Message:** `Failed to parse configuration file` **Cause:** Invalid YAML syntax in configuration file. **Solution:** 1. Validate YAML syntax: `yamllint backup.yaml` 2. Check for missing colons, incorrect indentation 3. Ensure proper quoting of special characters ``` # Example: Missing colon bootstrap_servers # Wrong - broker:9092 bootstrap_servers: # Correct - broker:9092 ``` **Message:** `Missing required configuration: {field}` **Cause:** Required field not specified in configuration. **Solution:** Add the missing field to your configuration: ``` # Required fields for backup: mode: backup # Required backup_id: "my-backup" # Required source: bootstrap_servers: # Required - broker:9092 storage: backend: filesystem # Required path: "/data" # Required for filesystem ``` **Message:** `Invalid value for {field}: {value}` **Cause:** Field value is not valid. **Common cases:** - Invalid compression algorithm - Invalid security protocol - Invalid SASL mechanism - Compression level out of range ``` # Valid compression values: zstd, lz4, none compression: gzip # Invalid # Valid security protocols security_protocol: INVALID # Invalid security_protocol: SASL_SSL # Valid # Valid zstd compression levels: 1-22 compression_level: 25 # Invalid (max is 22) ``` **Message:** `At least one topic must be specified` **Cause:** No topics configured for backup. **Solution:** ``` source: topics: include: - my-topic # Add at least one topic - "pattern-*" # Or a pattern ``` * * * **Message:** `Failed to connect to Kafka cluster: {details}` **Cause:** Cannot establish connection to Kafka brokers. **Solutions:** 1. **Verify brokers are reachable:** ``` nc -zv broker-1 9092 telnet broker-1 9092 ``` 2. **Check DNS resolution:** ``` nslookup broker-1.kafka.svc ``` 3. **Verify bootstrap servers in config:** ``` source: bootstrap_servers: - broker-1.kafka.svc:9092 # Use correct hostname/port ``` 4. **Check network policies/firewalls** **Message:** `Authentication failed: {details}` **Cause:** SASL authentication failed. **Solutions:** 1. **Verify credentials:** ``` source: security: sasl_username: correct-user sasl_password: correct-password ``` 2. **Check SASL mechanism matches broker:** ``` sasl_mechanism: SCRAM-SHA256 # Must match broker config ``` 3. **Verify user exists on Kafka cluster:** ``` kafka-configs --bootstrap-server broker:9092 \ --describe --entity-type users --entity-name backup-user ``` **Message:** `SSL/TLS error: {details}` **Cause:** TLS certificate or configuration error. **Solutions:** 1. **Verify certificate paths exist:** ``` ls -la /etc/kafka/ca.crt ls -la /etc/kafka/client.crt ls -la /etc/kafka/client.key ``` 2. **Check certificate validity:** ``` openssl x509 -in /etc/kafka/client.crt -text -noout openssl verify -CAfile /etc/kafka/ca.crt /etc/kafka/client.crt ``` 3. **Verify certificate permissions:** ``` chmod 600 /etc/kafka/client.key ``` **Message:** `Topic not found: {topic}` **Cause:** Specified topic doesn't exist on the cluster. **Solutions:** 1. **List available topics:** ``` kafka-topics --bootstrap-server broker:9092 --list ``` 2. **Check topic name spelling and case** 3. **Verify topic patterns:** ``` topics: include: - "events-*" # Check pattern matches actual topics ``` **Message:** `Unauthorized: {details}` **Cause:** User lacks required permissions. **Required ACLs for backup:** ``` Topic: Read, Describe Group: Read (for offset tracking) Cluster: DescribeConfigs (optional) ``` **Solution:** ``` kafka-acls --bootstrap-server broker:9092 \ --add --allow-principal User:backup-user \ --operation Read --operation Describe \ --topic '*' ``` * * * **Message:** `Access denied to storage: {path}` **Cause:** Insufficient permissions for storage backend. **Solutions by backend:** **Filesystem:** ``` # Check permissions ls -la /var/lib/kafka-backup/ # Fix permissions sudo chown -R kafka-backup:kafka-backup /var/lib/kafka-backup/ chmod -R 755 /var/lib/kafka-backup/ ``` **S3:** ``` # Test S3 access aws s3 ls s3://my-bucket/ # Check IAM permissions (need s3:GetObject, s3:PutObject, s3:ListBucket) ``` **Azure:** ``` # Test Azure access az storage blob list --container-name kafka-backups --account-name myaccount ``` **Message:** `Storage location not found: {path}` **Cause:** Specified path, bucket, or container doesn't exist. **Solutions:** ``` # Filesystem: Create directory mkdir -p /var/lib/kafka-backup/data # S3: Create bucket aws s3 mb s3://my-kafka-backups # Azure: Create container az storage container create --name kafka-backups ``` **Message:** `Storage quota exceeded` **Cause:** Insufficient storage space. **Solutions:** 1. **Check available space:** ``` df -h /var/lib/kafka-backup/ ``` 2. **Clean up old backups:** ``` kafka-backup list --path /data # Remove old backups manually or implement retention ``` 3. **Increase storage allocation** **Message:** `Invalid storage credentials` **Cause:** Cloud storage credentials are incorrect or expired. **Solutions:** **S3:** ``` # Test credentials aws sts get-caller-identity # Check environment variables echo $AWS_ACCESS_KEY_ID echo $AWS_SECRET_ACCESS_KEY ``` **Azure:** ``` # Test credentials az storage account show --name myaccount ``` **GCS:** ``` # Test credentials gcloud auth application-default print-access-token ``` * * * **Message:** `Backup not found: {backup_id}` **Cause:** Specified backup ID doesn't exist in storage. **Solutions:** 1. **List available backups:** ``` kafka-backup list --path /data ``` 2. **Check backup ID spelling** 3. **Verify storage path is correct** **Message:** `Backup manifest is corrupted` **Cause:** The manifest.json file is missing or invalid. **Solutions:** 1. **Check if manifest exists:** ``` cat /data/backup-001/manifest.json | jq . ``` 2. **Restore from a different backup** 3. **Contact support if critical data** **Message:** `Segment validation failed: {segment}` **Cause:** Segment file is corrupted (checksum mismatch). **Solutions:** 1. **Run deep validation:** ``` kafka-backup validate --path /data --backup-id backup-001 --deep ``` 2. **Check for partial writes (backup may have been interrupted)** 3. **Restore from a different backup point** **Message:** `Invalid PITR time window: start must be before end` **Cause:** time\_window\_start is greater than time\_window\_end. **Solution:** ``` restore: time_window_start: 1701417600000 # Earlier timestamp time_window_end: 1701504000000 # Later timestamp ``` * * * **Message:** `Failed to reset consumer offsets: {details}` **Cause:** Error resetting consumer group offsets. **Solutions:** 1. **Stop consumers first:** ``` # Consumer group must be inactive kafka-consumer-groups --bootstrap-server broker:9092 \ --group my-group --describe ``` 2. **Check group permissions:** ``` kafka-acls --bootstrap-server broker:9092 \ --add --allow-principal User:backup-user \ --operation Read --operation Describe \ --group my-group ``` **Message:** `Offset snapshot not found: {snapshot_id}` **Cause:** Referenced snapshot doesn't exist. **Solution:** ``` # List available snapshots kafka-backup offset-rollback list --path /data/snapshots ``` **Message:** `Offset mapping not available for backup` **Cause:** Backup was created with `include_offset_headers: false` (the default is `true`), or the archive predates offset headers. **Solution:** Re-run backup with offset headers enabled: ``` backup: include_offset_headers: true ``` * * * **Message:** `Failed to read file: {path}` **Cause:** File read operation failed. **Solutions:** 1. **Check file exists and permissions:** ``` ls -la /path/to/file ``` 2. **Check disk health:** ``` dmesg | grep -i error smartctl -a /dev/sda ``` **Message:** `Failed to write file: {path}` **Cause:** File write operation failed. **Solutions:** 1. **Check available disk space:** ``` df -h ``` 2. **Check directory permissions:** ``` ls -la /path/to/directory/ ``` 3. **Check if filesystem is read-only:** ``` mount | grep /data ``` * * * | Code | Meaning | | --- | --- | | 0 | Success | | 1 | General error | | 2 | Configuration error | | 3 | Connection error | | 4 | Authentication error | | 5 | Storage error | | 6 | Validation error | If you encounter an error not listed here: 1. **Enable verbose logging:** ``` kafka-backup -vv backup --config backup.yaml ``` 2. **Check logs for stack traces** 3. **Search [GitHub Issues](https://github.com/osodevops/kafka-backup/issues)** 4. **Open a new issue** with: - Error message - Configuration (sanitized) - Verbose output - Environment details --- title: Metrics Reference description: Prometheus metrics emitted by OSO Kafka Backup v0.19.2 source_url: html: https://kafkabackup.com/reference/metrics md: https://kafkabackup.com/reference/metrics.md --- # Metrics Reference OSO Kafka Backup `v0.19.2` exposes Prometheus text-format metrics while a backup or restore is running. ``` metrics: enabled: true port: 8080 bind_address: "0.0.0.0" path: "/metrics" update_interval_ms: 500 keep_alive_seconds: 60 max_partition_labels: 100 ``` | Field | Default | Description | | --- | --- | --- | | `enabled` | `true` | Start the HTTP metrics server | | `port` | `8080` | Listener port | | `bind_address` | `0.0.0.0` | Listener address | | `path` | `/metrics` | Metrics path | | `update_interval_ms` | `500` | Metrics recalculation interval | | `keep_alive_seconds` | `0` | Continue serving after a one-shot operation completes | | `max_partition_labels` | `100` | Maximum unique per-partition label sets; `0` is unlimited | Use `keep_alive_seconds` for short-lived Jobs so the final values live through at least one scrape. A practical minimum is twice your Prometheus scrape interval. Shutdown signals still stop the process promptly during this window. The partition limit applies to unique `topic`/`partition`/`backup_id` label sets used by `kafka_backup_lag_records`, `kafka_backup_lag_bytes`, and `kafka_backup_lag_seconds`. - Repeated updates to an admitted partition do not consume more slots. - Admitted partitions continue updating after the limit is reached. - `0` disables the limit; use it only when the number of partitions is bounded. - Aggregate lag and finite snapshot progress include all partitions, including those omitted by the per-partition limit. | Metric | Type | Labels | Description | | --- | --- | --- | --- | | `kafka_backup_lag_records` | Gauge | `topic`, `partition`, `backup_id` | Current records behind for a partition | | `kafka_backup_lag_bytes` | Gauge | `topic`, `partition`, `backup_id` | Estimated bytes behind for a partition | | `kafka_backup_lag_seconds` | Gauge | `topic`, `partition`, `backup_id` | Time lag for a partition | | `kafka_backup_lag_records_max` | Gauge | `backup_id` | Maximum current partition lag | | `kafka_backup_lag_records_sum` | Gauge | `backup_id` | Current lag summed across every partition | | `kafka_backup_snapshot_records_target` | Gauge | `backup_id` | Records this run will fetch: the captured finite-backup range minus offsets earlier runs already archived (checkpoints) | | `kafka_backup_snapshot_records_remaining` | Gauge | `backup_id` | Records this run has still to fetch; `0` when it completes | | `kafka_backup_throughput_records_per_sec` | Gauge | `backup_id`, `topic` | Current record throughput | | `kafka_backup_throughput_bytes_per_sec` | Gauge | `backup_id`, `topic` | Current byte throughput | | `kafka_backup_records_total` | Counter | `backup_id` | Records backed up | | `kafka_backup_bytes_total` | Counter | `backup_id` | Uncompressed bytes backed up | | `kafka_backup_duration_seconds` | Histogram | `backup_id`, `status` | Backup duration | | `kafka_backup_last_successful_commit` | Gauge | `backup_id`, `topic`, `partition` | Unix timestamp of the last offset commit | | `kafka_backup_consumer_rebalance_total` | Counter | `backup_id` | Consumer group rebalances | > [!NOTE] > > [!NOTE] > > note > > [!NOTE] > > kafka-backup versions before v0.15.12 emitted counters with a doubled suffix (for example `kafka_backup_records_total_total`). If you are running an older version, append an extra `_total` to the counter names below. | Metric | Type | Labels | Description | | --- | --- | --- | --- | | `kafka_backup_storage_write_latency_seconds` | Histogram | `backend`, `operation` | Storage write latency | | `kafka_backup_storage_read_latency_seconds` | Histogram | `backend`, `operation` | Storage read latency | | `kafka_backup_storage_write_bytes_total` | Counter | `backend`, `backup_id` | Bytes written to storage | | `kafka_backup_storage_read_bytes_total` | Counter | `backend`, `backup_id` | Bytes read from storage | | `kafka_backup_storage_errors_total` | Counter | `backend`, `error_type` | Storage errors | | `kafka_backup_compression_ratio` | Gauge | `algorithm`, `backup_id` | Uncompressed-to-compressed ratio | | `kafka_backup_compressed_bytes_total` | Counter | `algorithm`, `backup_id` | Compressed bytes written | | `kafka_backup_uncompressed_bytes_total` | Counter | `algorithm`, `backup_id` | Bytes before compression | | `kafka_backup_errors_total` | Counter | `backup_id`, `error_type` | Operation errors | | `kafka_backup_retries_total` | Counter | `backup_id`, `operation` | Retry attempts | Storage operations are labelled `segment`, `manifest`, `checkpoint`, or `other`. Error categories include `broker_connection`, `consumer_timeout`, `deserialization`, `storage_io`, `codec`, `offset_invalid`, `auth`, `timeout`, `not_found`, `quota`, and `unknown`. | Metric | Type | Labels | Description | | --- | --- | --- | --- | | `kafka_restore_duration_seconds` | Histogram | `backup_id`, `status` | Restore duration | | `kafka_restore_progress_percent` | Gauge | `restore_id`, `backup_id` | Completion percentage | | `kafka_restore_eta_seconds` | Gauge | `restore_id` | Estimated seconds remaining | | `kafka_restore_throughput_records_per_sec` | Gauge | `restore_id` | Current record throughput | | `kafka_restore_records_total` | Counter | `restore_id` | Records restored | Finite snapshot completion percentage: ``` 100 * ( 1 - kafka_backup_snapshot_records_remaining / clamp_min(kafka_backup_snapshot_records_target, 1) ) ``` Both gauges are scoped to the run that exports them. An incremental one-shot backup (`offset_storage` plus `stop_at_current_offsets`) plans only the records after its checkpoints, so a run with a large archive but few new records reports a small target that counts down to `0` — not the size of the whole archive (kafka-backup v0.19.1+). A run with nothing new reports `0` / `0`. Each scheduled run is a new process, so group by pod (or `job`) in dashboards rather than reading the last sample of an older series. Current total lag without a high-cardinality sum: ``` kafka_backup_lag_records_sum ``` Record throughput over five minutes: ``` rate(kafka_backup_records_total[5m]) ``` 99th-percentile storage write latency: ``` histogram_quantile( 0.99, sum by (le, backend, operation) ( rate(kafka_backup_storage_write_latency_seconds_bucket[5m]) ) ) ``` Errors during the last five minutes: ``` sum by (backup_id, error_type) ( increase(kafka_backup_errors_total[5m]) ) ``` ``` scrape_configs: - job_name: kafka-backup static_configs: - targets: ["kafka-backup-host:8080"] metrics_path: /metrics scrape_interval: 30s ``` For Kubernetes Jobs managed by the Strimzi Backup Operator, use its [Job PodMonitor configuration](https://kafkabackup.com/strimzi-operator/metrics.md). A ServiceMonitor cannot discover a short-lived Job pod unless a matching Service is also created. --- title: Validation Config Reference description: Complete reference for all validation.yaml configuration options source_url: html: https://kafkabackup.com/reference/validation-config md: https://kafkabackup.com/reference/validation-config.md --- # Validation Config Reference Complete reference for the `validation.yaml` configuration file used with `kafka-backup validation run`. ``` # Required: Backup to validate against backup_id: "my-backup-001" # Required: Where the backup is stored storage: backend: s3 bucket: my-bucket # Required: The restored Kafka cluster target: bootstrap_servers: [] # Optional: Which checks to run checks: {} # Optional: Evidence report settings evidence: {} # Optional: Notification settings notifications: {} # Optional: PITR timestamp and trigger metadata pitr_timestamp: null triggered_by: null ``` * * * **Required.** The backup ID to load the manifest from. ``` backup_id: "production-daily-001" ``` **Required.** Storage backend configuration. Uses the same format as backup/restore configs. ``` storage: backend: s3 bucket: my-kafka-backups region: us-west-2 prefix: production/daily ``` See the [Storage Configuration](https://kafkabackup.com/reference/config-yaml.md#storage-configuration) section for all backend options (S3, Azure, GCS, Filesystem). **Required.** The Kafka cluster where the backup was restored (the cluster to validate). ``` target: bootstrap_servers: - restored-kafka:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA-512 sasl_username: validator sasl_password: "${KAFKA_PASSWORD}" ``` Supports the same authentication options as the backup/restore `source`/`target` config. See [Kafka Cluster Configuration](https://kafkabackup.com/reference/config-yaml.md#kafka-cluster-configuration). **Optional.** PITR timestamp (epoch milliseconds) used during the restore. Included in the evidence report metadata. ``` pitr_timestamp: 1711929600000 ``` Can also be set via `--pitr` on the CLI. **Optional.** Human-readable string recording who or what triggered this validation run. Appears in the evidence report for chain of custody. ``` triggered_by: "weekly-cron-job" triggered_by: "External auditor KPMG - Q1 2026 review" ``` Can also be set via `--triggered-by` on the CLI. * * * Compares per-topic/partition record counts between the backup manifest and the restored cluster. | Option | Type | Default | Description | | --- | --- | --- | --- | | `enabled` | bool | `true` | Enable this check | | `mode` | string | `exact` | `exact` (all partitions) or `sample` (random subset) | | `sample_percentage` | int | `100` | Percentage of partitions to sample (1-100). Only used in `sample` mode | | `topics` | list | `[]` | Topic filter. Empty = all topics in the backup | | `fail_threshold` | int | `0` | Number of records difference allowed before failing. 0 = exact match | ``` checks: message_count: enabled: true mode: exact topics: - orders - payments fail_threshold: 0 ``` Verifies that high watermark and low watermark for each partition match the backup manifest. | Option | Type | Default | Description | | --- | --- | --- | --- | | `enabled` | bool | `true` | Enable this check | | `verify_high_watermark` | bool | `true` | Verify the high watermark matches | | `verify_low_watermark` | bool | `true` | Verify the low watermark matches | ``` checks: offset_range: enabled: true verify_high_watermark: true verify_low_watermark: true ``` Verifies that consumer group offsets are present and valid in the restored cluster. | Option | Type | Default | Description | | --- | --- | --- | --- | | `enabled` | bool | `true` | Enable this check | | `verify_all_groups` | bool | `true` | Verify all groups found on the cluster | | `groups` | list | `[]` | Specific groups to check. Empty + `verify_all_groups=true` = all | ``` checks: consumer_group_offsets: enabled: true verify_all_groups: false groups: - order-processor - payment-service ``` Call external HTTP endpoints for custom validation logic. Each webhook is a separate check in the report. | Option | Type | Default | Description | | --- | --- | --- | --- | | `name` | string | _required_ | Display name for this check in the report | | `url` | string | _required_ | URL to POST the validation payload to | | `timeout_seconds` | int | `120` | Request timeout in seconds | | `expected_status_code` | int | `200` | Expected HTTP status code for success | | `fail_on_timeout` | bool | `true` | Whether to treat a timeout as failure | ``` checks: custom_webhooks: - name: application-health-check url: "https://internal.example.com/kafka-validation-hook" timeout_seconds: 120 expected_status_code: 200 fail_on_timeout: true ``` The webhook receives a JSON POST body: ``` { "event": "kafka_backup_validation", "backup_id": "production-daily-001", "pitr_timestamp": null, "restored_cluster": { "bootstrap_servers": ["restored-kafka:9092"] } } ``` Expected response: ``` { "result": "passed", "detail": "All health checks passed", "data": {} } ``` Valid `result` values: `passed`, `failed`, `warning`, `skipped`. * * * List of output formats to generate. | Value | Description | | --- | --- | | `json` | Machine-readable JSON evidence report (canonical format when signing enabled) | | `pdf` | Auditor-ready branded PDF report | ``` evidence: formats: - json - pdf ``` Cryptographic signing configuration. | Option | Type | Default | Description | | --- | --- | --- | --- | | `enabled` | bool | `false` | Enable ECDSA-P256-SHA256 signing | | `private_key_path` | string | `null` | Path to PEM-encoded PKCS#8 private key | | `public_key_path` | string | `null` | Path to PEM-encoded public key (optional, for reference) | ``` evidence: signing: enabled: true private_key_path: "/etc/kafka-backup/signing-key.pem" ``` > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > The private key **must** be in PKCS#8 format (header: `-----BEGIN PRIVATE KEY-----`). See the [Evidence Signing Guide](https://kafkabackup.com/guides/evidence-signing.md) for key generation instructions. Controls where evidence reports are uploaded in object storage. | Option | Type | Default | Description | | --- | --- | --- | --- | | `prefix` | string | `evidence-reports/` | Storage key prefix for evidence files | | `retention_days` | int | `2555` | Retention period in days (~7 years for SOX) | ``` evidence: storage: prefix: "evidence-reports/" retention_days: 2555 # ~7 years ``` Evidence files are stored at: `{prefix}{run-id}/{YYYY}/{MM}/{run-id}.{json|pdf|sig}` * * * Send Slack notifications via incoming webhook. | Option | Type | Required | Description | | --- | --- | --- | --- | | `webhook_url` | string | Yes | Slack incoming webhook URL | ``` notifications: slack: webhook_url: "https://hooks.slack.com/services/T00/B00/xxxxx" ``` Send PagerDuty alerts via Events API v2. | Option | Type | Default | Description | | --- | --- | --- | --- | | `integration_key` | string | _required_ | PagerDuty Events API v2 integration key | | `severity` | string | `critical` | Alert severity: `critical`, `error`, `warning`, `info` | ``` notifications: pagerduty: integration_key: "your-integration-key" severity: critical ``` On failure, a trigger event is sent. On the next success, a resolve event is sent (using the backup ID as dedup key). - [Backup Validation Guide](https://kafkabackup.com/guides/validation-compliance.md) — step-by-step walkthrough - [Evidence Report Schema](https://kafkabackup.com/reference/evidence-report-schema.md) — JSON report structure - [CLI Reference](https://kafkabackup.com/reference/cli-reference.md) — `validation` command options --- title: Evidence Report Schema description: JSON schema reference for backup validation evidence reports source_url: html: https://kafkabackup.com/reference/evidence-report-schema md: https://kafkabackup.com/reference/evidence-report-schema.md --- # Evidence Report Schema The JSON evidence report is the canonical machine-readable output of a validation run. All other outputs (PDF, signature) are derived from it. Current schema version: **1.0** The `schema_version` field enables forward compatibility. Consumers should check this field and handle unknown versions gracefully. ``` { "schema_version": "1.0", "report_id": "validation-9275b4aa-2aeb-4910-a3a6-9e4aa1dc016a", "generated_at": "2026-04-09T05:37:49.204+00:00", "tool": { }, "backup": { }, "restore": { }, "validation": { }, "integrity": { }, "compliance_mappings": { }, "triggered_by": "weekly-cron-job" } ``` | Field | Type | Description | | --- | --- | --- | | `schema_version` | string | Schema version for compatibility (`"1.0"`) | | `report_id` | string | Unique identifier for this validation run | | `generated_at` | string | ISO-8601 timestamp of report generation | | `tool` | object | Information about the tool that generated the report | | `backup` | object | Backup metadata | | `restore` | object | Restore details (nullable) | | `validation` | object | Validation check results | | `integrity` | object | Checksums and signature information | | `compliance_mappings` | object | Regulatory framework mappings | | `triggered_by` | string | Who/what triggered this run (nullable) | * * * ``` { "tool": { "name": "kafka-backup", "version": "0.11.0" } } ``` ``` { "backup": { "id": "production-daily-001", "source_cluster_id": "prod-kafka-01", "source_brokers": ["broker-1:9092", "broker-2:9092"], "storage_backend": "s3://my-kafka-backups/production/daily", "pitr_timestamp": 1711929600000, "created_at": 1711929600000, "total_topics": 3, "total_partitions": 9, "total_segments": 9, "total_records": 1000 } } ``` | Field | Type | Description | | --- | --- | --- | | `id` | string | Backup ID | | `source_cluster_id` | string | Source cluster identifier (nullable) | | `source_brokers` | array | Source bootstrap servers | | `storage_backend` | string | Storage backend description | | `pitr_timestamp` | int | PITR timestamp in epoch millis (nullable) | | `created_at` | int | Backup creation timestamp in epoch millis | | `total_topics` | int | Number of topics in the backup | | `total_partitions` | int | Total partitions across all topics | | `total_segments` | int | Total segment files | | `total_records` | int | Total records across all segments | Present when validation was run against a restored cluster. ``` { "restore": { "target_bootstrap_servers": ["restored-kafka:9092"], "start_time": "2026-04-09T02:00:00Z", "end_time": "2026-04-09T02:31:14Z", "duration_seconds": 1874 } } ``` ``` { "validation": { "checks_total": 3, "checks_passed": 3, "checks_failed": 0, "checks_skipped": 0, "checks_warned": 0, "overall_result": "PASSED", "results": [ ], "total_duration_ms": 18 } } ``` | Field | Type | Description | | --- | --- | --- | | `overall_result` | string | `PASSED`, `FAILED`, `WARNING`, or `SKIPPED` | | `checks_total` | int | Total checks executed | | `checks_passed` | int | Checks that passed | | `checks_failed` | int | Checks that failed | | `total_duration_ms` | int | Total validation duration in milliseconds | | `results` | array | Individual check results (see below) | ``` { "check_name": "MessageCountCheck", "outcome": "PASSED", "detail": "3 topics; 1000 messages expected, 1000 restored; 0 discrepancies", "data": { "check": "MessageCountCheck", "topics_verified": 3, "total_messages_expected": 1000, "total_messages_restored": 1000, "discrepancies": [], "sha256_offset_summary": "e3b0c44298fc1c149afb..." }, "duration_ms": 12 } ``` | Field | Type | Description | | --- | --- | --- | | `check_name` | string | Check identifier | | `outcome` | string | `PASSED`, `FAILED`, `WARNING`, or `SKIPPED` | | `detail` | string | Human-readable summary | | `data` | object | Machine-readable data (varies by check type) | | `duration_ms` | int | Check execution time in milliseconds | | Field | Type | Description | | --- | --- | --- | | `topics_verified` | int | Number of topics checked | | `total_messages_expected` | int | Expected record count from manifest | | `total_messages_restored` | int | Actual record count on restored cluster | | `discrepancies` | array | Per-partition discrepancies (empty when passed) | | `sha256_offset_summary` | string | SHA-256 of the offset summary for evidence | | Field | Type | Description | | --- | --- | --- | | `partitions_checked` | int | Total partitions verified | | `partitions_passed` | int | Partitions with correct offset ranges | | `issues` | array | Per-partition issues (empty when passed) | | Field | Type | Description | | --- | --- | --- | | `groups_checked` | int | Consumer groups verified | | `total_offsets` | int | Total committed offsets checked | | `issues` | array | Per-group issues (empty when passed) | ``` { "integrity": { "backup_manifest_sha256": "a1b2c3d4e5f6...", "report_sha256": "2482bbdfa113146e...", "checksums_valid": true, "signature_algorithm": "ECDSA-P256-SHA256", "signed_by": "production-signing-key" } } ``` | Field | Type | Description | | --- | --- | --- | | `backup_manifest_sha256` | string | SHA-256 of the backup manifest.json | | `report_sha256` | string | SHA-256 of the canonical JSON report | | `checksums_valid` | bool | Whether checksums are consistent | | `signature_algorithm` | string | `ECDSA-P256-SHA256` or `none` | | `signed_by` | string | Signing identity (nullable) | ``` { "compliance_mappings": { "sox_itgc": { "control": "IT General Controls - Backup and Recovery", "satisfied_by": ["MessageCountCheck", "OffsetRangeCheck"], "evidence_retention_required_years": 7, "evidence_retention_configured_days": 2555 }, "cmmc_l2": { "control": "RE.3.139", "description": "Regularly perform and test data back-ups", "satisfied_by": ["MessageCountCheck", "OffsetRangeCheck", "ConsumerGroupOffsetCheck"] }, "gdpr_art32": { "control": "Article 32 - Testing technical measures", "satisfied_by": ["MessageCountCheck", "OffsetRangeCheck"], "test_frequency": "on-demand", "rto_demonstrated_seconds": 1874 } } } ``` Each framework mapping identifies which validation checks satisfy which regulatory control, enabling auditors to trace evidence directly to compliance requirements. The PDF is generated from the JSON report and contains: | Page | Content | | --- | --- | | 1 | Cover — title, report ID, timestamp, overall result (PASSED/FAILED), tool version | | 2 | Validation results — check summary table, per-check detail | | 3 | Integrity & compliance — SHA-256 checksums, signature details, SOX/CMMC/GDPR mappings | The PDF includes the `report_sha256` so auditors can independently verify the JSON report matches by computing its SHA-256. - [Validation Config Reference](https://kafkabackup.com/reference/validation-config.md) — configuration options - [Evidence Signing Guide](https://kafkabackup.com/guides/evidence-signing.md) — signing and verification - [Backup Validation Guide](https://kafkabackup.com/guides/validation-compliance.md) — complete walkthrough --- title: Index source_url: html: https://kafkabackup.com/operator/index md: https://kafkabackup.com/operator/index.md --- --- title: Installation description: Install the OSO Kafka Backup Operator source_url: html: https://kafkabackup.com/operator/installation md: https://kafkabackup.com/operator/installation.md --- # Installation Install the OSO Kafka Backup Operator in your Kubernetes cluster. - Kubernetes 1.26+ - Helm 3.0+ - kubectl configured to access your cluster - Cluster admin permissions ``` helm repo add oso https://osodevops.github.io/helm-charts/ helm repo update ``` ``` # Create namespace kubectl create namespace kafka-backup # Install with default values helm install kafka-backup-operator oso/kafka-backup-operator \ --namespace kafka-backup ``` ``` # Check operator pod kubectl get pods -n kafka-backup # Check CRDs installed kubectl get crds | grep kafka.oso.sh # Expected output: # kafkabackups.kafka.oso.sh # kafkarestores.kafka.oso.sh # kafkaoffsetresets.kafka.oso.sh # kafkaoffsetrollbacks.kafka.oso.sh # kafkabackupvalidations.kafka.oso.sh ``` ``` helm install kafka-backup-operator oso/kafka-backup-operator \ --namespace kafka-backup \ --set replicaCount=2 \ --set resources.requests.memory=256Mi \ --set metrics.enabled=true ``` values.yaml ``` replicaCount: 2 image: repository: ghcr.io/osodevops/kafka-backup-operator tag: "1.3.0" pullPolicy: Always resources: requests: cpu: 100m memory: 256Mi limits: cpu: 500m memory: 512Mi metrics: enabled: true serviceMonitor: enabled: true podSecurityContext: runAsNonRoot: true runAsUser: 1000 fsGroup: 1000 ``` ``` helm install kafka-backup-operator oso/kafka-backup-operator \ --namespace kafka-backup \ --values values.yaml ``` ``` helm install kafka-backup-operator oso/kafka-backup-operator \ --namespace kafka-backup \ --version 1.3.0 ``` CRDs are installed automatically by Helm. If you need to install CRDs separately (e.g., for GitOps): ``` # Install CRDs only kubectl apply -f https://github.com/osodevops/kafka-backup-operator/releases/download/v1.3.0/crds.yaml ``` By default, the chart keeps CRDs on uninstall. To make that explicit: ``` crds: install: true keep: true # Don't delete on uninstall ``` The operator creates a service account with required permissions: ``` apiVersion: v1 kind: ServiceAccount metadata: name: kafka-backup-operator namespace: kafka-backup ``` The operator needs these permissions: ``` apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: kafka-backup-operator rules: # Manage CRDs - apiGroups: ["kafka.oso.sh"] resources: ["kafkabackups", "kafkarestores", "kafkaoffsetresets", "kafkaoffsetrollbacks", "kafkabackupvalidations"] verbs: ["get", "list", "watch", "create", "update", "patch", "delete"] - apiGroups: ["kafka.oso.sh"] resources: ["kafkabackups/status", "kafkarestores/status", "kafkaoffsetresets/status", "kafkaoffsetrollbacks/status", "kafkabackupvalidations/status"] verbs: ["get", "patch", "update"] - apiGroups: ["kafka.oso.sh"] resources: ["kafkabackups/finalizers", "kafkarestores/finalizers", "kafkaoffsetresets/finalizers", "kafkaoffsetrollbacks/finalizers", "kafkabackupvalidations/finalizers"] verbs: ["update"] # Read secrets and manage PVCs used by storage-backed workflows - apiGroups: [""] resources: ["secrets"] verbs: ["get", "list", "watch"] - apiGroups: [""] resources: ["persistentvolumeclaims"] verbs: ["get", "list", "watch", "create", "update", "patch", "delete"] # Events - apiGroups: [""] resources: ["events"] verbs: ["create", "patch"] # Required only when leaderElection.enabled=true - apiGroups: ["coordination.k8s.io"] resources: ["leases"] verbs: ["get", "list", "watch", "create", "update", "patch", "delete"] ``` ``` serviceAccount: create: true name: "my-kafka-backup-sa" annotations: eks.amazonaws.com/role-arn: arn:aws:iam::123456789:role/kafka-backup-role ``` ``` # Create IAM policy aws iam create-policy \ --policy-name KafkaBackupPolicy \ --policy-document file://policy.json # Create IAM role with OIDC eksctl create iamserviceaccount \ --name kafka-backup-operator \ --namespace kafka-backup \ --cluster my-cluster \ --attach-policy-arn arn:aws:iam::123456789:policy/KafkaBackupPolicy \ --approve ``` policy.json ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "s3:PutObject", "s3:GetObject", "s3:ListBucket", "s3:DeleteObject" ], "Resource": [ "arn:aws:s3:::kafka-backups", "arn:aws:s3:::kafka-backups/*" ] } ] } ``` ``` serviceAccount: create: false name: kafka-backup-operator # Service account created by eksctl ``` ``` # Enable workload identity on AKS az aks update \ --resource-group myResourceGroup \ --name myAKSCluster \ --enable-oidc-issuer \ --enable-workload-identity # Create managed identity az identity create \ --name kafka-backup-identity \ --resource-group myResourceGroup # Get identity client ID CLIENT_ID=$(az identity show \ --name kafka-backup-identity \ --resource-group myResourceGroup \ --query clientId -o tsv) # Assign Storage Blob Data Contributor role az role assignment create \ --assignee $CLIENT_ID \ --role "Storage Blob Data Contributor" \ --scope /subscriptions//resourceGroups//providers/Microsoft.Storage/storageAccounts/ # Create federated credential az identity federated-credential create \ --name kafka-backup-federated \ --identity-name kafka-backup-identity \ --resource-group myResourceGroup \ --issuer $(az aks show --name myAKSCluster --resource-group myResourceGroup --query oidcIssuerProfile.issuerUrl -o tsv) \ --subject system:serviceaccount:kafka-backup:kafka-backup-operator \ --audience api://AzureADTokenExchange ``` Install the CSI Secrets Store Driver for Azure Key Vault integration: ``` # Add the Helm repo helm repo add csi-secrets-store-provider-azure \ https://azure.github.io/secrets-store-csi-driver-provider-azure/charts helm repo update # Install the Azure provider helm install csi-secrets-store-provider-azure \ csi-secrets-store-provider-azure/csi-secrets-store-provider-azure \ --namespace kube-system ``` Grant the managed identity access to Key Vault: ``` # Get the managed identity principal ID PRINCIPAL_ID=$(az identity show \ --name kafka-backup-identity \ --resource-group myResourceGroup \ --query principalId -o tsv) # Grant secret access to Key Vault az keyvault set-policy \ --name \ --object-id $PRINCIPAL_ID \ --secret-permissions get list ``` ``` serviceAccount: create: true name: kafka-backup-operator azureWorkloadIdentity: enabled: true clientId: extraEnv: - name: AZURE_CLIENT_ID value: ``` ``` # Enable workload identity on GKE gcloud container clusters update my-cluster \ --workload-pool=PROJECT_ID.svc.id.goog # Create GCP service account gcloud iam service-accounts create kafka-backup-sa # Bind to Kubernetes service account gcloud iam service-accounts add-iam-policy-binding \ kafka-backup-sa@PROJECT_ID.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:PROJECT_ID.svc.id.goog[kafka-backup/kafka-backup-operator]" ``` ``` replicaCount: 2 # Leader election is enabled by default leaderElection: enabled: true leaseDuration: 15s renewDeadline: 10s retryPeriod: 2s ``` ``` podDisruptionBudget: enabled: true minAvailable: 1 ``` ``` affinity: podAntiAffinity: preferredDuringSchedulingIgnoredDuringExecution: - weight: 100 podAffinityTerm: labelSelector: matchLabels: app.kubernetes.io/name: kafka-backup-operator topologyKey: kubernetes.io/hostname ``` ``` # Update the CRDs first; Helm does not upgrade existing CRDs from crds/ kubectl apply -f https://github.com/osodevops/kafka-backup-operator/releases/download/v1.3.0/crds.yaml helm repo update helm upgrade kafka-backup-operator oso/kafka-backup-operator \ --namespace kafka-backup \ --version 1.3.0 \ --values values.yaml ``` Helm installs CRDs on the initial chart installation, but deliberately does not upgrade existing CRDs from a chart's `crds/` directory. Apply the release bundle before upgrading the operator: ``` kubectl apply -f https://github.com/osodevops/kafka-backup-operator/releases/download/v1.3.0/crds.yaml ``` ``` helm uninstall kafka-backup-operator --namespace kafka-backup ``` If CRDs were not kept: ``` kubectl delete crds kafkabackups.kafka.oso.sh kubectl delete crds kafkarestores.kafka.oso.sh kubectl delete crds kafkaoffsetresets.kafka.oso.sh kubectl delete crds kafkaoffsetrollbacks.kafka.oso.sh kubectl delete crds kafkabackupvalidations.kafka.oso.sh ``` ``` kubectl delete namespace kafka-backup ``` ``` # Check pod logs kubectl logs -n kafka-backup deployment/kafka-backup-operator # Check events kubectl get events -n kafka-backup --sort-by='.lastTimestamp' ``` ``` # Verify CRDs exist kubectl get crds | grep kafka # Reinstall CRDs helm upgrade kafka-backup-operator oso/kafka-backup-operator \ --namespace kafka-backup \ --set crds.install=true ``` ``` # Check RBAC kubectl auth can-i create kafkabackups --as=system:serviceaccount:kafka-backup:kafka-backup-operator # Check cluster role binding kubectl get clusterrolebinding | grep kafka-backup ``` - [Configuration](https://kafkabackup.com/operator/configuration.md) - Configure operator settings - [Secrets Guide](https://kafkabackup.com/operator/guides/secrets.md) - Set up credentials - [KafkaBackup CRD](https://kafkabackup.com/operator/crds/kafkabackup.md) - Create your first backup --- title: Configuration description: Configure the OSO Kafka Backup Operator source_url: html: https://kafkabackup.com/operator/configuration md: https://kafkabackup.com/operator/configuration.md --- # Configuration Configure the OSO Kafka Backup Operator with Helm values. values.yaml ``` replicaCount: 1 image: repository: ghcr.io/osodevops/kafka-backup-operator pullPolicy: Always tag: "" # Defaults to the chart appVersion imagePullSecrets: [] nameOverride: "" fullnameOverride: "" serviceAccount: create: true annotations: {} name: "" azureWorkloadIdentity: enabled: false clientId: "" podAnnotations: prometheus.io/scrape: "true" prometheus.io/port: "8080" prometheus.io/path: "/metrics" podSecurityContext: runAsNonRoot: true runAsUser: 1000 fsGroup: 1000 securityContext: allowPrivilegeEscalation: false readOnlyRootFilesystem: true capabilities: drop: - ALL service: type: ClusterIP port: 8080 annotations: {} resources: requests: cpu: 100m memory: 128Mi limits: cpu: 500m memory: 512Mi nodeSelector: kubernetes.io/os: linux tolerations: [] affinity: {} extraVolumes: [] extraVolumeMounts: [] extraEnv: [] logging: level: "info,kafka_backup_operator=debug" format: json metrics: enabled: true serviceMonitor: enabled: false interval: 30s scrapeTimeout: 10s labels: {} crds: install: true keep: true leaderElection: enabled: false leaseDuration: 15s renewDeadline: 10s retryPeriod: 2s ``` ``` image: repository: ghcr.io/osodevops/kafka-backup-operator tag: "1.3.0" pullPolicy: Always ``` The chart defaults the image tag to `appVersion`, so chart `1.3.0` deploys `ghcr.io/osodevops/kafka-backup-operator:1.3.0` unless you override `image.tag`. ``` azureWorkloadIdentity: enabled: true clientId: ``` When `azureWorkloadIdentity.enabled` is true, the chart annotates the service account, labels the pod template with `azure.workload.identity/use: "true"`, and the operator can use federated identity for Azure Blob Storage when `storage.azure.useWorkloadIdentity: true` is set on a CRD. Use `extraVolumes`, `extraVolumeMounts`, and `extraEnv` for custom CA bundles or S3-compatible endpoint settings. ``` extraVolumes: - name: internal-ca secret: secretName: internal-ca-bundle extraVolumeMounts: - name: internal-ca mountPath: /etc/internal-certs readOnly: true extraEnv: - name: SSL_CERT_FILE value: /etc/internal-certs/ca.crt ``` ``` metrics: enabled: true serviceMonitor: enabled: true interval: 30s scrapeTimeout: 10s labels: release: prometheus ``` The operator exposes metrics on port `8080`. The Helm chart also sets default Prometheus scrape annotations on the pod. ``` replicaCount: 2 leaderElection: enabled: true leaseDuration: 15s renewDeadline: 10s retryPeriod: 2s ``` Enable leader election when running more than one replica. ``` helm repo add oso https://osodevops.github.io/helm-charts/ helm repo update helm upgrade --install kafka-backup-operator oso/kafka-backup-operator \ --namespace kafka-backup \ --create-namespace \ --values values.yaml ``` ``` helm template kafka-backup-operator oso/kafka-backup-operator \ --namespace kafka-backup \ --values values.yaml ``` ``` helm get values kafka-backup-operator -n kafka-backup --all ``` - [Metrics](https://kafkabackup.com/operator/metrics.md) - Prometheus metrics reference - [Secrets Guide](https://kafkabackup.com/operator/guides/secrets.md) - Configure credentials - [KafkaBackup CRD](https://kafkabackup.com/operator/crds/kafkabackup.md) - Create backups --- title: Metrics description: Prometheus metrics for the OSO Kafka Backup Operator source_url: html: https://kafkabackup.com/operator/metrics md: https://kafkabackup.com/operator/metrics.md --- # Metrics The OSO Kafka Backup Operator exposes controller and completed-operation metrics on port `8080`. | Endpoint | Description | | --- | --- | | `/metrics` | Prometheus metrics | | `/healthz` or `/health` | Liveness response | | `/readyz` or `/ready` | Readiness response | > [!NOTE] > > [!NOTE] > > Operator and CLI metrics are different > > [!NOTE] > > The operator runs `kafka-backup-core` in its own process. Its `/metrics` endpoint exposes the `kafka_backup_operator_*` families listed below; it does not expose the CLI's live `kafka_backup_lag_records`, throughput, or snapshot progress registry. Use the [Strimzi operator](https://kafkabackup.com/strimzi-operator/metrics.md) when you need separately scrapeable backup/restore Job pods. Metrics are enabled by default. With the Prometheus Operator installed, enable the chart's ServiceMonitor: values.yaml ``` metrics: enabled: true serviceMonitor: enabled: true interval: 30s scrapeTimeout: 10s labels: release: prometheus ``` ``` helm upgrade --install kafka-backup-operator oso/kafka-backup-operator \ --namespace kafka-backup \ --create-namespace \ --version 1.3.0 \ --values values.yaml ``` The chart also adds Prometheus scrape annotations to the operator pod by default. To inspect the endpoint directly: ``` kubectl port-forward -n kafka-backup \ svc/kafka-backup-operator-metrics 8080:8080 curl --fail http://127.0.0.1:8080/metrics ``` | Metric | Type | Labels | Description | | --- | --- | --- | --- | | `kafka_backup_operator_reconciliations_total` | Counter | `kind` | Reconciliation attempts | | `kafka_backup_operator_reconciliation_errors_total` | Counter | `kind` | Reconciliation errors | | `kafka_backup_operator_reconcile_duration_seconds` | Histogram | `kind` | Reconciliation duration | | `kafka_backup_operator_backups_total` | Counter | `outcome`, `namespace`, `name` | Completed backups by outcome | | `kafka_backup_operator_backup_size_bytes` | Gauge | `namespace`, `name` | Last backup size | | `kafka_backup_operator_backup_records_total` | Gauge | `namespace`, `name` | Records processed in the last backup | | `kafka_backup_operator_restores_total` | Counter | `outcome`, `namespace`, `name` | Completed restores by outcome | | `kafka_backup_operator_offset_resets_total` | Counter | `outcome`, `namespace` | Offset resets by outcome | | `kafka_backup_operator_offset_reset_duration_seconds` | Histogram | `namespace` | Offset reset duration | | `kafka_backup_operator_validations_total` | Counter | `outcome`, `namespace`, `name` | Validations by outcome | | `kafka_backup_operator_cleanups_total` | Counter | `kind` | Finalizer cleanup operations | | `kafka_backup_operator_health` | Gauge | none | Operator health (`1` is healthy) | `kafka_backup_operator_backup_records_total` is a gauge despite its historical name: it represents the most recently completed backup, not a monotonically increasing counter. Reconciliation error rate: ``` sum by (kind) ( rate(kafka_backup_operator_reconciliation_errors_total[5m]) ) ``` Failed backups during the last hour: ``` sum by (namespace, name) ( increase(kafka_backup_operator_backups_total{outcome="failure"}[1h]) ) ``` 99th-percentile reconciliation duration: ``` histogram_quantile( 0.99, sum by (le, kind) ( rate(kafka_backup_operator_reconcile_duration_seconds_bucket[5m]) ) ) ``` Operator unavailable: ``` absent(kafka_backup_operator_health) or kafka_backup_operator_health != 1 ``` operator-alerts.yaml ``` groups: - name: kafka-backup-operator rules: - alert: KafkaBackupOperatorDown expr: absent(kafka_backup_operator_health) or kafka_backup_operator_health != 1 for: 5m labels: severity: critical annotations: summary: Kafka Backup Operator is unavailable - alert: KafkaBackupOperatorReconcileErrors expr: increase(kafka_backup_operator_reconciliation_errors_total[10m]) > 0 for: 5m labels: severity: warning annotations: summary: Kafka Backup Operator reconciliation is failing - alert: KafkaBackupFailed expr: increase(kafka_backup_operator_backups_total{outcome="failure"}[1h]) > 0 labels: severity: critical annotations: summary: Kafka backup failed ``` `KafkaBackup.spec.metrics` is retained in the CRD and passed into the embedded core configuration. It does not create a second HTTP listener or add live core progress families to the operator endpoint. Monitor completion through CR status and the operator metric families above. --- title: Operator Troubleshooting description: Troubleshoot the OSO Kafka Backup Operator source_url: html: https://kafkabackup.com/operator/troubleshooting md: https://kafkabackup.com/operator/troubleshooting.md --- # Operator Troubleshooting This guide helps troubleshoot common issues with the Kafka Backup Operator. ``` # Operator pods kubectl get pods -n kafka-backup -l app.kubernetes.io/name=kafka-backup-operator # Operator logs kubectl logs -n kafka-backup deployment/kafka-backup-operator # Follow logs kubectl logs -n kafka-backup deployment/kafka-backup-operator -f # Previous container logs (if restarted) kubectl logs -n kafka-backup deployment/kafka-backup-operator --previous ``` ``` # List all backup resources kubectl get kafkabackups -A # Describe specific backup kubectl describe kafkabackup my-backup -n kafka-backup # Get backup YAML with status kubectl get kafkabackup my-backup -n kafka-backup -o yaml # List all CRDs kubectl get crds | grep kafka.oso.sh ``` ``` # Events in namespace kubectl get events -n kafka-backup --sort-by='.lastTimestamp' # Events for specific resource kubectl get events -n kafka-backup --field-selector involvedObject.name=my-backup ``` **Symptoms:** - Operator pod in CrashLoopBackOff or Error state **Check:** ``` kubectl describe pod -n kafka-backup -l app.kubernetes.io/name=kafka-backup-operator kubectl logs -n kafka-backup deployment/kafka-backup-operator --previous ``` **Common Causes:** 1. **Missing CRDs** ``` # Check CRDs exist kubectl get crds | grep kafka.oso.sh # Reinstall CRDs helm upgrade kafka-backup-operator oso/kafka-backup-operator \ --set crds.install=true ``` 2. **Insufficient permissions** ``` # Check RBAC kubectl auth can-i list kafkabackups --as=system:serviceaccount:kafka-backup:kafka-backup-operator # Check ClusterRoleBinding kubectl get clusterrolebinding | grep kafka-backup ``` 3. **Resource limits too low** ``` resources: requests: memory: 256Mi # Increase if OOMKilled limits: memory: 512Mi ``` **Symptoms:** - KafkaBackup resource created but no backup job running - Status shows "Pending" indefinitely **Check:** ``` kubectl get kafkabackup my-backup -o yaml | grep -A 20 status kubectl get jobs -n kafka-backup | grep my-backup ``` **Common Causes:** 1. **Schedule not triggered yet** ``` # Check schedule spec: schedule: "0 0 * * * * *" # Next hour # For immediate backup, remove schedule or create one-time backup ``` 2. **Previous job still running** ``` # Check for running jobs kubectl get jobs -n kafka-backup # Delete stuck job kubectl delete job my-backup-job-xyz -n kafka-backup ``` 3. **Invalid configuration** ``` # Check events kubectl get events -n kafka-backup --field-selector involvedObject.name=my-backup # Common: missing secret, wrong topic name, invalid storage config ``` **Symptoms:** - KafkaBackup status shows "Failed" - Job completed with error **Check:** ``` # Get job pod logs kubectl logs -n kafka-backup job/my-backup-job-xyz # Describe job kubectl describe job my-backup-job-xyz -n kafka-backup ``` **Common Causes:** 1. **Kafka connection failed** ``` Error: Failed to connect to Kafka broker ``` ``` # Verify bootstrap servers spec: kafkaCluster: bootstrapServers: - kafka.default.svc:9092 # Correct service name and port ``` 2. **Authentication failed** ``` Error: SASL authentication failed ``` ``` # Check secret exists and has correct keys kubectl get secret kafka-credentials -n kafka-backup -o yaml # Verify secret reference spec: kafkaCluster: saslSecret: name: kafka-credentials usernameKey: username # Must match secret key passwordKey: password # Must match secret key ``` 3. **Storage access denied** ``` Error: Access denied to S3 bucket ``` ``` # Check IAM/credentials kubectl get secret s3-credentials -n kafka-backup # For IRSA, check service account kubectl describe sa kafka-backup-operator -n kafka-backup ``` 4. **Topic not found** ``` Error: Topic 'my-topic' not found ``` ``` # Verify topic exists in Kafka kubectl exec -it kafka-0 -- kafka-topics --list --bootstrap-server localhost:9092 ``` **Symptoms:** - KafkaRestore stuck in "Running" or fails **Check:** ``` kubectl get kafkarestore my-restore -o yaml kubectl logs -n kafka-backup job/my-restore-job-xyz ``` **Common Causes:** 1. **Backup not found** ``` Error: Backup 'backup-xyz' not found ``` ``` # List available backups kafka-backup list --path s3://bucket/prefix # Check backupId matches spec: backupRef: name: production-backup backupId: "backup-20241201-120000" # Exact ID ``` 2. **Target cluster unreachable** ``` Error: Failed to connect to target cluster ``` ``` spec: kafkaCluster: bootstrapServers: - target-kafka:9092 # Verify accessibility ``` 3. **PITR time range invalid** ``` Error: No records found in time range ``` ``` spec: pitr: # Verify timestamp is within backup range endTime: "2024-12-01T12:00:00Z" ``` **Symptoms:** - Offset reset completes but consumers still at wrong position **Check:** ``` kubectl get kafkaoffsetreset my-reset -o yaml # Verify consumer group positions kafka-consumer-groups --bootstrap-server kafka:9092 --group my-group --describe ``` **Common Causes:** 1. **Consumers still running** ``` # Stop consumers before reset kubectl scale deployment my-consumer --replicas=0 # Then run reset ``` 2. **Wrong strategy** ``` spec: resetStrategy: from-mapping # Use restore offset mapping when available # Or reset to a timestamp: resetStrategy: to-timestamp resetTimestamp: 1701432000000 ``` 3. **Consumer group doesn't exist** ``` spec: consumerGroups: - my-consumer-group # Must match exact group ID ``` **Symptoms:** - "Secret not found" or "key not found" errors **Check:** ``` # List secrets kubectl get secrets -n kafka-backup # Check secret contents kubectl get secret kafka-credentials -n kafka-backup -o yaml ``` **Fix:** ``` # Ensure secret exists in same namespace as KafkaBackup apiVersion: v1 kind: Secret metadata: name: kafka-credentials namespace: kafka-backup # Same namespace! type: Opaque stringData: username: myuser password: mypassword ``` **Symptoms:** - Error when creating/updating resources - "spec.xxx: Invalid value" messages **Check:** ``` # Validate YAML kubectl apply -f backup.yaml --dry-run=client # Check CRD schema kubectl explain kafkabackup.spec kubectl explain kafkabackup.spec.storage ``` **Common validation errors:** ``` # Wrong: Missing required field spec: kafkaCluster: # bootstrapServers is required! # Right: spec: kafkaCluster: bootstrapServers: - kafka:9092 ``` ``` # Wrong: Invalid enum value spec: storage: storageType: s4 # Invalid # Right: spec: storage: storageType: s3 # s3, azure, gcs, pvc ``` ``` # Helm values logging: level: debug # Or via environment env: - name: RUST_LOG value: "kafka_backup_operator=debug,kafka_backup=debug" ``` ``` kubectl rollout restart deployment/kafka-backup-operator -n kafka-backup kubectl logs -n kafka-backup deployment/kafka-backup-operator -f ``` | Phase | Description | | --- | --- | | `Pending` | Waiting for next scheduled run | | `Running` | Backup in progress | | `Completed` | Backup finished successfully | | `Failed` | Backup failed | | Phase | Description | | --- | --- | | `Pending` | Waiting to start | | `Running` | Restore in progress | | `Completed` | Restore finished successfully | | `Failed` | Restore failed | ``` status: conditions: - type: Ready status: "True" reason: BackupCompleted message: "Backup completed successfully" lastTransitionTime: "2024-12-01T12:00:00Z" ``` ``` #!/bin/bash # collect-diagnostics.sh NAMESPACE="kafka-backup" OUTPUT_DIR="kafka-backup-diagnostics" mkdir -p $OUTPUT_DIR # Operator logs kubectl logs -n $NAMESPACE deployment/kafka-backup-operator > $OUTPUT_DIR/operator.log # All resources kubectl get kafkabackups -A -o yaml > $OUTPUT_DIR/kafkabackups.yaml kubectl get kafkarestores -A -o yaml > $OUTPUT_DIR/kafkarestores.yaml kubectl get kafkaoffsetresets -A -o yaml > $OUTPUT_DIR/kafkaoffsetresets.yaml kubectl get kafkaoffsetrollbacks -A -o yaml > $OUTPUT_DIR/kafkaoffsetrollbacks.yaml # Jobs kubectl get jobs -n $NAMESPACE -o yaml > $OUTPUT_DIR/jobs.yaml # Events kubectl get events -n $NAMESPACE --sort-by='.lastTimestamp' > $OUTPUT_DIR/events.txt # Pods kubectl get pods -n $NAMESPACE -o yaml > $OUTPUT_DIR/pods.yaml # Secrets (names only) kubectl get secrets -n $NAMESPACE > $OUTPUT_DIR/secrets.txt # Create archive tar -czf kafka-backup-diagnostics.tar.gz $OUTPUT_DIR echo "Diagnostics saved to kafka-backup-diagnostics.tar.gz" ``` - [GitHub Issues](https://github.com/osodevops/kafka-backup-operator/issues) - [Support Guide](https://kafkabackup.com/troubleshooting/support.md) - [Common Errors](https://kafkabackup.com/troubleshooting/common-errors.md) - CLI error reference - [Metrics](https://kafkabackup.com/operator/metrics.md) - Monitor operator health - [Configuration](https://kafkabackup.com/operator/configuration.md) - Operator settings --- title: KafkaBackup CRD description: KafkaBackup Custom Resource Definition reference source_url: html: https://kafkabackup.com/operator/crds/kafkabackup md: https://kafkabackup.com/operator/crds/kafkabackup.md --- # KafkaBackup CRD The `KafkaBackup` custom resource defines a backup run or recurring backup schedule for Kafka topics. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: production-backup namespace: kafka-backup spec: kafkaCluster: bootstrapServers: - kafka-0.kafka.svc:9092 - kafka-1.kafka.svc:9092 securityProtocol: SASL_SSL caSecret: # optional: separate CA secret (e.g. Strimzi) name: cluster-ca-cert caKey: ca.crt tlsSecret: name: kafka-tls caKey: ca.crt certKey: tls.crt keyKey: tls.key saslSecret: name: kafka-credentials mechanism: SCRAM-SHA-512 usernameKey: username passwordKey: password connection: tcpKeepalive: true keepaliveTimeSecs: 60 keepaliveIntervalSecs: 20 tcpNodelay: true connectionsPerBroker: 4 topics: - orders - payments - "events-*" storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 prefix: production/hourly endpoint: https://s3.us-west-2.amazonaws.com pathStyle: false allowHttp: false credentialsSecret: name: s3-credentials accessKeyIdKey: AWS_ACCESS_KEY_ID secretAccessKeyKey: AWS_SECRET_ACCESS_KEY schedule: "0 0 * * * * *" # Every hour, cron format includes seconds suspend: false compression: zstd compressionLevel: 3 segmentMaxBytes: 134217728 segmentMaxIntervalMs: 60000 # Backup mode continuous: false stopAtCurrentOffsets: true pollIntervalMs: 100 includeOffsetHeaders: true sourceClusterId: production-us-west-2 consumerGroupSnapshot: true checkpoint: enabled: true intervalSecs: 30 retention: enabled: false maxAgeDays: 30 keepLast: 3 dryRun: true rateLimiting: recordsPerSec: 0 bytesPerSec: 0 maxConcurrentPartitions: 4 circuitBreaker: enabled: true failureThreshold: 5 resetTimeoutSecs: 60 successThreshold: 3 operationTimeoutMs: 30000 metrics: enabled: true port: 9090 bindAddress: "0.0.0.0" path: /metrics updateIntervalMs: 500 maxPartitionLabels: 100 # unique topic/partition label sets; 0 is unlimited template: pod: hostAliases: - ip: "10.10.0.5" hostnames: - s3.internal - minio.internal ``` > [!NOTE] > > [!NOTE] > > Metrics endpoint > > [!NOTE] > > The standalone operator exposes its own `kafka_backup_operator_*` registry on port `8080`. Resource-level `spec.metrics` is passed to the embedded core configuration, but does not start another HTTP endpoint or expose live per-partition progress. See [operator metrics](https://kafkabackup.com/operator/metrics.md). `checkpoint.enabled` no longer makes a backup run continuously. Use these fields explicitly: | Field | Default | Description | | --- | --- | --- | | `continuous` | `false` | Keep polling and writing new records after the initial pass | | `stopAtCurrentOffsets` | `false` | Snapshot mode: capture starting high watermarks and exit after all partitions catch up | | `consumerGroupSnapshot` | `false` | Write `consumer-groups-snapshot.json` after each backup cycle | For scheduled point-in-time backups, set `stopAtCurrentOffsets: true`. For streaming backups, set `continuous: true`. Do not set `continuous` and `stopAtCurrentOffsets` together. Operator-managed `KafkaBackup` retention is disabled by default. When `spec.retention` is absent, or when `spec.retention.enabled: false`, the operator does not delete backup data. Enable retention per `KafkaBackup` resource when you want the operator to prune complete backup sets after a successful backup run: ``` spec: retention: enabled: true maxAgeDays: 30 keepLast: 3 dryRun: false ``` Retention deletes whole backup IDs, not individual segments, manifests, or offset files from a still-retained backup set. This avoids creating partially pruned manifests that would break point-in-time restore unexpectedly. When retention is enabled: - Set at least one of `maxAgeDays` or `keepLast`. - `maxAgeDays` and `keepLast` must be greater than `0` when set. - The current backup ID is always retained. - If only `maxAgeDays` is set, the operator still keeps at least the newest backup set. - `dryRun: true` reports what would be deleted without deleting data. - Retention failures are reported in status but do not turn a successful backup into a failed backup. Operator-managed retention is supported for PVC/local storage, S3/S3-compatible storage, and Azure Blob Storage. GCS retention is not currently wired through the operator; use a GCS bucket lifecycle policy for that backend. Storage lifecycle policies are still a good option when retention should be enforced outside Kubernetes, when you need backend-native legal hold or object-lock controls, or when you use GCS. Keep all retention windows aligned with restore requirements because deleting old backup sets makes older point-in-time restore windows unavailable. > [!NOTE] > > [!NOTE] > > info > > [!NOTE] > > `KafkaBackupValidation.spec.evidence.retentionDays` controls validation evidence retention only. It does not control `KafkaBackup` data retention. | Field | Type | Required | Description | | --- | --- | --- | --- | | `bootstrapServers` | `[]string` | Yes | Kafka broker addresses | | `securityProtocol` | `string` | No | `PLAINTEXT`, `SSL`, `SASL_PLAINTEXT`, or `SASL_SSL` | | `tlsSecret` | object | No | TLS certificate secret reference | | `caSecret` | object | No | Separate CA certificate secret (e.g. Strimzi cluster CA). Overrides `caKey` in `tlsSecret` when both are set | | `saslSecret` | object | No | SASL credentials secret reference | | `connection` | object | No | Kafka TCP connection tuning | | Field | Type | Default | Description | | --- | --- | --- | --- | | `tcpKeepalive` | bool | `true` | Enable TCP keepalive | | `keepaliveTimeSecs` | int | `60` | Seconds before the first keepalive probe | | `keepaliveIntervalSecs` | int | `20` | Seconds between keepalive probes | | `tcpNodelay` | bool | `true` | Enable TCP\_NODELAY | | `connectionsPerBroker` | int | `4` | TCP connections to maintain per broker | | Field | Type | Required | Description | | --- | --- | --- | --- | | `storageType` | `string` | No | `pvc`, `s3`, `azure`, or `gcs`; defaults to `pvc` | | `pvc` | object | When `storageType: pvc` | PVC storage configuration | | `s3` | object | When `storageType: s3` | S3 or S3-compatible storage configuration | | `azure` | object | When `storageType: azure` | Azure Blob Storage configuration | | `gcs` | object | When `storageType: gcs` | Google Cloud Storage configuration | | Field | Type | Required | Description | | --- | --- | --- | --- | | `bucket` | string | Yes | Bucket name | | `region` | string | Yes | AWS region | | `prefix` | string | No | Object key prefix | | `endpoint` | string | No | Custom S3-compatible endpoint | | `pathStyle` | bool | No | Force path-style addressing for MinIO, Ceph, and similar endpoints | | `allowHttp` | bool | No | Allow HTTP endpoint traffic; the operator logs a warning when enabled | | `credentialsSecret` | object | Yes | Secret containing access key credentials | | Field | Type | Required | Description | | --- | --- | --- | --- | | `accountName` | string | Yes | Storage account name | | `container` | string | Yes | Blob container name | | `prefix` | string | No | Blob prefix | | `endpoint` | string | No | Custom endpoint for sovereign cloud or private endpoint use | | `useWorkloadIdentity` | bool | No | Use AKS Workload Identity | | `credentialsSecret` | object | No | Account key secret | | `sasTokenSecret` | object | No | SAS token secret | | `servicePrincipalSecret` | object | No | Service principal secret | Azure authentication methods are mutually exclusive. If none are set, the adapter falls back to Azure default credentials when the runtime environment supports them. | Field | Type | Default | Description | | --- | --- | --- | --- | | `compression` | string | `zstd` | `none`, `lz4`, or `zstd` | | `compressionLevel` | int | `3` | Compression level | | `segmentMaxBytes` | int | `134217728` | Rotate backup segments after this many bytes | | `segmentMaxIntervalMs` | int | `60000` | Rotate backup segments after this many milliseconds | | `pollIntervalMs` | int | `100` | Poll interval for continuous mode | | `includeOffsetHeaders` | bool | `true` | Add `x-original-offset` / `x-original-timestamp` (and `x-source-cluster` when `sourceClusterId` is set) to every archived record. Because it defaults to `true`, archived records are not header-for-header identical to the source — set `false` here, or `stripOffsetHeaders` on the restore | | `sourceClusterId` | string | unset | Source cluster ID recorded in manifests and offset headers | | `schedule` | string | unset | Cron schedule. The operator uses the Rust `cron` syntax with seconds, for example `0 0 * * * * *` | | `suspend` | bool | `false` | Pause scheduled backups | | Field | Type | Default | Description | | --- | --- | --- | --- | | `enabled` | bool | `true` | Enable resumable checkpoints | | `intervalSecs` | int | `30` | Checkpoint interval | | `storage` | object | unset | Optional PVC checkpoint storage override | | Field | Type | Default | Description | | --- | --- | --- | --- | | `enabled` | bool | `false` | Enable operator-managed retention pruning | | `maxAgeDays` | int | unset | Delete complete backup sets older than this many days | | `keepLast` | int | unset | Keep at least this many newest backup sets | | `dryRun` | bool | `false` | Report eligible backup sets without deleting data | `retention.enabled: true` requires at least one of `maxAgeDays` or `keepLast`. Retention runs after a backup completes successfully. For long-running `continuous: true` backups, retention does not prune while the backup engine is still running. | Field | Type | Required | Description | | --- | --- | --- | --- | | `hostAliases` | array | No | Pod-level Kubernetes `hostAliases` entries injected into backup Job and CronJob pods. Use this for split-DNS or private endpoint storage access | Each `hostAliases` item uses the Kubernetes `HostAlias` shape: | Field | Type | Required | Description | | --- | --- | --- | --- | | `ip` | string | Yes | IP address to write into the pod `/etc/hosts` file | | `hostnames` | `[]string` | No | Hostnames mapped to `ip` | > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Prefer DNS or Kubernetes Services when possible. Use `hostAliases` as a narrow pod-local override for cases such as VPN-only S3-compatible endpoints or split-DNS storage names. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: hourly-snapshot spec: kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders storage: storageType: pvc pvc: claimName: kafka-backups schedule: "0 0 * * * * *" stopAtCurrentOffsets: true compression: zstd ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: hourly-snapshot spec: kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 prefix: hourly credentialsSecret: name: s3-credentials schedule: "0 0 * * * * *" stopAtCurrentOffsets: true retention: enabled: true maxAgeDays: 30 keepLast: 3 dryRun: true ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: continuous-backup spec: kafkaCluster: bootstrapServers: - kafka:9092 topics: - "*" storage: storageType: azure azure: accountName: kafkabackups123456 container: kafka-backups prefix: production useWorkloadIdentity: true continuous: true includeOffsetHeaders: true sourceClusterId: production consumerGroupSnapshot: true ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: minio-backup spec: kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders storage: storageType: s3 s3: bucket: kafka-backups region: us-east-1 endpoint: http://minio.storage.svc.cluster.local:9000 pathStyle: true allowHttp: true credentialsSecret: name: minio-credentials ``` Use `template.pod.hostAliases` when backup pods must reach an S3-compatible endpoint by its public hostname but resolve it to an internal IP address from inside the cluster. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: private-s3-backup spec: kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders storage: storageType: s3 s3: bucket: kafka-backups region: us-east-1 endpoint: https://s3.internal pathStyle: true credentialsSecret: name: s3-credentials template: pod: hostAliases: - ip: "10.10.0.5" hostnames: - s3.internal - minio.internal ``` ``` status: phase: Completed message: "Backup completed successfully" lastBackupTime: "2026-04-13T12:00:00Z" nextScheduledBackup: "2026-04-13T13:00:00Z" recordsProcessed: 1000000 bytesProcessed: 1073741824 segmentsCompleted: 8 checkpointEnabled: true lastCheckpointTime: "2026-04-13T12:05:00Z" resumable: true backupId: "production-backup-20260413-120000" lastRetentionTime: "2026-04-13T12:06:00Z" retentionInspectedBackups: 12 retentionEligibleBackups: 2 retentionDeletedBackups: 0 retentionReclaimedBytes: 536870912 retentionDryRun: true retentionError: null ``` - [KafkaRestore](https://kafkabackup.com/operator/crds/kafkarestore.md) - Restore from backups - [Scheduled Backups Guide](https://kafkabackup.com/operator/guides/scheduled-backups.md) - Scheduling strategies - [Backup Retention Guide](https://kafkabackup.com/operator/guides/backup-retention.md) - Configure opt-in retention - [Secrets Guide](https://kafkabackup.com/operator/guides/secrets.md) - Configure credentials --- title: KafkaRestore CRD description: KafkaRestore Custom Resource Definition reference source_url: html: https://kafkabackup.com/operator/crds/kafkarestore md: https://kafkabackup.com/operator/crds/kafkarestore.md --- # KafkaRestore CRD The `KafkaRestore` custom resource restores a backup into a target Kafka cluster. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaRestore metadata: name: production-restore namespace: kafka-backup spec: backupRef: name: production-backup namespace: kafka-backup backupId: "production-backup-20260413-120000" # storage can be set here instead of name/namespace for external backups. kafkaCluster: bootstrapServers: - kafka-0.kafka.svc:9092 - kafka-1.kafka.svc:9092 securityProtocol: SASL_SSL caSecret: # optional: separate CA secret (e.g. Strimzi) name: cluster-ca-cert caKey: ca.crt tlsSecret: name: kafka-tls caKey: ca.crt saslSecret: name: kafka-credentials mechanism: SCRAM-SHA-512 connection: tcpKeepalive: true keepaliveTimeSecs: 60 keepaliveIntervalSecs: 20 tcpNodelay: true connectionsPerBroker: 4 topics: - orders - payments topicMapping: orders: restored-orders partitionMapping: 0: 0 repartitioning: restored-orders: strategy: murmur2 targetPartitions: 6 pitr: startTime: "2026-04-13T10:00:00Z" endTime: "2026-04-13T12:00:00Z" offsetReset: enabled: true consumerGroups: - order-processor strategy: manual rollback: snapshotBeforeRestore: true snapshotRetentionHours: 24 autoRollbackOnFailure: false dryRun: false produceBatchSize: 1000 produceAcks: -1 produceTimeoutMs: 30000 purgeTopics: false autoConsumerGroups: false createTopics: true defaultReplicationFactor: 3 template: pod: hostAliases: - ip: "10.10.0.5" hostnames: - s3.internal - minio.internal ``` | Field | Type | Required | Description | | --- | --- | --- | --- | | `name` | string | Yes | `KafkaBackup` resource to restore from; use `""` when using direct `storage` | | `namespace` | string | No | Namespace of the `KafkaBackup`; defaults to the restore namespace | | `backupId` | string | No | Specific backup ID | | `storage` | object | No | Direct storage reference for external backups | | Field | Type | Required | Description | | --- | --- | --- | --- | | `bootstrapServers` | `[]string` | Yes | Target Kafka broker addresses | | `securityProtocol` | string | No | `PLAINTEXT`, `SSL`, `SASL_PLAINTEXT`, or `SASL_SSL` | | `tlsSecret` | object | No | TLS certificate secret reference | | `caSecret` | object | No | Separate CA certificate secret (e.g. Strimzi cluster CA). Overrides `caKey` in `tlsSecret` when both are set | | `saslSecret` | object | No | SASL credentials secret reference | | `connection` | object | No | Kafka TCP connection tuning, including `connectionsPerBroker` | | Field | Type | Required | Description | | --- | --- | --- | --- | | `startTimestamp` | int | No | Start timestamp in epoch milliseconds | | `endTimestamp` | int | No | End timestamp in epoch milliseconds | | `startTime` | string | No | Start timestamp as ISO 8601 | | `endTime` | string | No | End timestamp as ISO 8601 | | Field | Type | Default | Description | | --- | --- | --- | --- | | `topics` | `[]string` | `[]` | Topics to restore; empty restores all topics in the backup | | `topicMapping` | map | `{}` | Source topic to target topic names | | `partitionMapping` | map | `{}` | Source partition to target partition IDs | | `repartitioning` | map | `{}` | Per-target-topic repartitioning config | | `dryRun` | bool | `false` | Validate without executing the restore | | `produceBatchSize` | int | `1000` | Producer batch size | | `produceAcks` | int | `-1` | `-1` for all ISR, `1` for leader, `0` for no acknowledgements | | `produceTimeoutMs` | int | `30000` | Broker-side produce timeout | | `purgeTopics` | bool | `false` | Delete records from target topics before restore | | `autoConsumerGroups` | bool | `false` | Load consumer groups from `consumer-groups-snapshot.json` and reset offsets after restore | | `createTopics` | bool | `false` | Create missing target topics | | `defaultReplicationFactor` | int | unset | Replication factor for created topics | | `includeOriginalOffsetHeader` | bool | `true` | Add `x-original-offset`, `x-original-timestamp` and `x-source-partition` headers to every restored record (v1.3.0; this was always on before) | | `stripOffsetHeaders` | bool | `false` | Remove the headers kafka-backup added at backup time before producing; with `includeOriginalOffsetHeader: false` the restored records match the source header-for-header (v1.3.0) | ``` repartitioning: restored-orders: strategy: murmur2 # murmur2 or automatic targetPartitions: 6 ``` `targetPartitions` must be greater than zero. `strategy` must be `murmur2` or `automatic`. | Field | Type | Required | Description | | --- | --- | --- | --- | | `enabled` | bool | No | Enable offset reset as part of restore | | `consumerGroups` | `[]string` | Yes | Consumer groups to reset | | `strategy` | string | No | `manual`, `auto`, or `dry_run`; defaults to `manual` | | Field | Type | Default | Description | | --- | --- | --- | --- | | `snapshotBeforeRestore` | bool | `true` | Create an offset snapshot before restore | | `snapshotRetentionHours` | int | `24` | Snapshot retention in hours | | `autoRollbackOnFailure` | bool | `false` | Roll back offsets automatically on restore failure | | `snapshotStorage` | object | unset | Optional PVC storage override for snapshots | | Field | Type | Required | Description | | --- | --- | --- | --- | | `hostAliases` | array | No | Pod-level Kubernetes `hostAliases` entries injected into restore Job pods. Use this for split-DNS or private endpoint storage access | Each `hostAliases` item uses the Kubernetes `HostAlias` shape: | Field | Type | Required | Description | | --- | --- | --- | --- | | `ip` | string | Yes | IP address to write into the pod `/etc/hosts` file | | `hostnames` | `[]string` | No | Hostnames mapped to `ip` | ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaRestore metadata: name: restore-orders spec: backupRef: name: production-backup backupId: "production-backup-20260413-120000" kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders createTopics: true produceAcks: -1 ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaRestore metadata: name: restore-external spec: backupRef: name: "" backupId: "external-backup-20260413" storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 prefix: production credentialsSecret: name: s3-credentials kafkaCluster: bootstrapServers: - kafka:9092 ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaRestore metadata: name: restore-with-groups spec: backupRef: name: production-backup kafkaCluster: bootstrapServers: - dr-kafka:9092 autoConsumerGroups: true createTopics: true ``` Use `template.pod.hostAliases` when restore pods need the same private storage endpoint mapping as backup pods. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaRestore metadata: name: restore-private-s3 spec: backupRef: name: production-backup backupId: "production-backup-20260413-120000" kafkaCluster: bootstrapServers: - dr-kafka:9092 createTopics: true template: pod: hostAliases: - ip: "10.10.0.5" hostnames: - s3.internal - minio.internal ``` ``` status: phase: Completed message: "Restore completed successfully" startTime: "2026-04-13T12:00:00Z" completionTime: "2026-04-13T12:30:00Z" progressPercent: 100 recordsRestored: 1000000 bytesRestored: 1073741824 segmentsProcessed: 8 offsetMappingPath: "/tmp/offset-mapping.json" ``` - [KafkaBackup](https://kafkabackup.com/operator/crds/kafkabackup.md) - Create backups - [KafkaOffsetReset](https://kafkabackup.com/operator/crds/kafkaoffsetreset.md) - Reset offsets after restore --- title: KafkaOffsetReset CRD description: KafkaOffsetReset Custom Resource Definition reference source_url: html: https://kafkabackup.com/operator/crds/kafkaoffsetreset md: https://kafkabackup.com/operator/crds/kafkaoffsetreset.md --- # KafkaOffsetReset CRD The `KafkaOffsetReset` custom resource resets consumer group offsets on a Kafka cluster. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaOffsetReset metadata: name: offset-reset namespace: kafka-backup spec: kafkaCluster: bootstrapServers: - kafka-0.kafka.svc:9092 - kafka-1.kafka.svc:9092 securityProtocol: SASL_SSL caSecret: # optional: separate CA secret (e.g. Strimzi) name: cluster-ca-cert caKey: ca.crt tlsSecret: name: kafka-tls caKey: ca.crt saslSecret: name: kafka-credentials mechanism: SCRAM-SHA-512 connection: connectionsPerBroker: 4 consumerGroups: - order-processor - payment-handler topics: - orders - payments resetStrategy: to-timestamp resetTimestamp: 1713009600000 resetOffset: null parallelism: 50 dryRun: false continueOnError: false snapshotBeforeReset: true offsetMappingRef: restoreName: production-restore path: /tmp/offset-mapping.json pvcName: restore-workspace ``` | Field | Type | Required | Description | | --- | --- | --- | --- | | `bootstrapServers` | `[]string` | Yes | Kafka broker addresses | | `securityProtocol` | string | No | `PLAINTEXT`, `SSL`, `SASL_PLAINTEXT`, or `SASL_SSL` | | `tlsSecret` | object | No | TLS certificate secret reference | | `caSecret` | object | No | Separate CA certificate secret (e.g. Strimzi cluster CA). Overrides `caKey` in `tlsSecret` when both are set | | `saslSecret` | object | No | SASL credentials secret reference | | `connection` | object | No | Kafka TCP connection tuning, including `connectionsPerBroker` | | Value | Description | | --- | --- | | `to-earliest` | Reset to earliest available offset | | `to-latest` | Reset to latest offset | | `to-timestamp` | Reset to the first offset at or after `resetTimestamp` | | `to-offset` | Reset to `resetOffset` | | `from-mapping` | Reset using an offset mapping from a restore operation | | Field | Type | Default | Description | | --- | --- | --- | --- | | `consumerGroups` | `[]string` | Required | Consumer groups to reset | | `topics` | `[]string` | `[]` | Topics to reset; empty resets all topics for the group | | `resetTimestamp` | int | unset | Epoch milliseconds for `to-timestamp` | | `resetOffset` | int | unset | Target offset for `to-offset` | | `parallelism` | int | `50` | Bulk reset parallelism | | `dryRun` | bool | `false` | Validate without committing offsets | | `continueOnError` | bool | `false` | Continue if one group fails | | `snapshotBeforeReset` | bool | `true` | Create a rollback snapshot first | | `offsetMappingRef` | object | unset | Restore offset mapping reference for `from-mapping` | ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaOffsetReset metadata: name: reset-to-earliest spec: kafkaCluster: bootstrapServers: - kafka:9092 consumerGroups: - my-consumer-group resetStrategy: to-earliest ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaOffsetReset metadata: name: reset-to-timestamp spec: kafkaCluster: bootstrapServers: - kafka:9092 consumerGroups: - order-processor topics: - orders resetStrategy: to-timestamp resetTimestamp: 1713009600000 snapshotBeforeReset: true ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaOffsetReset metadata: name: reset-from-restore spec: kafkaCluster: bootstrapServers: - kafka:9092 consumerGroups: - order-processor resetStrategy: from-mapping offsetMappingRef: restoreName: production-restore ``` ``` status: phase: Completed groupsTotal: 2 groupsReset: 2 groupsFailed: 0 snapshotId: "offset-reset-20260413" snapshotPath: "/tmp/offset-reset-20260413.json" ``` --- title: KafkaOffsetRollback CRD description: KafkaOffsetRollback Custom Resource Definition reference source_url: html: https://kafkabackup.com/operator/crds/kafkaoffsetrollback md: https://kafkabackup.com/operator/crds/kafkaoffsetrollback.md --- # KafkaOffsetRollback CRD The `KafkaOffsetRollback` custom resource rolls consumer group offsets back to a saved snapshot. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaOffsetRollback metadata: name: offset-rollback namespace: kafka-backup spec: snapshotRef: name: pre-restore-snapshot pvcName: restore-workspace path: snapshots/pre-restore-snapshot.json restoreRef: production-restore offsetResetRef: reset-from-restore kafkaCluster: bootstrapServers: - kafka-0.kafka.svc:9092 - kafka-1.kafka.svc:9092 securityProtocol: SASL_SSL caSecret: # optional: separate CA secret (e.g. Strimzi) name: cluster-ca-cert caKey: ca.crt tlsSecret: name: kafka-tls saslSecret: name: kafka-credentials mechanism: SCRAM-SHA-512 connection: connectionsPerBroker: 4 consumerGroups: - order-processor - payment-handler dryRun: false verifyAfterRollback: true ``` | Field | Type | Required | Description | | --- | --- | --- | --- | | `name` | string | Yes | Snapshot name or ID | | `pvcName` | string | No | PVC containing the snapshot | | `path` | string | No | Path to the snapshot file within the PVC | | `restoreRef` | string | No | `KafkaRestore` that created the snapshot | | `offsetResetRef` | string | No | `KafkaOffsetReset` that created the snapshot | | Field | Type | Required | Description | | --- | --- | --- | --- | | `bootstrapServers` | `[]string` | Yes | Kafka broker addresses | | `securityProtocol` | string | No | `PLAINTEXT`, `SSL`, `SASL_PLAINTEXT`, or `SASL_SSL` | | `tlsSecret` | object | No | TLS certificate secret reference | | `caSecret` | object | No | Separate CA certificate secret (e.g. Strimzi cluster CA). Overrides `caKey` in `tlsSecret` when both are set | | `saslSecret` | object | No | SASL credentials secret reference | | `connection` | object | No | Kafka TCP connection tuning, including `connectionsPerBroker` | | Field | Type | Default | Description | | --- | --- | --- | --- | | `consumerGroups` | `[]string` | `[]` | Groups to roll back; empty uses all groups in the snapshot | | `dryRun` | bool | `false` | Validate without committing offsets | | `verifyAfterRollback` | bool | `true` | Verify offsets after rollback | ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaOffsetRollback metadata: name: rollback-restore spec: snapshotRef: name: pre-restore-snapshot restoreRef: production-restore kafkaCluster: bootstrapServers: - kafka:9092 consumerGroups: - order-processor verifyAfterRollback: true ``` ``` status: phase: Completed groupsRolledBack: 2 groupsFailed: 0 verification: allMatched: true totalGroups: 2 matchedGroups: 2 ``` --- title: KafkaBackupValidation CRD description: KafkaBackupValidation Custom Resource Definition reference source_url: html: https://kafkabackup.com/operator/crds/kafkabackupvalidation md: https://kafkabackup.com/operator/crds/kafkabackupvalidation.md --- # KafkaBackupValidation CRD The `KafkaBackupValidation` custom resource validates backups and can produce evidence reports for compliance workflows. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackupValidation metadata: name: daily-backup-validation namespace: kafka-backup spec: backupRef: name: production-backup backupId: "production-backup-20260413-120000" kafkaCluster: bootstrapServers: - kafka:9092 connection: connectionsPerBroker: 4 checks: messageCount: enabled: true failThreshold: 0 topics: - orders offsetRange: enabled: true consumerGroupOffsets: enabled: true consumerGroups: - order-processor evidence: formats: - json - pdf retentionDays: 90 signing: enabled: true keySecret: name: evidence-signing-key privateKeyKey: signing-key.pem publicKeyKey: signing-key-pub.pem schedule: "0 0 2 * * * *" suspend: false ``` | Field | Type | Required | Description | | --- | --- | --- | --- | | `name` | string | Yes | `KafkaBackup` resource to validate; use `""` when using direct `storage` | | `namespace` | string | No | Backup namespace; defaults to the validation namespace | | `backupId` | string | No | Specific backup ID to validate | | `storage` | object | No | Direct storage reference for external backups | | Field | Type | Description | | --- | --- | --- | | `messageCount` | object | Compare record counts against the backup manifest | | `offsetRange` | object | Verify backup offset ranges | | `consumerGroupOffsets` | object | Verify consumer group positions are backed up | | `customWebhooks` | array | POST validation payloads to external endpoints | `kafkaCluster` is required when `consumerGroupOffsets` is enabled. The shared `kafkaCluster.connection` tuning fields are supported here, along with `tlsSecret`, `caSecret`, and `saslSecret` for secure clusters. | Field | Type | Default | Description | | --- | --- | --- | --- | | `formats` | `[]string` | `["json"]` | Evidence output formats | | `signing` | object | unset | ECDSA-P256-SHA256 signing configuration | | `storage` | object | unset | Evidence storage location | | `retentionDays` | int | `90` | Evidence retention period | `schedule` uses the same cron parser as `KafkaBackup` and includes seconds. For example, `0 0 2 * * * *` runs daily at 02:00 UTC. ``` status: phase: Completed validationResult: Pass checksTotal: 3 checksCompleted: 3 checksPassed: 3 evidenceReportPath: "evidence/daily-backup-validation.json" evidenceReportSigned: true lastValidationTime: "2026-04-13T02:00:00Z" nextScheduledValidation: "2026-04-14T02:00:00Z" ``` --- title: Scheduled Backups description: Configure automated backup schedules with the Kubernetes Operator source_url: html: https://kafkabackup.com/operator/guides/scheduled-backups md: https://kafkabackup.com/operator/guides/scheduled-backups.md --- # Scheduled Backups Configure recurring backup schedules with the `KafkaBackup` CRD. The operator uses the Rust `cron` parser. Schedules include seconds: ``` ┌───────────── second (0 - 59) │ ┌───────────── minute (0 - 59) │ │ ┌───────────── hour (0 - 23) │ │ │ ┌───────────── day of month (1 - 31) │ │ │ │ ┌───────────── month (1 - 12) │ │ │ │ │ ┌───────────── day of week (0 - 6) (Sunday to Saturday) │ │ │ │ │ │ ┌───────────── year (optional) │ │ │ │ │ │ │ * * * * * * * ``` | Schedule | Cron Expression | Description | | --- | --- | --- | | Every hour | `0 0 * * * * *` | At minute 0 of every hour | | Every 15 minutes | `0 */15 * * * * *` | Every 15 minutes | | Every 6 hours | `0 0 */6 * * * *` | At minute 0 past every 6th hour | | Daily at midnight | `0 0 0 * * * *` | Every day at 00:00 | | Daily at 2 AM | `0 0 2 * * * *` | Every day at 02:00 | | Weekly on Sunday | `0 0 0 * * 0 *` | Every Sunday at 00:00 | | Monthly | `0 0 0 1 * * *` | First day of month at 00:00 | Set `stopAtCurrentOffsets: true` for scheduled point-in-time backups. The backup captures the current high watermarks when it starts and exits after all partitions reach them. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: hourly-backup namespace: kafka-backup spec: schedule: "0 0 * * * * *" stopAtCurrentOffsets: true kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders - payments storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 prefix: hourly credentialsSecret: name: s3-credentials compression: zstd compressionLevel: 3 includeOffsetHeaders: true sourceClusterId: production ``` For S3-compatible storage reached through VPN or split DNS, add `template.pod.hostAliases`. The operator copies these aliases into the CronJob pod template, so every scheduled backup pod receives the same `/etc/hosts` entries. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: private-s3-hourly namespace: kafka-backup spec: schedule: "0 0 * * * * *" stopAtCurrentOffsets: true kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders - payments storage: storageType: s3 s3: bucket: kafka-backups region: us-east-1 endpoint: https://s3.internal pathStyle: true credentialsSecret: name: s3-credentials template: pod: hostAliases: - ip: "10.10.0.5" hostnames: - s3.internal - minio.internal ``` Set `continuous: true` for streaming backups. In v1.0.0 and later, `checkpoint.enabled` does not imply continuous mode. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: continuous-backup namespace: kafka-backup spec: continuous: true kafkaCluster: bootstrapServers: - kafka:9092 topics: - "*" storage: storageType: azure azure: accountName: kafkabackups123456 container: kafka-backups prefix: continuous useWorkloadIdentity: true checkpoint: enabled: true intervalSecs: 30 consumerGroupSnapshot: true ``` Retention is disabled by default. If `spec.retention` is omitted, scheduled backups continue to create new backup IDs and the operator does not delete old backup data. Enable retention per `KafkaBackup` when you want the operator to prune complete backup sets after a successful run: ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: hourly-backup namespace: kafka-backup spec: schedule: "0 0 * * * * *" stopAtCurrentOffsets: true kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders - payments storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 prefix: hourly credentialsSecret: name: s3-credentials retention: enabled: true maxAgeDays: 30 keepLast: 3 dryRun: true ``` Start with `dryRun: true` and inspect `.status.retentionEligibleBackups`, `.status.retentionDeletedBackups`, `.status.retentionReclaimedBytes`, and `.status.retentionError`. Change `dryRun` to `false` only after the reported policy matches your restore and compliance requirements. Retention deletes whole backup IDs, not individual segments. The current backup ID is always retained, and if only `maxAgeDays` is set the operator still keeps at least the newest backup set. Use `keepLast` as an additional safety guard for scheduled backups. Operator-managed retention works with PVC/local storage, S3/S3-compatible storage, and Azure Blob Storage. For GCS, or when you need backend-native legal hold or object-lock controls, use storage-level lifecycle policies scoped to the backup prefix. For long-running `continuous: true` backups, retention does not prune while the backup engine is still running. Treat retention as primarily useful for scheduled and on-demand backup sets. > [!NOTE] > > [!NOTE] > > Available from v0.13.5 Set `checkpoint.enabled: true` with `stopAtCurrentOffsets: true` to create **incremental scheduled backups**. Each scheduled run picks up where the previous one stopped, backing up only new messages. This is the recommended pattern for most scheduled backup workloads — it combines the consistency guarantees of snapshot mode with the efficiency of incremental backups. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: incremental-hourly namespace: kafka-backup spec: schedule: "0 0 * * * * *" stopAtCurrentOffsets: true kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders - payments storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 prefix: incremental-hourly credentialsSecret: name: s3-credentials checkpoint: enabled: true intervalSecs: 30 compression: zstd compressionLevel: 3 includeOffsetHeaders: true ``` With this configuration: - **First run** backs up all existing data, saves the offset checkpoint, exits - **Subsequent runs** load the checkpoint, skip already-backed-up data, and back up only new messages - The manifest is **merged** across runs — all segments from all runs are preserved - If a run fails, the next run resumes from the last successful checkpoint > [!TIP] > > [!NOTE] > > When to use incremental vs full scheduled backups > > [!NOTE] > > Use **incremental** when your topics have high throughput and you want fast, efficient scheduled backups. Use **full** (without `checkpoint.enabled`) when you want each backup to be a standalone, self-contained snapshot — for example, for compliance evidence where each backup must independently cover a specific time window. ``` # Tier 1: frequent critical-topic snapshots --- apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: tier1-frequent spec: schedule: "0 */15 * * * * *" stopAtCurrentOffsets: true kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders - payments storage: storageType: s3 s3: bucket: kafka-backups prefix: tier1-frequent region: us-west-2 credentialsSecret: name: s3-credentials compression: lz4 --- # Tier 2: hourly all-topic snapshots apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: tier2-hourly spec: schedule: "0 0 * * * * *" stopAtCurrentOffsets: true kafkaCluster: bootstrapServers: - kafka:9092 topics: - "*" storage: storageType: s3 s3: bucket: kafka-backups prefix: tier2-hourly region: us-west-2 credentialsSecret: name: s3-credentials compression: zstd ``` For manual or on-demand backups, omit `schedule`. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: manual-backup-20260413 spec: kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders - payments storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 prefix: manual credentialsSecret: name: s3-credentials stopAtCurrentOffsets: true ``` ``` spec: suspend: true ``` ``` kubectl get kafkabackup hourly-backup \ -o jsonpath='{.status.nextScheduledBackup}' ``` ``` kubectl get kafkabackup hourly-backup \ -o jsonpath='{.status.lastBackupTime}' ``` 1. Set `stopAtCurrentOffsets: true` for recurring point-in-time backups. 2. Set `continuous: true` only for streaming backups. 3. Keep `includeOffsetHeaders: true` when you plan to restore and reset offsets. 4. Use opt-in operator retention or storage lifecycle policies for backup data retention. 5. Enable `consumerGroupSnapshot: true` when consumer group recovery matters. - [Backup Retention](https://kafkabackup.com/operator/guides/backup-retention.md) - Configure opt-in retention - [GitOps Integration](https://kafkabackup.com/operator/guides/gitops.md) - Version control your backup configs - [Secrets Guide](https://kafkabackup.com/operator/guides/secrets.md) - Configure credentials - [KafkaBackup CRD](https://kafkabackup.com/operator/crds/kafkabackup.md) - Full specification --- title: Backup Retention description: Configure opt-in backup retention with the Kubernetes Operator source_url: html: https://kafkabackup.com/operator/guides/backup-retention md: https://kafkabackup.com/operator/guides/backup-retention.md --- # Backup Retention Operator-managed retention is opt-in. The operator does not delete backup data unless `spec.retention.enabled: true` is set on a `KafkaBackup`. Retention runs after a backup completes successfully. It deletes complete backup sets only, not individual segment files inside a retained backup. This keeps restore metadata consistent and avoids partial point-in-time restore windows. | Storage backend | Operator-managed retention | | --- | --- | | PVC/local | Supported | | S3 and S3-compatible storage | Supported | | Azure Blob Storage | Supported | | GCS | Not currently supported; use a GCS lifecycle policy | Use storage lifecycle policies instead of operator-managed retention when you need backend-native object lock, legal hold, cross-account enforcement, or GCS support. ``` spec: retention: enabled: true maxAgeDays: 30 keepLast: 3 dryRun: true ``` | Field | Default | Description | | --- | --- | --- | | `enabled` | `false` | Enables operator-managed pruning | | `maxAgeDays` | unset | Deletes complete backup sets older than this many days | | `keepLast` | unset | Keeps at least this many newest backup sets | | `dryRun` | `false` | Reports eligible backup sets without deleting data | When retention is enabled, set at least one of `maxAgeDays` or `keepLast`. Values must be greater than `0`. The current backup ID is always retained. If only `maxAgeDays` is set, the operator still keeps at least the newest backup set as a safety guard. Start with `dryRun: true`: ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: production-nightly namespace: kafka-backup spec: schedule: "0 0 2 * * * *" stopAtCurrentOffsets: true kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders - payments storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 prefix: production/nightly credentialsSecret: name: s3-credentials retention: enabled: true maxAgeDays: 30 keepLast: 3 dryRun: true ``` After a backup run completes, inspect the retention status: ``` kubectl get kafkabackup production-nightly -n kafka-backup \ -o jsonpath='{.status.retentionInspectedBackups}{" inspected, "}{.status.retentionEligibleBackups}{" eligible, "}{.status.retentionDeletedBackups}{" deleted, dryRun="}{.status.retentionDryRun}{"\n"}' ``` Check for retention errors: ``` kubectl get kafkabackup production-nightly -n kafka-backup \ -o jsonpath='{.status.retentionError}{"\n"}' ``` Switch to active deletion only after the dry-run output matches the intended policy: ``` spec: retention: enabled: true maxAgeDays: 30 keepLast: 3 dryRun: false ``` | Field | Description | | --- | --- | | `lastRetentionTime` | Time of the last retention run | | `retentionInspectedBackups` | Number of backup sets considered for this `KafkaBackup` | | `retentionEligibleBackups` | Number of backup sets selected by policy | | `retentionDeletedBackups` | Number of backup sets deleted | | `retentionReclaimedBytes` | Bytes reclaimed, or bytes that would be reclaimed during dry run | | `retentionDryRun` | Whether the last retention run was a dry run | | `retentionError` | Last retention error, if pruning failed after a successful backup | Retention errors do not make a successful backup fail. Alert on `retentionError` if retention is part of your cost or compliance controls. The operator's retention (this page) deletes whole backup IDs and does not run while a `continuous: true` engine is still running. For continuous and incremental sets that keep one `backup_id`, use the engine's own segment-level retention instead — `backup.retention` in the engine config (passed through via `spec.backup.config`) or `kafka-backup prune` — which is safe to run against a live set and records what it removed in the manifest. See [Retention and Erasure](https://kafkabackup.com/guides/retention-and-erasure.md). Deleting a backup set removes that backup ID from the restoreable history. Keep `maxAgeDays` and `keepLast` aligned with recovery objectives, audit requirements, and any legal hold process. Do not enable operator-managed retention on backup prefixes that are also managed by an incompatible external cleanup process. If both the operator and the storage backend enforce retention, make sure the policies have the same restore window. --- title: GitOps Integration description: Manage Kafka backups using GitOps workflows source_url: html: https://kafkabackup.com/operator/guides/gitops md: https://kafkabackup.com/operator/guides/gitops.md --- # GitOps Integration Manage your Kafka backup configurations using GitOps principles with tools like ArgoCD, Flux, or similar. - **Version controlled** - Track all changes to backup configurations - **Auditable** - Know who changed what and when - **Reproducible** - Same configuration across environments - **Declarative** - Desired state in Git - **Automated** - Changes automatically applied ``` kafka-backup-gitops/ ├── base/ │ ├── kustomization.yaml │ ├── namespace.yaml │ └── operator/ │ └── kustomization.yaml ├── overlays/ │ ├── dev/ │ │ ├── kustomization.yaml │ │ ├── kafka-backup.yaml │ │ └── secrets.yaml │ ├── staging/ │ │ ├── kustomization.yaml │ │ ├── kafka-backup.yaml │ │ └── secrets.yaml │ └── production/ │ ├── kustomization.yaml │ ├── kafka-backup.yaml │ └── secrets.yaml └── README.md ``` base/kustomization.yaml ``` apiVersion: kustomize.config.k8s.io/v1beta1 kind: Kustomization resources: - namespace.yaml - operator/ ``` base/namespace.yaml ``` apiVersion: v1 kind: Namespace metadata: name: kafka-backup ``` overlays/production/kustomization.yaml ``` apiVersion: kustomize.config.k8s.io/v1beta1 kind: Kustomization namespace: kafka-backup resources: - ../../base - kafka-backup.yaml secretGenerator: - name: kafka-credentials files: - username=secrets/kafka-username.txt - password=secrets/kafka-password.txt configMapGenerator: - name: backup-config literals: - ENVIRONMENT=production ``` argocd/kafka-backup-app.yaml ``` apiVersion: argoproj.io/v1alpha1 kind: Application metadata: name: kafka-backup-production namespace: argocd spec: project: default source: repoURL: https://github.com/myorg/kafka-backup-gitops.git targetRevision: main path: overlays/production destination: server: https://kubernetes.default.svc namespace: kafka-backup syncPolicy: automated: prune: true selfHeal: true syncOptions: - CreateNamespace=true ``` argocd/kafka-backup-appset.yaml ``` apiVersion: argoproj.io/v1alpha1 kind: ApplicationSet metadata: name: kafka-backup namespace: argocd spec: generators: - list: elements: - environment: dev cluster: dev-cluster - environment: staging cluster: staging-cluster - environment: production cluster: production-cluster template: metadata: name: 'kafka-backup-{{environment}}' spec: project: default source: repoURL: https://github.com/myorg/kafka-backup-gitops.git targetRevision: main path: 'overlays/{{environment}}' destination: server: '{{cluster}}' namespace: kafka-backup syncPolicy: automated: prune: true selfHeal: true ``` flux/git-repo.yaml ``` apiVersion: source.toolkit.fluxcd.io/v1 kind: GitRepository metadata: name: kafka-backup namespace: flux-system spec: interval: 1m url: https://github.com/myorg/kafka-backup-gitops.git ref: branch: main secretRef: name: github-credentials ``` flux/kustomization.yaml ``` apiVersion: kustomize.toolkit.fluxcd.io/v1 kind: Kustomization metadata: name: kafka-backup-production namespace: flux-system spec: interval: 10m path: ./overlays/production prune: true sourceRef: kind: GitRepository name: kafka-backup healthChecks: - apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup name: production-backup namespace: kafka-backup ``` overlays/production/sealed-secrets.yaml ``` apiVersion: bitnami.com/v1alpha1 kind: SealedSecret metadata: name: kafka-credentials namespace: kafka-backup spec: encryptedData: username: AgBy8hCi... # Encrypted password: AgDk4kLm... # Encrypted template: metadata: name: kafka-credentials ``` Create sealed secret: ``` kubeseal --format yaml < secret.yaml > sealed-secret.yaml ``` overlays/production/external-secret.yaml ``` apiVersion: external-secrets.io/v1beta1 kind: ExternalSecret metadata: name: kafka-credentials namespace: kafka-backup spec: refreshInterval: 1h secretStoreRef: name: vault-backend kind: ClusterSecretStore target: name: kafka-credentials data: - secretKey: username remoteRef: key: kafka/production property: username - secretKey: password remoteRef: key: kafka/production property: password ``` overlays/production/secrets.enc.yaml ``` apiVersion: v1 kind: Secret metadata: name: kafka-credentials type: Opaque stringData: username: ENC[AES256_GCM,data:...,type:str] password: ENC[AES256_GCM,data:...,type:str] sops: kms: - arn: arn:aws:kms:us-west-2:123456789:key/abc123 version: 3.7.3 ``` Configure Flux for SOPS: ``` apiVersion: kustomize.toolkit.fluxcd.io/v1 kind: Kustomization metadata: name: kafka-backup-production spec: decryption: provider: sops secretRef: name: sops-gpg ``` overlays/dev/kafka-backup.yaml ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: dev-backup spec: schedule: "0 0 */4 * * * *" # Every 4 hours stopAtCurrentOffsets: true kafkaCluster: bootstrapServers: - kafka.dev.svc:9092 topics: - "*" storage: storageType: s3 s3: bucket: kafka-backups-dev region: us-west-2 prefix: dev compression: lz4 # Fast ``` overlays/production/kafka-backup.yaml ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: production-backup spec: schedule: "0 0 * * * * *" # Hourly stopAtCurrentOffsets: true kafkaCluster: bootstrapServers: - kafka-0.kafka.svc:9092 - kafka-1.kafka.svc:9092 - kafka-2.kafka.svc:9092 securityProtocol: SASL_SSL saslSecret: name: kafka-credentials mechanism: SCRAM-SHA256 tlsSecret: name: kafka-tls topics: - "*" storage: storageType: s3 s3: bucket: kafka-backups-production region: us-west-2 prefix: production/hourly compression: zstd compressionLevel: 3 includeOffsetHeaders: true sourceClusterId: "production" ``` .github/pull\_request\_template.md ``` ## Kafka Backup Change ### Type of Change - [ ] New backup configuration - [ ] Schedule change - [ ] Topic change - [ ] Storage configuration change - [ ] Other ### Description ### Checklist - [ ] Tested in dev environment - [ ] Retention policy reviewed - [ ] Storage costs estimated - [ ] Security reviewed (no plain text secrets) ``` .github/workflows/validate.yaml ``` name: Validate Kafka Backup Config on: pull_request: paths: - 'overlays/**' jobs: validate: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - name: Setup Kustomize uses: imranismail/setup-kustomize@v2 - name: Validate Kustomize run: | for dir in overlays/*/; do echo "Validating $dir" kustomize build "$dir" > /dev/null done - name: Validate CRDs run: | # Install kubeconform wget https://github.com/yannh/kubeconform/releases/latest/download/kubeconform-linux-amd64.tar.gz tar xf kubeconform-linux-amd64.tar.gz for dir in overlays/*/; do echo "Validating CRDs in $dir" kustomize build "$dir" | ./kubeconform -strict done ``` ``` # Dev -> Staging git checkout staging git merge dev git push # Staging -> Production (after validation) git checkout main git merge staging git tag release-$(date +%Y%m%d) git push --tags ``` ``` # ArgoCD argocd app rollback kafka-backup-production # Flux flux suspend kustomization kafka-backup-production git revert HEAD git push flux resume kustomization kafka-backup-production # Or simply git revert HEAD git push ``` ``` # Alert on sync failures - alert: ArgoCDSyncFailed expr: argocd_app_info{sync_status!="Synced"} == 1 for: 5m labels: severity: warning ``` ``` apiVersion: notification.toolkit.fluxcd.io/v1beta2 kind: Alert metadata: name: kafka-backup-alerts namespace: flux-system spec: providerRef: name: slack eventSeverity: error eventSources: - kind: Kustomization name: kafka-backup-production ``` 1. **Never commit plain text secrets** - Use sealed secrets, external secrets, or SOPS 2. **Use environment overlays** - Different configs for dev/staging/production 3. **Review all changes** - Require PR reviews for production changes 4. **Automated validation** - CI pipeline validates configs 5. **Gradual rollout** - Test in dev before production 6. **Document changes** - Clear commit messages and PR descriptions 7. **Tag releases** - Version your production deployments - [Secrets Guide](https://kafkabackup.com/operator/guides/secrets.md) - Secure credential management - [Scheduled Backups](https://kafkabackup.com/operator/guides/scheduled-backups.md) - Backup scheduling - [Installation](https://kafkabackup.com/operator/installation.md) - Operator setup --- title: Secrets Configuration description: Configure Kubernetes Secrets for Kafka backup credentials source_url: html: https://kafkabackup.com/operator/guides/secrets md: https://kafkabackup.com/operator/guides/secrets.md --- # Secrets Configuration Configure Kubernetes Secrets for Kafka authentication and cloud storage credentials. | Secret | Purpose | | --- | --- | | Kafka SASL credentials | Username/password for Kafka | | Kafka TLS certificates | CA cert, client cert, client key | | S3 credentials | AWS access key and secret | | Azure credentials | Connection string or service principal | | GCS credentials | Service account JSON | ``` apiVersion: v1 kind: Secret metadata: name: kafka-credentials namespace: kafka-backup type: Opaque stringData: username: backup-user password: your-secure-password ``` Reference in KafkaBackup: ``` spec: kafkaCluster: bootstrapServers: - kafka:9092 securityProtocol: SASL_SSL saslSecret: name: kafka-credentials mechanism: SCRAM-SHA256 usernameKey: username passwordKey: password ``` ``` apiVersion: v1 kind: Secret metadata: name: kafka-tls namespace: kafka-backup type: Opaque data: ca.crt: tls.crt: tls.key: ``` Create from files: ``` kubectl create secret generic kafka-tls \ --namespace kafka-backup \ --from-file=ca.crt=./ca.crt \ --from-file=tls.crt=./client.crt \ --from-file=tls.key=./client.key ``` Reference in KafkaBackup: ``` spec: kafkaCluster: tlsSecret: name: kafka-tls caKey: ca.crt certKey: tls.crt keyKey: tls.key ``` When your Kafka cluster stores the CA certificate and client certificates in separate Kubernetes secrets, use `caSecret` alongside `tlsSecret`. This is the default layout for [Strimzi](https://strimzi.io/), where the cluster CA is in its own secret and client certificates are generated per KafkaUser. ``` spec: kafkaCluster: bootstrapServers: - my-cluster-kafka-bootstrap:9093 securityProtocol: SSL caSecret: name: my-cluster-cluster-ca-cert # Strimzi cluster CA secret caKey: ca.crt tlsSecret: name: my-kafka-user # Strimzi KafkaUser secret certKey: user.crt keyKey: user.key ``` This avoids creating a combined secret manually and preserves Strimzi's automatic certificate rotation. For SSL encryption that trusts a custom CA but does not present a client certificate: ``` spec: kafkaCluster: bootstrapServers: - kafka:9093 securityProtocol: SSL caSecret: name: my-cluster-cluster-ca-cert caKey: ca.crt ``` ``` spec: kafkaCluster: bootstrapServers: - kafka:9093 securityProtocol: SASL_SSL saslSecret: name: kafka-credentials mechanism: SCRAM-SHA256 tlsSecret: name: kafka-tls ``` ``` apiVersion: v1 kind: Secret metadata: name: s3-credentials namespace: kafka-backup type: Opaque stringData: accessKey: AKIAIOSFODNN7EXAMPLE secretKey: wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY ``` Reference in KafkaBackup: ``` spec: storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 credentialsSecret: name: s3-credentials accessKeyKey: accessKey secretKeyKey: secretKey ``` No secret needed - use service account annotations: ``` apiVersion: v1 kind: ServiceAccount metadata: name: kafka-backup-operator namespace: kafka-backup annotations: eks.amazonaws.com/role-arn: arn:aws:iam::123456789:role/kafka-backup-role ``` IAM Policy: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "s3:PutObject", "s3:GetObject", "s3:ListBucket", "s3:DeleteObject" ], "Resource": [ "arn:aws:s3:::kafka-backups", "arn:aws:s3:::kafka-backups/*" ] } ] } ``` KafkaBackup without credentials (uses IRSA): ``` spec: storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 # No credentialsSecret - uses IRSA ``` ``` apiVersion: v1 kind: Secret metadata: name: azure-storage namespace: kafka-backup type: Opaque stringData: connectionString: "DefaultEndpointsProtocol=https;AccountName=mystorageaccount;AccountKey=...;EndpointSuffix=core.windows.net" ``` Reference: ``` spec: storage: storageType: azure azure: container: kafka-backups connectionStringSecret: name: azure-storage key: connectionString ``` ``` apiVersion: v1 kind: ServiceAccount metadata: name: kafka-backup-operator namespace: kafka-backup annotations: azure.workload.identity/client-id: labels: azure.workload.identity/use: "true" ``` No secret needed - uses workload identity. ``` apiVersion: v1 kind: Secret metadata: name: gcs-credentials namespace: kafka-backup type: Opaque stringData: credentials.json: | { "type": "service_account", "project_id": "my-project", "private_key_id": "...", "private_key": "-----BEGIN PRIVATE KEY-----\n...\n-----END PRIVATE KEY-----\n", "client_email": "kafka-backup@my-project.iam.gserviceaccount.com", "client_id": "...", ... } ``` Create from file: ``` kubectl create secret generic gcs-credentials \ --namespace kafka-backup \ --from-file=credentials.json=./service-account.json ``` Reference: ``` spec: storage: storageType: gcs gcs: bucket: kafka-backups credentialsSecret: name: gcs-credentials key: credentials.json ``` ``` apiVersion: v1 kind: ServiceAccount metadata: name: kafka-backup-operator namespace: kafka-backup annotations: iam.gke.io/gcp-service-account: kafka-backup@my-project.iam.gserviceaccount.com ``` No secret needed - uses workload identity. Sync secrets from external providers: ``` apiVersion: external-secrets.io/v1beta1 kind: ExternalSecret metadata: name: kafka-credentials namespace: kafka-backup spec: refreshInterval: 1h secretStoreRef: name: vault-backend kind: ClusterSecretStore target: name: kafka-credentials template: type: Opaque data: username: "{{ .username }}" password: "{{ .password }}" data: - secretKey: username remoteRef: key: kafka/production/credentials property: username - secretKey: password remoteRef: key: kafka/production/credentials property: password ``` Encrypt secrets for GitOps: ``` # Create sealed secret kubeseal --format yaml < kafka-credentials.yaml > kafka-credentials-sealed.yaml ``` ``` apiVersion: bitnami.com/v1alpha1 kind: SealedSecret metadata: name: kafka-credentials namespace: kafka-backup spec: encryptedData: username: AgBy8hCiH... password: AgDk4kLmN... ``` Inject secrets from HashiCorp Vault: ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: backup-with-vault spec: podTemplate: annotations: vault.hashicorp.com/agent-inject: "true" vault.hashicorp.com/role: "kafka-backup" vault.hashicorp.com/agent-inject-secret-kafka: "secret/data/kafka/credentials" vault.hashicorp.com/agent-inject-template-kafka: | {{- with secret "secret/data/kafka/credentials" -}} export KAFKA_USERNAME="{{ .Data.data.username }}" export KAFKA_PASSWORD="{{ .Data.data.password }}" {{- end }} ``` ``` apiVersion: external-secrets.io/v1beta1 kind: ExternalSecret spec: refreshInterval: 1h # Check for updates hourly ``` ``` # Update secret kubectl create secret generic kafka-credentials \ --namespace kafka-backup \ --from-literal=username=backup-user \ --from-literal=password=new-secure-password \ --dry-run=client -o yaml | kubectl apply -f - # Restart operator to pick up new credentials kubectl rollout restart deployment/kafka-backup-operator -n kafka-backup ``` ``` kubectl get secret kafka-credentials -n kafka-backup ``` ``` kubectl get secret kafka-credentials -n kafka-backup -o jsonpath='{.data}' | jq 'keys' ``` ``` kubectl get secret kafka-credentials -n kafka-backup \ -o jsonpath='{.data.password}' | base64 -d ``` ``` Error: Secret 'kafka-credentials' not found ``` Solution: ``` # Check secret exists in correct namespace kubectl get secrets -n kafka-backup # Create if missing kubectl create secret generic kafka-credentials ... ``` ``` Error: Key 'user' not found in secret ``` Solution: ``` # Check secret has correct keys saslSecret: name: kafka-credentials usernameKey: username # Must match actual key in secret passwordKey: password ``` ``` Error: Cannot access secret ``` Solution: ``` # Ensure operator has RBAC to read secrets apiVersion: rbac.authorization.k8s.io/v1 kind: Role rules: - apiGroups: [""] resources: ["secrets"] verbs: ["get", "list", "watch"] ``` ``` # All secrets for production backup --- apiVersion: v1 kind: Secret metadata: name: kafka-credentials namespace: kafka-backup type: Opaque stringData: username: backup-user password: ${KAFKA_PASSWORD} --- apiVersion: v1 kind: Secret metadata: name: kafka-tls namespace: kafka-backup type: Opaque data: ca.crt: ${BASE64_CA_CERT} --- # S3 using IRSA, no secret needed --- apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: production-backup namespace: kafka-backup spec: kafkaCluster: bootstrapServers: - kafka:9093 securityProtocol: SASL_SSL saslSecret: name: kafka-credentials mechanism: SCRAM-SHA256 tlsSecret: name: kafka-tls caKey: ca.crt storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 # Uses IRSA from service account ``` - [GitOps Integration](https://kafkabackup.com/operator/guides/gitops.md) - Secure secrets in Git - [Scheduled Backups](https://kafkabackup.com/operator/guides/scheduled-backups.md) - Configure backup schedules - [Security Setup](https://kafkabackup.com/guides/security-setup.md) - Complete security configuration --- title: Index source_url: html: https://kafkabackup.com/strimzi-operator/index md: https://kafkabackup.com/strimzi-operator/index.md --- --- title: Restore Jobs and Retry Behavior description: How the Strimzi Backup Operator runs Jobs — retries, backoffLimit, status conditions, and cleanup source_url: html: https://kafkabackup.com/strimzi-operator/job-behavior md: https://kafkabackup.com/strimzi-operator/job-behavior.md --- # Restore Jobs and Retry Behavior The operator executes every backup and restore as a Kubernetes Job. This page describes how those Jobs behave: how many times they may run, how to control retries with `spec.backoffLimit`, what the CR status reports, and what happens when you delete a resource. A `KafkaRestore` is **one-shot**. The operator creates a single Job for it and never creates another — whether the Job succeeds or fails. Because a restore appends to (or purges) the target topics, implicitly re-running one could duplicate data, so every run must be intentional: - After a **successful** restore the CR reports `RestoreComplete=True` and is never re-executed. - After a **failed** restore the CR reports `Ready=False` / `RestoreFailed` and the operator does not retry it. - To run a restore again, delete the `KafkaRestore` and create a new one. _Available from operator `0.2.9`. Earlier versions could re-create a completed restore Job on every 5-minute reconcile ([#29](https://github.com/osodevops/strimzi-backup-operator/issues/29))._ Within its single Job, pod-level retries are governed by the Job's `backoffLimit`, which you can set on both CRDs: ``` apiVersion: kafkabackup.com/v1 kind: KafkaRestore metadata: name: restore-orders namespace: kafka spec: strimziClusterRef: name: my-cluster backupRef: name: daily-backup backoffLimit: 2 # allow up to 2 pod retries (3 attempts total) ``` | CRD | Default | Rationale | | --- | --- | --- | | `KafkaRestore` | `0` — exactly one attempt | A retried pod re-applies a partially completed restore, which can duplicate records. Opt in deliberately if your restore is idempotent (e.g. `dryRun`). | | `KafkaBackup` | `3` | Re-running a backup is safe; transient failures (broker restarts, network blips) are retried automatically. Applies to one-shot Jobs and scheduled CronJob runs. | ``` apiVersion: kafkabackup.com/v1 kind: KafkaBackup metadata: name: daily-backup namespace: kafka spec: strimziClusterRef: name: my-cluster schedule: cron: "0 2 * * *" backoffLimit: 1 # tighten scheduled runs to a single retry storage: type: s3 s3: bucket: my-kafka-backups region: eu-west-1 ``` _`spec.backoffLimit` is available from operator `0.2.10` ([#31](https://github.com/osodevops/strimzi-backup-operator/issues/31)). Before that, all Jobs used a fixed `backoffLimit: 3`._ The operator watches its Jobs, so the CR status reflects the outcome within seconds of the Job finishing: | Condition | Meaning | | --- | --- | | `Ready=False` / `RestoreRunning` | The restore Job is running (or waiting to start) | | `Ready=True` / `RestoreCompleted` and `RestoreComplete=True` | The restore Job succeeded; `status.restore` carries start/completion times | | `Ready=False` / `RestoreFailed` and `Error=True` | The restore Job exhausted its `backoffLimit`; the operator will not retry | `KafkaBackup` reports the analogous `BackupRunning` / `BackupCompleted` / `BackupFailed` reasons, plus `status.lastBackup` and `status.backupHistory` for completed runs. ``` kubectl get kafkarestore restore-orders -n kafka \ -o jsonpath='{range .status.conditions[*]}{.type}={.status} ({.reason}){"\n"}{end}' ``` Deleting a `KafkaBackup` or `KafkaRestore` removes everything the operator created for it — Jobs, scheduled CronJobs, generated ConfigMaps, **and the Jobs' pods**. Deletes are issued with Background propagation so the garbage collector removes dependents; completed pods are not left behind. _Available from operator `0.2.10`. Earlier versions left Completed pods orphaned after CR deletion ([#30](https://github.com/osodevops/strimzi-backup-operator/issues/30))._ ``` kubectl delete kafkarestore restore-orders -n kafka kubectl get jobs,pods -n kafka -l kafkabackup.com/restore=restore-orders # No resources found — Jobs and pods are garbage collected together ``` Add Strimzi's standard annotation to temporarily stop reconciliation for a `KafkaBackup` or `KafkaRestore`: ``` kubectl annotate kafkabackup daily-backup -n kafka \ strimzi.io/pause-reconciliation="true" ``` The operator reports `ReconciliationPaused=True`, but does not add a finalizer, resolve Strimzi dependencies, or create/update ConfigMaps, Jobs, or CronJobs. Deletion cleanup still runs for resources that already have the operator finalizer. Resume by removing the annotation (or setting it to `"false"`): ``` kubectl annotate kafkabackup daily-backup -n kafka \ strimzi.io/pause-reconciliation- ``` _Available from operator `0.2.15` ([#44](https://github.com/osodevops/strimzi-backup-operator/issues/44))._ --- title: Prometheus Metrics description: Scrape Strimzi Backup Operator controller and backup/restore Job metrics source_url: html: https://kafkabackup.com/strimzi-operator/metrics md: https://kafkabackup.com/strimzi-operator/metrics.md --- # Prometheus Metrics Operator `v0.3.1` exposes two independent classes of metrics. They must be scraped separately: | Target | Default endpoint | Discovery | | --- | --- | --- | | Operator controller | `:9090/metrics` | ServiceMonitor or pod annotations | | Backup/restore Job pod | `:8080/metrics` | PodMonitor | The operator does not proxy Job metrics. Enabling only `serviceMonitor` does not collect backup progress. The ServiceMonitor and PodMonitor require the Prometheus Operator CRDs. values.yaml ``` metrics: enabled: true serviceMonitor: enabled: true interval: 30s scrapeTimeout: 10s jobPodMonitor: enabled: true interval: 30s scrapeTimeout: 10s ``` ``` helm upgrade --install strimzi-backup-operator \ oso-devops/strimzi-backup-operator \ --namespace kafka \ --create-namespace \ --version 0.3.1 \ --values values.yaml ``` The Job PodMonitor selects pods labelled `app.kubernetes.io/managed-by=kafka-backup-operator` and `kafkabackup.com/metrics=enabled`. Its default `namespaceSelector.any: true` is intentional: custom resources can live outside the Helm release namespace. If you set a custom `spec.metrics.path`, also set `metrics.jobPodMonitor.path` to the same value. Metrics are enabled by default for Job pods. For one-shot operations, keep the endpoint alive long enough for Prometheus to scrape the final values. A useful minimum is twice the scrape interval: ``` spec: metrics: enabled: true port: 8080 bindAddress: "0.0.0.0" path: /metrics updateIntervalMs: 500 keepAliveSeconds: 60 maxPartitionLabels: 100 ``` `maxPartitionLabels` limits unique `topic`/`partition`/`backup_id` label sets, not metric updates. Set it to `0` only if unlimited per-partition cardinality is acceptable. Aggregate lag and snapshot progress remain available when the per-partition limit is reached. The controller endpoint currently publishes these populated metric families: | Metric | Labels | Description | | --- | --- | --- | | `strimzi_backup_operator_build_info` | `version` | Running operator build | | `strimzi_backup_operator_reconciliations_total` | `controller`, `result` | Reconciliations by controller and result | | `strimzi_backup_operator_reconciliation_duration_seconds` | `controller`, `result` | Reconciliation duration histogram | | `strimzi_backup_operator_engine_image_info` | `image`, `source` | `1` for the default `kafka-backup` image Jobs run when `spec.image` is unset; `source` is `compiled-in` or `env` (Helm `backupJobs.image`). Operator v0.2.25+ | | `strimzi_backup_operator_engine_version_unsupported_total` | `controller` | Reconciliations that resolved an engine image older than the minimum supported `kafka-backup` release (see [Compatibility](https://kafkabackup.com/strimzi-operator.md#compatibility)). Operator v0.2.25+ | Check it without Prometheus: ``` kubectl port-forward -n kafka \ svc/strimzi-backup-operator-metrics 9090:9090 curl --fail http://127.0.0.1:9090/metrics ``` The Job endpoint uses the `kafka-backup` runtime metrics. Useful v0.19.0 families include: | Metric | Description | | --- | --- | | `kafka_backup_lag_records` | Current lag for an admitted topic/partition series | | `kafka_backup_lag_records_sum` | Current lag summed across all partitions | | `kafka_backup_snapshot_records_target` | Records in the captured finite-backup offset range | | `kafka_backup_snapshot_records_remaining` | Records left before the finite backup completes | | `kafka_backup_records_total` | Records backed up | | `kafka_backup_bytes_total` | Uncompressed bytes backed up | | `kafka_backup_errors_total` | Errors by backup and category | | `kafka_restore_progress_percent` | Restore completion percentage | Example progress query: ``` 100 * ( 1 - kafka_backup_snapshot_records_remaining / clamp_min(kafka_backup_snapshot_records_target, 1) ) ``` Use CR status for durable success/failure history. A completed one-shot Job's in-process metrics disappear when its pod exits, even when Prometheus scraped the final sample successfully. --- title: Index source_url: html: https://kafkabackup.com/enterprise/index md: https://kafkabackup.com/enterprise/index.md --- --- title: Enterprise Installation description: Install OSO Kafka Backup Enterprise via Homebrew, Docker, or binary download source_url: html: https://kafkabackup.com/enterprise/installation md: https://kafkabackup.com/enterprise/installation.md --- # Enterprise Installation OSO Kafka Backup Enterprise is distributed as a separate binary that is a **drop-in replacement** for the open-source version. It contains all OSS functionality plus enterprise features gated by a license file. Without a license, it operates identically to the OSS version. > [!TIP] > > [!NOTE] > > 14-day free trial > > [!NOTE] > > The enterprise binary includes a **14-day free trial** of all enterprise features on first run — no signup required. After the trial, [purchase a license](https://enterprise.kafkabackup.com) to continue using enterprise features, or keep using it as your OSS binary. | | Open Source | Enterprise | | --- | --- | --- | | **Homebrew** | `brew install osodevops/tap/kafka-backup` | `brew install osodevops/tap/kafka-backup-enterprise` | | **Docker** | `ghcr.io/osodevops/kafka-backup` | `osodevops/kafka-backup-enterprise` | | **Binary name** | `kafka-backup` | `kafka-backup` | | **License required** | No | Yes (trial included) | | **Source** | [GitHub (public)](https://github.com/osodevops/kafka-backup) | [GitHub (private)](https://github.com/osodevops/kafka-backup-enterprise) | Both editions produce the same `kafka-backup` binary. You only need one installed at a time. * * * The enterprise formula is published to the same Homebrew tap as the OSS version: ``` # Install enterprise edition brew install osodevops/tap/kafka-backup-enterprise # Verify kafka-backup --version kafka-backup license info ``` > [!NOTE] > > [!NOTE] > > Switching from OSS > > [!NOTE] > > If you already have the OSS version installed via Homebrew, uninstall it first to avoid conflicts: > > > > ``` > > brew uninstall kafka-backup > > brew install osodevops/tap/kafka-backup-enterprise > > ``` > > > > Both formulas install a binary named `kafka-backup`, so only one can be active at a time. The enterprise Docker image is published to Docker Hub: ``` # Pull the latest enterprise image docker pull osodevops/kafka-backup-enterprise:latest # Verify docker run --rm osodevops/kafka-backup-enterprise:latest --version docker run --rm osodevops/kafka-backup-enterprise:latest license info ``` | Tag | Description | | --- | --- | | `latest` | Most recent stable release | | `0.3.2` | Specific version (semver) | | `0.3` | Latest patch within a minor version | Mount your license file into the container: ``` docker run --rm \ -v /path/to/license.key:/etc/kafka-backup/license.key:ro \ -v $(pwd)/backup.yaml:/config/backup.yaml:ro \ osodevops/kafka-backup-enterprise:latest \ backup --config /config/backup.yaml ``` Or pass the license as an environment variable: ``` docker run --rm \ -e ENTERPRISE_LICENSE_KEY=$(base64 < license.key) \ -v $(pwd)/backup.yaml:/config/backup.yaml:ro \ osodevops/kafka-backup-enterprise:latest \ backup --config /config/backup.yaml ``` docker-compose.yml ``` services: kafka-backup: image: osodevops/kafka-backup-enterprise:latest volumes: - ./config:/config:ro - ./backups:/data/backups - ./license.key:/etc/kafka-backup/license.key:ro command: ["backup", "--config", "/config/backup.yaml"] ``` Download pre-built binaries from [GitHub Releases](https://github.com/osodevops/kafka-backup-enterprise/releases). ``` # Download latest release curl -L -o kafka-backup.tar.gz \ https://github.com/osodevops/kafka-backup-enterprise/releases/latest/download/kafka-backup-x86_64-linux.tar.gz # Verify checksum curl -L -o kafka-backup.tar.gz.sha256 \ https://github.com/osodevops/kafka-backup-enterprise/releases/latest/download/kafka-backup-x86_64-linux.tar.gz.sha256 sha256sum -c kafka-backup.tar.gz.sha256 # Extract and install tar xzf kafka-backup.tar.gz sudo mv kafka-backup /usr/local/bin/ # Verify kafka-backup --version kafka-backup license info ``` ``` curl -L -o kafka-backup.tar.gz \ https://github.com/osodevops/kafka-backup-enterprise/releases/latest/download/kafka-backup-aarch64-linux.tar.gz tar xzf kafka-backup.tar.gz sudo mv kafka-backup /usr/local/bin/ ``` ``` curl -L -o kafka-backup.tar.gz \ https://github.com/osodevops/kafka-backup-enterprise/releases/latest/download/kafka-backup-x86_64-macos.tar.gz tar xzf kafka-backup.tar.gz sudo mv kafka-backup /usr/local/bin/ ``` ``` curl -L -o kafka-backup.tar.gz \ https://github.com/osodevops/kafka-backup-enterprise/releases/latest/download/kafka-backup-aarch64-macos.tar.gz tar xzf kafka-backup.tar.gz sudo mv kafka-backup /usr/local/bin/ ``` Download `kafka-backup-x86_64-windows.zip` from the [releases page](https://github.com/osodevops/kafka-backup-enterprise/releases) and extract `kafka-backup.exe` to a directory in your `PATH`. The recommended way to deploy on Kubernetes is the official **Helm chart**, which creates a CronJob for scheduled backups with built-in license injection, credential management, and Prometheus metrics. ``` helm install kafka-backup \ oci://ghcr.io/osodevops/charts/kafka-backup-enterprise \ --namespace kafka-backup \ --values values.yaml ``` See the **[Helm Chart guide](https://kafkabackup.com/enterprise/helm-chart.md)** for complete installation instructions, configuration examples, cloud IAM setup, and troubleshooting. If you prefer raw manifests, swap the image reference and mount the license as a Secret: license-secret.yaml ``` apiVersion: v1 kind: Secret metadata: name: kafka-backup-license namespace: kafka-backup type: Opaque stringData: license.key: | -----BEGIN KAFKA-BACKUP LICENSE----- eyJsaWNlbnNlX2lkIjoiYWJjZGVm... -----END KAFKA-BACKUP LICENSE----- ``` ``` spec: containers: - name: kafka-backup image: osodevops/kafka-backup-enterprise:latest volumeMounts: - name: license mountPath: /etc/kafka-backup readOnly: true volumes: - name: license secret: secretName: kafka-backup-license ``` For full Kubernetes deployment examples, see the [Kubernetes Deployment](https://kafkabackup.com/deployment/kubernetes.md) guide. After installing the enterprise binary, verify it is working: ``` # Check version kafka-backup --version # Check license status kafka-backup license info ``` On first run without a license, the output will show the auto-trial: ``` License: 14-day trial (auto-activated) Expires: 2026-04-24 (14 days remaining) Features: schema_registry, rbac, encryption, masking, audit, wasm Status: Valid ``` The enterprise binary is a drop-in replacement. To upgrade: 1. **Replace the binary** — install the enterprise edition using any method above 2. **Your existing config works as-is** — all OSS configuration is fully compatible 3. **Add enterprise features** — optionally add an `enterprise:` section to your config YAML 4. **Apply a license** — when you're ready, [obtain and apply a license](https://kafkabackup.com/enterprise/licensing.md) No data migration is needed. Your existing backups remain fully compatible. After your 14-day trial, purchase a license to continue using enterprise features: - **Buy online:** [enterprise.kafkabackup.com](https://enterprise.kafkabackup.com) - **Contact sales:** [sales@oso.sh](mailto:sales@oso.sh) See [Licensing](https://kafkabackup.com/enterprise/licensing.md) for details on license tiers, activation, and management. 1. [Apply a license](https://kafkabackup.com/enterprise/licensing.md) (or use the 14-day trial) 2. [Configure Schema Registry backup](https://kafkabackup.com/enterprise/schema-registry.md) 3. [Configure Confluent RBAC backup](https://kafkabackup.com/enterprise/confluent-rbac.md) 4. [Review the full enterprise feature set](https://kafkabackup.com/enterprise.md) --- title: Helm Chart description: Deploy OSO Kafka Backup Enterprise on Kubernetes using the official Helm chart source_url: html: https://kafkabackup.com/enterprise/helm-chart md: https://kafkabackup.com/enterprise/helm-chart.md --- # Helm Chart Deploy OSO Kafka Backup Enterprise on Kubernetes as a scheduled CronJob using the official Helm chart. The chart packages everything needed for production Kafka backups — CronJob scheduling, license injection, credential management, metrics, and cloud IAM integration. > [!TIP] > > [!NOTE] > > 14-day free trial > > [!NOTE] > > The enterprise image includes a **14-day free trial** of all enterprise features — no signup, no license file needed. Install the chart, point it at your Kafka cluster, and it just works. - Kubernetes 1.26+ - Helm 3.8+ (OCI registry support) - kubectl configured to access your cluster - Access to your Kafka cluster from within the Kubernetes cluster - A storage backend (S3, Azure Blob, GCS, or a PersistentVolume) The chart is published as an OCI artifact to GitHub Container Registry. ``` # Create namespace kubectl create namespace kafka-backup # Install with default values helm install kafka-backup \ oci://ghcr.io/osodevops/charts/kafka-backup-enterprise \ --namespace kafka-backup ``` ``` # Check CronJob was created kubectl get cronjob -n kafka-backup # Expected output: # NAME SCHEDULE SUSPEND ACTIVE LAST SCHEDULE # kafka-backup-enterprise 0 2 * * * False 0 # Trigger a manual backup to test kubectl create job --from=cronjob/kafka-backup-kafka-backup-enterprise \ kafka-backup-manual-test -n kafka-backup # Watch the pod kubectl get pods -n kafka-backup -w # Check logs kubectl logs -n kafka-backup -l job-name=kafka-backup-manual-test -f ``` The chart uses a two-layer configuration approach: 1. **`values.yaml`** — controls Kubernetes resources (scheduling, secrets, resources, monitoring) 2. **`config.backupConfig`** — the raw `kafka-backup` YAML config, rendered into a ConfigMap Credentials in the backup config use `${ENV_VAR}` placeholders. The binary expands them at runtime from environment variables injected by Kubernetes Secrets. values.yaml ``` cronjob: schedule: "0 2 * * *" # Daily at 2 AM config: backupConfig: | mode: backup backup_id: "daily-backup" source: bootstrap_servers: - kafka-broker-0.kafka:9092 - kafka-broker-1.kafka:9092 storage: backend: s3 bucket: my-kafka-backups region: eu-west-1 serviceAccount: annotations: eks.amazonaws.com/role-arn: arn:aws:iam::123456789:role/kafka-backup ``` ``` helm install kafka-backup \ oci://ghcr.io/osodevops/charts/kafka-backup-enterprise \ --namespace kafka-backup \ --values values.yaml ``` This example enables Schema Registry backup, Confluent RBAC backup, and Prometheus metrics: values.yaml ``` cronjob: schedule: "0 */6 * * *" # Every 6 hours concurrencyPolicy: Forbid activeDeadlineSeconds: 3600 image: tag: "0.3.2" resources: requests: cpu: 100m memory: 256Mi limits: cpu: "1" memory: 1Gi license: existingSecret: kafka-backup-license credentials: existingSecret: kafka-backup-credentials config: backupConfig: | mode: backup backup_id: "enterprise-backup" source: bootstrap_servers: - kafka:9092 storage: backend: s3 bucket: kafka-backups region: us-east-1 enterprise: schema_registry: url: "https://schema-registry:8081" auth: type: basic username: ${SR_USER} password: ${SR_PASS} backup: subjects: ["*"] exclude: ["_*"] confluent_rbac: mds_url: "https://mds:8090" auth: username: ${MDS_USER} password: ${MDS_PASS} backup: principals: ["*"] metrics: enabled: true serviceMonitor: enabled: true labels: release: prometheus ``` > [!NOTE] > > [!NOTE] > > Credentials as environment variables > > [!NOTE] > > Notice how the backup config uses `${SR_USER}`, `${SR_PASS}`, etc. These are **not** Helm template variables — they are literal `${ENV_VAR}` placeholders that the `kafka-backup` binary expands at runtime. The actual values come from the Kubernetes Secret referenced by `credentials.existingSecret`. * * * Enterprise features are gated by an Ed25519-signed license file, validated entirely offline — no license server, no network calls. Without a license, the binary operates in one of three modes: | Mode | When | Enterprise Features | | --- | --- | --- | | **Auto-trial** | First 14 days, no signup needed | All features enabled | | **OSS mode** | After trial expires, no license | Disabled (warns and skips) | | **Licensed** | Valid `.lic` or `.key` file provided | Per-license features enabled | > [!NOTE] > > [!NOTE] > > Graceful degradation > > [!NOTE] > > Enterprise features **never block your backup**. If Schema Registry backup is configured but not licensed, the binary logs a warning, skips that feature, and completes the Kafka data backup normally with exit code 0. On startup, the `kafka-backup` binary checks these sources in order: | Priority | Source | How the Helm chart uses it | | --- | --- | --- | | 1 | `ENTERPRISE_LICENSE_KEY` env var | **Default** — chart injects this from a Kubernetes Secret | | 2 | `ENTERPRISE_LICENSE_FILE` env var | Can be set via `extraEnv` if mounting a file | | 3 | `/etc/kafka-backup/license.key` or `.lic` | Mount via `extraVolumes` | | 4 | `~/.config/kafka-backup/license.lic` | Not typical in containers | | 5 | Auto-trial (14 days) | Fallback when no license is found | The Helm chart uses **priority 1** by default — it sets the `ENTERPRISE_LICENSE_KEY` environment variable from a Kubernetes Secret. The value must be the **base64-encoded content** of your license file (`.lic` or `.key`). The binary accepts two formats, auto-detected by the PEM header: **Keygen.sh `.lic` format** (standard — issued via [enterprise.kafkabackup.com](https://enterprise.kafkabackup.com)): ``` -----BEGIN LICENSE FILE----- eyJlbmMiOiJleUowZVhBaU9pSktWMVFp... -----END LICENSE FILE----- ``` **Custom PEM `.key` format** (purchased via [enterprise.kafkabackup.com](https://enterprise.kafkabackup.com)): ``` -----BEGIN KAFKA-BACKUP LICENSE----- eyJwYXlsb2FkIjp7ImxpY2Vuc2VfaWQi... -----END KAFKA-BACKUP LICENSE----- ``` Both formats are verified using Ed25519 signatures with the public key embedded in the binary at compile time. The binary auto-detects the format from the PEM header. See [Licensing](https://kafkabackup.com/enterprise/licensing.md) for full details on the license payload, feature flags, and security model. Create the Secret, then reference it in your values. This works with both `.lic` and `.key` files: ``` # Create the license secret (works with either license.lic or license.key) kubectl create secret generic kafka-backup-license \ --from-literal=license-b64="$(base64 < license.key)" \ --namespace kafka-backup ``` values.yaml ``` license: existingSecret: kafka-backup-license existingSecretKey: license-b64 # default key name ``` The chart sets `ENTERPRISE_LICENSE_KEY` on the pod from this Secret. At startup, the binary base64-decodes the value, detects the PEM format, verifies the Ed25519 signature, and extracts the licensed features. > [!TIP] > > [!NOTE] > > External Secrets Operator > > [!NOTE] > > For production, consider using [External Secrets Operator](https://external-secrets.io/) or [Sealed Secrets](https://sealed-secrets.netlify.app/) to manage the license Secret. The chart works with any Secret — it just references the name: > > > > ``` > > license: > > existingSecret: my-externalsecret-license > > ``` For quick testing, pass the base64-encoded license directly: ``` helm install kafka-backup \ oci://ghcr.io/osodevops/charts/kafka-backup-enterprise \ --namespace kafka-backup \ --set license.key="$(base64 < license.key)" ``` > [!WARNING] > > [!NOTE] > > Production deployments > > [!NOTE] > > Never commit license keys to version control. Use `existingSecret` with a Secret created out-of-band. If you prefer file-based license discovery (priority 3), mount the license file directly: values.yaml ``` extraVolumes: - name: license secret: secretName: kafka-backup-license-file extraVolumeMounts: - name: license mountPath: /etc/kafka-backup readOnly: true ``` ``` # Create the secret containing the raw license file kubectl create secret generic kafka-backup-license-file \ --from-file=license.key=/path/to/your/license.key \ --namespace kafka-backup ``` With this approach, don't set `license.existingSecret` or `license.key` — the binary will find the file at `/etc/kafka-backup/license.key` automatically. When no license is configured, the binary activates a **14-day auto-trial** with all enterprise features enabled. The trial state is tracked in a file at `/var/lib/kafka-backup/.trial`. In a CronJob context, each pod starts fresh with an empty filesystem. This means: - **Without `trialPersistence`**: each pod creates a new trial state file, so the trial effectively never expires across CronJob runs - **With `trialPersistence.enabled: true`**: a PVC persists the trial state across runs, so the 14-day countdown is accurate For production, always use a real license. To disable the auto-trial entirely: values.yaml ``` extraEnv: - name: KAFKA_BACKUP_NO_TRIAL value: "1" ``` Run the built-in Helm test to verify the license was injected correctly: ``` helm test kafka-backup -n kafka-backup ``` This creates a pod that runs `kafka-backup license info` and reports the license status. ``` # Exec into a running backup pod (during a CronJob run) kubectl exec -it -n kafka-backup -- kafka-backup license info ``` Expected output when licensed: ``` License ID: abcdef12-3456-7890-abcd-ef1234567890 Customer: Acme Corp (acme@example.com) Tier: Enterprise Features: encryption, schema_registry, rbac, audit, support Expires: 2027-01-15 (285 days remaining) Status: Valid ``` The `kafka-backup` binary provides three license management commands: | Command | Description | | --- | --- | | `kafka-backup license info` | Show current license status (what the binary is using) | | `kafka-backup license verify --file license.lic` | Verify a license file without applying it | | `kafka-backup license apply --file license.lic` | Validate and save to `~/.config/kafka-backup/` | In Kubernetes, `license info` is the most useful — it shows exactly what license the pod detected at startup. The `apply` command saves to the user config directory, which isn't persistent in containers — use Secrets instead. See [Licensing](https://kafkabackup.com/enterprise/licensing.md) for the full licensing guide including obtaining licenses, license tiers, offline validation, and FAQ. * * * Service credentials (Kafka SASL, Schema Registry auth, MDS auth, S3 keys) are injected as environment variables. The backup config YAML uses `${ENV_VAR}` placeholders, and the binary expands them at runtime. Create a Secret with all credential environment variables: ``` kubectl create secret generic kafka-backup-credentials \ --from-literal=SR_USER=sr-admin \ --from-literal=SR_PASS=secret123 \ --from-literal=MDS_USER=mds-admin \ --from-literal=MDS_PASS=secret456 \ --from-literal=AWS_ACCESS_KEY_ID=AKIA... \ --from-literal=AWS_SECRET_ACCESS_KEY=wJal... \ --namespace kafka-backup ``` values.yaml ``` credentials: existingSecret: kafka-backup-credentials ``` All keys in the Secret are injected as environment variables into the pod via `envFrom`. values.yaml ``` credentials: inline: SR_USER: sr-admin SR_PASS: secret123 MDS_USER: mds-admin MDS_PASS: secret456 ``` > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Inline credentials are stored in a Kubernetes Secret created by the chart, but the values are visible in your `values.yaml` file. Use `existingSecret` for production. For S3, Azure Blob, or GCS storage, use your cloud provider's workload identity instead of static credentials. See [Cloud Provider Setup](#cloud-provider-setup) below. * * * The chart creates a Kubernetes CronJob that runs `kafka-backup backup` on a schedule: values.yaml ``` cronjob: enabled: true schedule: "0 2 * * *" # Daily at 2 AM UTC timeZone: "Europe/London" # Kubernetes 1.27+ concurrencyPolicy: Forbid # Never overlap backoffLimit: 2 # Retry twice on failure activeDeadlineSeconds: 7200 # Kill after 2 hours ttlSecondsAfterFinished: 86400 # Clean up after 24 hours ``` | Value | Default | Description | | --- | --- | --- | | `cronjob.schedule` | `0 2 * * *` | Cron schedule expression | | `cronjob.timeZone` | `""` | IANA timezone (requires K8s 1.27+) | | `cronjob.concurrencyPolicy` | `Forbid` | Prevent overlapping backup runs | | `cronjob.backoffLimit` | `2` | Number of retries on failure | | `cronjob.activeDeadlineSeconds` | `7200` | Maximum runtime before kill | | `cronjob.suspend` | `false` | Pause without deleting | ``` kubectl create job --from=cronjob/kafka-backup-kafka-backup-enterprise \ kafka-backup-manual-$(date +%s) -n kafka-backup ``` Pause scheduled backups without deleting the CronJob: ``` kubectl patch cronjob kafka-backup-kafka-backup-enterprise \ -n kafka-backup \ -p '{"spec":{"suspend":true}}' ``` Or via Helm: ``` helm upgrade kafka-backup \ oci://ghcr.io/osodevops/charts/kafka-backup-enterprise \ --namespace kafka-backup \ --set cronjob.suspend=true ``` * * * For one-shot operations like restores, enable the Job workload instead of the CronJob: ``` helm install kafka-restore \ oci://ghcr.io/osodevops/charts/kafka-backup-enterprise \ --namespace kafka-backup \ --set cronjob.enabled=false \ --set job.enabled=true \ --set job.command=restore \ --values restore-values.yaml ``` restore-values.yaml ``` cronjob: enabled: false job: enabled: true command: restore backoffLimit: 0 activeDeadlineSeconds: 7200 config: backupConfig: | mode: restore backup_id: "daily-backup" target: bootstrap_servers: - kafka:9092 storage: backend: s3 bucket: kafka-backups region: us-east-1 ``` ``` # Watch the restore kubectl logs -n kafka-backup -l job-name=kafka-restore-kafka-backup-enterprise-job -f ``` ``` helm install schema-backup \ oci://ghcr.io/osodevops/charts/kafka-backup-enterprise \ --namespace kafka-backup \ --set cronjob.enabled=false \ --set job.enabled=true \ --set "job.extraArgs={--schema-only}" \ --values values.yaml ``` > [!TIP] > > [!NOTE] > > Clean up after one-shot jobs > > [!NOTE] > > Jobs auto-delete after `ttlSecondsAfterFinished` (default: 24 hours). To clean up immediately: > > > > ``` > > helm uninstall kafka-restore -n kafka-backup > > ``` * * * Use workload identity to grant the backup pod access to cloud storage without static credentials. ``` # Create IAM policy for S3 access cat > /tmp/kafka-backup-policy.json <<'EOF' { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "s3:PutObject", "s3:GetObject", "s3:ListBucket", "s3:DeleteObject" ], "Resource": [ "arn:aws:s3:::kafka-backups", "arn:aws:s3:::kafka-backups/*" ] } ] } EOF aws iam create-policy \ --policy-name KafkaBackupPolicy \ --policy-document file:///tmp/kafka-backup-policy.json # Create IAM service account with IRSA eksctl create iamserviceaccount \ --name kafka-backup-enterprise \ --namespace kafka-backup \ --cluster my-cluster \ --attach-policy-arn arn:aws:iam::123456789:policy/KafkaBackupPolicy \ --approve ``` values.yaml ``` serviceAccount: create: false name: kafka-backup-enterprise # Service account created by eksctl with IRSA annotation ``` ``` # Create managed identity az identity create \ --name kafka-backup-identity \ --resource-group myResourceGroup CLIENT_ID=$(az identity show \ --name kafka-backup-identity \ --resource-group myResourceGroup \ --query clientId -o tsv) # Grant Storage Blob Data Contributor role az role assignment create \ --assignee $CLIENT_ID \ --role "Storage Blob Data Contributor" \ --scope /subscriptions//resourceGroups//providers/Microsoft.Storage/storageAccounts/ # Create federated credential az identity federated-credential create \ --name kafka-backup-federated \ --identity-name kafka-backup-identity \ --resource-group myResourceGroup \ --issuer $(az aks show --name myAKSCluster --resource-group myResourceGroup --query oidcIssuerProfile.issuerUrl -o tsv) \ --subject system:serviceaccount:kafka-backup:kafka-backup-enterprise \ --audience api://AzureADTokenExchange ``` values.yaml ``` serviceAccount: create: true annotations: azure.workload.identity/client-id: "" podLabels: azure.workload.identity/use: "true" ``` ``` # Create GCP service account gcloud iam service-accounts create kafka-backup-sa # Grant Storage Object Admin gcloud projects add-iam-policy-binding PROJECT_ID \ --member "serviceAccount:kafka-backup-sa@PROJECT_ID.iam.gserviceaccount.com" \ --role "roles/storage.objectAdmin" # Bind to Kubernetes service account gcloud iam service-accounts add-iam-policy-binding \ kafka-backup-sa@PROJECT_ID.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:PROJECT_ID.svc.id.goog[kafka-backup/kafka-backup-enterprise]" ``` values.yaml ``` serviceAccount: create: true annotations: iam.gke.io/gcp-service-account: kafka-backup-sa@PROJECT_ID.iam.gserviceaccount.com ``` * * * When `metrics.enabled` is true, the chart adds Prometheus scrape annotations to the CronJob pods: values.yaml ``` metrics: enabled: true config: backupConfig: | mode: backup # ... your config ... metrics: enabled: true port: 8080 path: /metrics ``` > [!NOTE] > > [!NOTE] > > note > > [!NOTE] > > Prometheus pod annotations require your Prometheus instance to be configured for pod service discovery. CronJob pods are short-lived — metrics are only scrapeable while a backup is running. For clusters using the Prometheus Operator: values.yaml ``` metrics: enabled: true serviceMonitor: enabled: true interval: 15s labels: release: prometheus # Must match your Prometheus selector ``` The `kafka-backup` binary exposes Prometheus metrics during execution: | Metric | Type | Description | | --- | --- | --- | | `kafka_backup_records_processed` | Counter | Total records backed up | | `kafka_backup_bytes_processed` | Counter | Total bytes backed up | | `kafka_backup_segments_completed` | Counter | Completed segment files | | `kafka_backup_duration_seconds` | Histogram | Backup duration | | `kafka_backup_errors_total` | Counter | Errors during backup | * * * The chart enforces a hardened security context by default: ``` # Defaults — no changes needed podSecurityContext: runAsNonRoot: true runAsUser: 1000 runAsGroup: 1000 fsGroup: 1000 securityContext: allowPrivilegeEscalation: false readOnlyRootFilesystem: true capabilities: drop: [ALL] ``` The container runs as non-root (UID 1000) with a read-only filesystem. Writable paths: - `/tmp` — emptyDir for temporary files - `/var/lib/kafka-backup` — optional PVC for trial state persistence To connect to a Kafka cluster using mTLS, mount your certificates via `extraVolumes`: values.yaml ``` extraVolumes: - name: kafka-certs secret: secretName: kafka-client-certs extraVolumeMounts: - name: kafka-certs mountPath: /certs/kafka readOnly: true config: backupConfig: | mode: backup source: bootstrap_servers: ["kafka:9093"] security_protocol: SSL ssl: ca_location: /certs/kafka/ca.crt certificate_location: /certs/kafka/client.crt key_location: /certs/kafka/client.key storage: backend: s3 bucket: kafka-backups ``` * * * ``` helm upgrade kafka-backup \ oci://ghcr.io/osodevops/charts/kafka-backup-enterprise \ --namespace kafka-backup \ --values values.yaml ``` To upgrade to a specific version: ``` helm upgrade kafka-backup \ oci://ghcr.io/osodevops/charts/kafka-backup-enterprise \ --namespace kafka-backup \ --version 0.2.0 ``` ``` helm uninstall kafka-backup --namespace kafka-backup ``` > [!NOTE] > > [!NOTE] > > PVC retention > > [!NOTE] > > If `trialPersistence.enabled` was set, the PVC is **not** deleted on uninstall (Helm default). Delete it manually if no longer needed: > > > > ``` > > kubectl delete pvc kafka-backup-kafka-backup-enterprise-trial -n kafka-backup > > ``` * * * ``` # Check CronJob status kubectl get cronjob -n kafka-backup # Check recent jobs kubectl get jobs -n kafka-backup --sort-by='.metadata.creationTimestamp' # Check events kubectl get events -n kafka-backup --sort-by='.lastTimestamp' | tail -20 ``` ``` # Get logs from the most recent job kubectl logs -n kafka-backup -l app.kubernetes.io/instance=kafka-backup --tail=200 # Or find the specific pod kubectl get pods -n kafka-backup kubectl logs -n kafka-backup ``` ``` # Verify service account annotations kubectl get sa -n kafka-backup -o yaml # Test IAM from within the pod (AWS) kubectl run aws-test --rm -it --image=amazon/aws-cli \ --overrides='{"spec":{"serviceAccountName":"kafka-backup-enterprise"}}' \ -n kafka-backup -- sts get-caller-identity ``` ``` # Run the Helm test helm test kafka-backup -n kafka-backup # Or manually check kubectl run license-check --rm -it \ --image=osodevops/kafka-backup-enterprise:0.3.2 \ --env="ENTERPRISE_LICENSE_KEY=$(kubectl get secret kafka-backup-license -n kafka-backup -o jsonpath='{.data.license-b64}' | base64 -d)" \ -n kafka-backup -- license info ``` ``` # Check if suspended kubectl get cronjob -n kafka-backup -o jsonpath='{.items[0].spec.suspend}' # Check startingDeadlineSeconds — if the controller was down when the job # was due, it may have been skipped kubectl describe cronjob -n kafka-backup kafka-backup-kafka-backup-enterprise ``` * * * | Value | Default | Description | | --- | --- | --- | | `image.repository` | `osodevops/kafka-backup-enterprise` | Docker image | | `image.tag` | Chart appVersion | Image tag | | `cronjob.enabled` | `true` | Create a CronJob | | `cronjob.schedule` | `0 2 * * *` | Cron schedule | | `cronjob.concurrencyPolicy` | `Forbid` | Overlap prevention | | `cronjob.activeDeadlineSeconds` | `7200` | Maximum runtime | | `job.enabled` | `false` | Create a one-shot Job | | `job.command` | `backup` | `backup` or `restore` | | `license.existingSecret` | `""` | Secret with license key | | `license.key` | `""` | Inline license (dev only) | | `config.backupConfig` | _(minimal example)_ | Raw kafka-backup YAML | | `config.existingConfigMap` | `""` | Use pre-existing ConfigMap | | `credentials.existingSecret` | `""` | Secret with credential env vars | | `credentials.inline` | `{}` | Inline credentials (dev only) | | `metrics.enabled` | `false` | Enable Prometheus metrics | | `metrics.serviceMonitor.enabled` | `false` | Create ServiceMonitor | | `trialPersistence.enabled` | `false` | PVC for trial state file | | `serviceAccount.create` | `true` | Create ServiceAccount | | `serviceAccount.annotations` | `{}` | SA annotations (IRSA, WI) | For the complete `values.yaml` with all options, see the [chart source](https://github.com/osodevops/kafka-backup-enterprise/tree/main/deploy/helm/kafka-backup-enterprise). - [Licensing](https://kafkabackup.com/enterprise/licensing.md) — obtain and manage your license - [Schema Registry Backup](https://kafkabackup.com/enterprise/schema-registry.md) — configure Schema Registry backup - [Confluent RBAC Backup](https://kafkabackup.com/enterprise/confluent-rbac.md) — back up MDS role bindings - [Monitoring Setup](https://kafkabackup.com/guides/monitoring-setup.md) — configure alerts and dashboards --- title: Schema Registry Backup & Restore description: Back up and restore Confluent Schema Registry with OSO Kafka Backup Enterprise source_url: html: https://kafkabackup.com/enterprise/schema-registry md: https://kafkabackup.com/enterprise/schema-registry.md --- # Schema Registry Backup & Restore OSO Kafka Backup Enterprise backs up Confluent Schema Registry alongside your Kafka data, ensuring restored messages remain readable. Kafka messages serialized with Avro, JSON Schema, or Protobuf embed a 4-byte schema ID that references an external Schema Registry. Without backing up schemas, restored messages are useless: ``` Without Schema Backup: Kafka data ──── Backed up ✓ Schema IDs ──── NOT backed up ✗ Result: "Schema ID 42 not found" With Schema Backup: Kafka data ──── Backed up ✓ Schema IDs ──── Backed up ✓ Result: Consumers work correctly ``` Add the `enterprise.schema_registry` section to your existing backup config: backup-config.yaml ``` mode: backup backup_id: daily-2026-04-06 source: bootstrap_servers: ["kafka:9092"] storage: backend: s3 bucket: kafka-backups enterprise: schema_registry: url: "https://schema-registry:8081" auth: type: basic username: ${SR_USERNAME} password: ${SR_PASSWORD} ``` Run the backup — schemas are captured automatically alongside Kafka data: ``` kafka-backup backup --config backup-config.yaml ``` Or back up schemas only (no Kafka data): ``` kafka-backup backup --config backup-config.yaml --schema-only ``` | Item | Description | | --- | --- | | **Subjects** | All subject names matching filter patterns | | **Schema versions** | Every version of every subject: schema definition, type, ID, references | | **Global compatibility** | Global compatibility level (BACKWARD, FORWARD, FULL, NONE, etc.) | | **Per-subject compatibility** | Subject-specific compatibility overrides | | **Global mode** | Registry mode (READWRITE, READONLY, IMPORT) | | **Per-subject mode** | Subject-specific mode overrides | | **Schema references** | Cross-schema dependencies (Protobuf imports, Avro references) | Three authentication methods are supported: Basic Auth (most common) ``` enterprise: schema_registry: url: "https://schema-registry:8081" auth: type: basic username: ${SR_USERNAME} password: ${SR_PASSWORD} ``` Mutual TLS (mTLS) ``` enterprise: schema_registry: url: "https://schema-registry:8081" auth: type: mtls ca_cert: /certs/ca.crt client_cert: /certs/client.crt client_key: /certs/client.key ``` No Authentication ``` enterprise: schema_registry: url: "http://schema-registry:8081" # auth section omitted = no authentication ``` For HTTPS endpoints with custom CA certificates: ``` enterprise: schema_registry: url: "https://schema-registry:8081" tls: ca_cert: /certs/custom-ca.crt ``` Control which subjects are backed up using glob patterns: ``` enterprise: schema_registry: backup: # Include patterns (default: ["*"] = all subjects) subjects: - "orders-*" - "payments-*" # Exclude patterns (applied after include) exclude: - "*-test" - "*-internal" # Include soft-deleted subjects (default: false) include_soft_deleted: false # Back up all versions or only latest (default: all) include_versions: all # or "latest" # Auto-include referenced subjects outside filter (default: true) include_references: true ``` ``` enterprise: schema_registry: restore: # Restore strategy (default: preserve) strategy: preserve # preserve | overwrite | skip # Force original schema IDs using IMPORT mode (default: false) # Requires: empty target registry + mode.mutability=true force_ids: false # Rewrite schema IDs in Kafka message bytes (default: false) # Use when target registry assigns different IDs rewrite_ids: false # Subject name mapping for environment cloning subject_mapping: "orders-value": "staging-orders-value" # Dry-run mode (default: false) dry_run: false ``` **Restore strategies:** | Strategy | Behavior | Use Case | | --- | --- | --- | | **preserve** (default) | Keep existing schemas on target; add missing | Merge / partial recovery | | **overwrite** | Replace existing schemas from backup | Full DR to clean target | | **skip** | Only restore subjects that don't exist on target | Add new schemas only | ``` enterprise: schema_registry: connection: timeout_ms: 30000 # Request timeout (default: 30s) max_retries: 3 # Retry attempts on 429/5xx retry_backoff_ms: 1000 # Initial retry backoff rate_limit_rps: 25 # Max requests per second concurrent_requests: 4 # Max parallel API calls ``` Schema backups are stored alongside Kafka data in the same storage backend: ``` {backup_id}/ manifest.json # Main backup manifest topics/... # Kafka data segments schema-registry/ # Schema Registry backup _manifest.json # Schema manifest (subjects, counts, dependency order) _global_config.json # Global compatibility and mode subjects/ orders-value/ _metadata.json # Subject config, mode, version list v1.json # Schema version 1 (full definition) v2.json # Schema version 2 payments-value/ _metadata.json v1.json ``` Schemas can reference other schemas (e.g., Protobuf imports). The backup engine automatically: 1. Discovers all cross-subject references 2. Builds a dependency graph (DAG) 3. Performs a topological sort 4. Records the correct restore order in the manifest On restore, schemas are registered in dependency order — referenced schemas first, then schemas that reference them. Circular references are detected and reported as errors. Back up everything, restore to a fresh environment: ``` enterprise: schema_registry: url: "https://dr-schema-registry:8081" restore: strategy: preserve force_ids: true # Empty target — preserve original IDs ``` Schema IDs will differ on the target. Use ID rewriting: ``` enterprise: schema_registry: url: "https://psrc-xxxxx.confluent.cloud" auth: type: basic username: ${CCLOUD_SR_API_KEY} password: ${CCLOUD_SR_API_SECRET} restore: strategy: preserve rewrite_ids: true # Rewrite IDs in Kafka message bytes ``` Periodic schema snapshots for regulatory requirements: ``` kafka-backup backup --config schema-archive.yaml --schema-only ``` | Issue | Cause | Solution | | --- | --- | --- | | `Schema ID not found` after restore | Schemas not restored before data | Schema restore runs automatically before data restore | | `Subject not found` during backup | Subject was deleted | Enable `include_soft_deleted: true` | | Circular reference error | Schemas reference each other in a cycle | Fix the cycle in Schema Registry, then re-backup | | 429 rate limit errors | Too many API calls | Reduce `rate_limit_rps` or increase `retry_backoff_ms` | | Authentication failed | Invalid credentials | Check `SR_USERNAME`/`SR_PASSWORD` env vars | - Confluent Schema Registry (Platform or Cloud) - HTTP(S) access to Schema Registry REST API - Enterprise license with `schema_registry` feature enabled --- title: Confluent RBAC Backup & Restore description: Back up and restore Confluent Platform RBAC role bindings from the Metadata Service (MDS) source_url: html: https://kafkabackup.com/enterprise/confluent-rbac md: https://kafkabackup.com/enterprise/confluent-rbac.md --- # Confluent RBAC Backup & Restore OSO Kafka Backup Enterprise backs up Confluent Platform RBAC role bindings from the Metadata Service (MDS), ensuring your security posture is recoverable after a disaster. Confluent Platform customers running RBAC accumulate hundreds or thousands of role bindings across Kafka, Schema Registry, Connect, ksqlDB, and Flink. After a DR event: - **Without RBAC backup:** Every binding must be manually recreated. A cluster with 500+ bindings across 6 component types takes 4-8 hours and has a 15-25% human error rate. - **With RBAC backup:** All bindings are restored in under 5 minutes in the correct dependency order. Components start successfully on first attempt. rbac-backup-config.yaml ``` mode: backup backup_id: rbac-snapshot-2026-04-06 storage: backend: s3 bucket: kafka-backups prefix: production enterprise: confluent_rbac: mds_url: "https://mds.prod.company.com:8090" auth: username: ${MDS_USERNAME} password: ${MDS_PASSWORD} ``` ``` kafka-backup rbac-backup --config rbac-backup-config.yaml ``` Add `confluent_rbac` alongside your existing config — everything backs up together: full-backup-config.yaml ``` mode: backup backup_id: daily-2026-04-06 source: bootstrap_servers: ["kafka:9092"] storage: backend: s3 bucket: kafka-backups enterprise: schema_registry: url: "https://schema-registry:8081" auth: type: basic username: ${SR_USERNAME} password: ${SR_PASSWORD} confluent_rbac: mds_url: "https://mds:8090" auth: username: ${MDS_USERNAME} password: ${MDS_PASSWORD} ``` ``` kafka-backup backup --config full-backup-config.yaml ``` This runs Kafka data backup, Schema Registry backup, and RBAC backup in sequence. | Item | Description | | --- | --- | | **Cluster registry** | All clusters registered in MDS (Kafka, Connect, SR, ksqlDB, Flink) | | **Principals** | All users and groups with role bindings | | **Cluster-scoped bindings** | Roles granted at the cluster level (e.g., SystemAdmin, ClusterAdmin) | | **Resource-scoped bindings** | Roles on specific resources (e.g., DeveloperRead on Topic:orders) | | **Resource patterns** | Both LITERAL and PREFIXED pattern types | | **All predefined roles** | SystemAdmin, UserAdmin, ClusterAdmin, SecurityAdmin, AuditAdmin, Operator, ResourceOwner, DeveloperRead, DeveloperWrite, DeveloperManage | MDS uses a two-phase authentication model: Basic credentials are exchanged for a JWT bearer token, which is then used for all subsequent API calls. The token is automatically refreshed when it expires. ``` enterprise: confluent_rbac: mds_url: "https://mds:8090" auth: username: ${MDS_USERNAME} password: ${MDS_PASSWORD} ``` > [!NOTE] > > [!NOTE] > > Required MDS Permissions > > [!NOTE] > > - **For backup:** `SecurityAdmin` on all scopes (read-only access to bindings) > > - **For restore:** `UserAdmin` on the root kafka-cluster (can create/modify bindings) > > > > Recommended: Create a dedicated service account `User:kafka-backup-rbac` with these roles. ``` enterprise: confluent_rbac: mds_url: "https://mds:8090" auth: username: admin password: secret tls: ca_cert: /certs/ca.pem client_cert: /certs/client.pem # For mTLS client_key: /certs/client-key.pem # For mTLS ``` Control which principals are included in the backup using glob patterns: ``` enterprise: confluent_rbac: backup: # Include patterns (default: ["*"] = all principals) principals: - "User:app-*" - "Group:*-team" # Exclude patterns exclude_principals: - "User:kafka-backup-rbac" # Exclude the backup service account itself # Filter by specific Kafka clusters (default: all discovered) cluster_filter: - "lkc-abc123" # Include sub-clusters: Connect, SR, ksqlDB, Flink (default: true) include_sub_clusters: true ``` MDS has a rate limit of 15 requests per second. The default is 12 RPS to leave headroom: ``` enterprise: confluent_rbac: connection: timeout_ms: 30000 # Request timeout (default: 30s) max_retries: 3 # Retry attempts on 429/5xx retry_backoff_ms: 1000 # Initial retry backoff (doubles per retry) rate_limit_rps: 12 # Max requests per second (default: 12) ``` RBAC bindings have implicit dependencies. The restore engine applies bindings in 5 tiers to ensure correct ordering: | Tier | Priority | Roles | Purpose | | --- | --- | --- | --- | | **0** | Bootstrap | SystemAdmin, UserAdmin (kafka-cluster scope) | Root admin access — must exist first | | **1** | Components | Service accounts for SR, Connect, ksqlDB, Flink, C3 | Components need bindings to start | | **2** | Cluster Admin | ClusterAdmin, SecurityAdmin, AuditAdmin, Operator | Cluster-level administrative roles | | **3** | Resource Owner | ResourceOwner | Ownership of specific resources | | **4** | Developer | DeveloperRead, DeveloperWrite, DeveloperManage | Application-level access | This ordering ensures: - Admin accounts exist before they can delegate access - Component service accounts (Connect, Schema Registry, ksqlDB) have their required bindings before the components restart - Resource owners are established before developer access is granted RBAC snapshots are stored in the backup storage alongside Kafka data: ``` {backup_id}/ manifest.json # Main backup manifest topics/... # Kafka data schema-registry/... # Schema Registry backup enterprise/ confluent-rbac/ rbac-snapshot.json # Full RBAC snapshot (bindings + metadata) cluster-registry.json # MDS cluster registry export rbac-metadata.json # Backup statistics ``` The `rbac-snapshot.json` contains the complete security posture: ``` { "metadata": { "mds_url": "https://mds.prod:8090", "backup_timestamp": "2026-04-06T10:00:00Z", "backup_id": "daily-2026-04-06" }, "cluster_registry": [ { "clusterName": "prod-kafka", "scope": { "clusters": { "kafka-cluster": "lkc-abc123" } } } ], "bindings": [ { "principal": "User:admin", "role_name": "SystemAdmin", "scope": { "clusters": { "kafka-cluster": "lkc-abc123" } }, "binding_type": "ClusterScoped" }, { "principal": "User:alice", "role_name": "DeveloperRead", "scope": { "clusters": { "kafka-cluster": "lkc-abc123" } }, "binding_type": "ResourceScoped", "resource_patterns": [ { "resourceType": "Topic", "name": "orders", "patternType": "LITERAL" } ] } ], "stats": { "total_bindings": 347, "cluster_scoped_bindings": 23, "resource_scoped_bindings": 324, "unique_principals": 42, "unique_roles": 9, "cluster_scopes_count": 5, "duration_ms": 21400 } } ``` ``` # Backup from production enterprise: confluent_rbac: mds_url: "https://mds.prod:8090" auth: username: ${MDS_PROD_USER} password: ${MDS_PROD_PASS} # Restore to DR site (future milestone) # enterprise: # confluent_rbac: # restore: # mds_url: "https://mds.dr:8090" # cluster_id_mapping: # kafka_cluster: # "lkc-prod": "lkc-dr" ``` Periodic RBAC snapshots prove security posture continuity: ``` # Schedule daily RBAC snapshots kafka-backup rbac-backup --config rbac-audit.yaml # Compare snapshots to detect drift (future milestone) # kafka-backup rbac-diff --backup-id snap-1 --backup-id snap-2 ``` Clone security posture from production to staging with principal remapping (future milestone). The backup captures all required bindings for Confluent Platform components: | Component | Required Bindings | | --- | --- | | **Schema Registry** | SecurityAdmin on SR scope + ResourceOwner on `_schemas` topic and coordination group | | **Connect** | SecurityAdmin on Connect scope + ResourceOwner on configs, offsets, status topics + consumer group | | **ksqlDB** | SecurityAdmin on ksqlDB scope + ResourceOwner on ksqlDB cluster, command topic, processing log, consumer group prefix | | **Control Center** | SystemAdmin on kafka-cluster | | **Flink** | SecurityAdmin on CMF/Flink environment scope | These bindings are automatically classified as Tier 1 (ComponentService) and restored early to ensure components can start. | Issue | Cause | Solution | | --- | --- | --- | | Authentication failed | Invalid MDS credentials | Verify `MDS_USERNAME`/`MDS_PASSWORD` env vars | | 403 Forbidden | Backup user lacks SecurityAdmin | Grant SecurityAdmin on kafka-cluster scope to backup service account | | 429 Too Many Requests | MDS rate limit exceeded | Reduce `rate_limit_rps` (default 12 is usually safe) | | No clusters found | MDS cluster registry is empty | Verify MDS is running and clusters are registered | | Token expired errors | Long backup with token expiry | Token auto-refreshes; if persistent, check MDS token configuration | | Partial backup | Some principals failed to enumerate | Check MDS logs; backup continues with available data | - Confluent Platform 5.4+ with RBAC enabled - MDS running on Confluent Server brokers - HTTP(S) access to MDS REST API (port 8090 by default) - SecurityAdmin or SystemAdmin role for the backup service account - Enterprise license with `rbac` feature enabled --- title: Licensing description: OSO Kafka Backup Enterprise licensing and activation source_url: html: https://kafkabackup.com/enterprise/licensing md: https://kafkabackup.com/enterprise/licensing.md --- # Licensing OSO Kafka Backup Enterprise uses Ed25519-signed license files validated entirely offline. No license server, no phone-home, no network dependency. The license file is self-contained and cryptographically verifiable. | Tier | Target | Features | Pricing | | --- | --- | --- | --- | | **OSS** | Everyone | Core backup/restore, PITR, three-phase restore, all storage backends, retention (`prune`) | Free (MIT) | | **Enterprise** | Compliance-driven and multi-cluster platforms | All enterprise features (Schema Registry, Apicurio, Confluent RBAC, CSFLE metadata, MSK KRaft migration, GDPR erasure) and support | Annual **site licence** covering all your Kafka clusters — no per-cluster or per-broker count | | **Trial** | Evaluation | All enterprise features, 14 days, built into the binary | Free | Each license encodes which features are enabled: | Feature Flag | Description | | --- | --- | | `schema_registry` | Confluent Schema Registry backup & restore | | `rbac` | Confluent RBAC backup & restore (MDS) | | `migrations:msk-kraft` | AWS MSK ZooKeeper to KRaft migration execution | | `migrations:msk-kraft:plan` | AWS MSK ZooKeeper to KRaft plan/precheck entitlement | | `encryption` | Confluent CSFLE and DEK Registry metadata backup | | `masking` | Data masking for non-production (planned) | | `audit` | Audit logging (planned) | | `wasm` | WebAssembly plugin system (planned) | | `erasure` | GDPR erasure: restore-time key suppression (redaction and crypto-shredding to follow) | | `support` | Priority support SLA entitlement | The enterprise binary includes a **14-day free trial** on first run — no signup, no credit card. Just install the enterprise binary and start using all features immediately. For an extended 30-day trial, request one at: - **Web:** [enterprise.kafkabackup.com](https://enterprise.kafkabackup.com) - **Email:** [sales@oso.sh](mailto:sales@oso.sh) Buy a Standard or Enterprise license at **[enterprise.kafkabackup.com](https://enterprise.kafkabackup.com)** or contact sales directly: - **Web:** [enterprise.kafkabackup.com](https://enterprise.kafkabackup.com) - **Email:** [sales@oso.sh](mailto:sales@oso.sh) Pricing is per Kafka cluster per year. A cluster can grow from 3 to 30 brokers without changing the license cost. Licenses are distributed as PEM-encoded files containing a JSON payload signed with Ed25519: ``` -----BEGIN KAFKA-BACKUP LICENSE----- eyJsaWNlbnNlX2lkIjoiYWJjZGVmMTItMzQ1Ni03ODkwLWFiY2QtZWYxMjM0 NTY3ODkwIiwiY3VzdG9tZXJfaWQiOiJjdXN0XzEyMzQ1NiIsImN1c3RvbWVy ... -----END KAFKA-BACKUP LICENSE----- ``` The license payload contains: | Field | Description | | --- | --- | | `license_id` | Unique identifier (UUID) | | `customer_id` | Customer identifier | | `customer_name` | Organization name | | `customer_email` | Contact email | | `tier` | `standard` or `enterprise` | | `features` | List of enabled feature flags | | `max_clusters` | Maximum Kafka clusters (optional) | | `max_nodes` | Maximum brokers across all clusters (optional) | | `issued_at` | Issue date (ISO 8601) | | `expires_at` | Expiry date (optional — `null` = perpetual) | | `schema_version` | License format version | ``` kafka-backup license apply --file /path/to/license.key ``` This validates the license and saves it to `~/.config/kafka-backup/license.key`. Output: ``` License applied successfully. License ID: abcdef12-3456-7890-abcd-ef1234567890 Customer: Acme Corp Tier: Enterprise Saved to: /home/user/.config/kafka-backup/license.key ``` ``` kafka-backup license verify --file /path/to/license.key ``` Output: ``` License is valid. License ID: abcdef12-3456-7890-abcd-ef1234567890 Customer: Acme Corp Tier: Enterprise Features: encryption, schema_registry, rbac, audit, support Expires: 2027-01-15 (285 days remaining) ``` ``` kafka-backup license info ``` Output when licensed: ``` License ID: abcdef12-3456-7890-abcd-ef1234567890 Customer: Acme Corp (acme@example.com) Tier: Enterprise Features: encryption, schema_registry, rbac, audit, support Max clusters: 5 Issued: 2026-01-15 Expires: 2027-01-15 (285 days remaining) Status: Valid ``` Output when unlicensed: ``` No license found. Running in OSS mode. To activate enterprise features: kafka-backup license apply --file /path/to/license.key Get a license at https://enterprise.kafkabackup.com ``` The license is loaded automatically at startup. The tool searches these locations in order: | Priority | Source | Description | | --- | --- | --- | | 1 | `ENTERPRISE_LICENSE_KEY` env var | Base64-encoded PEM content (for CI/CD) | | 2 | `ENTERPRISE_LICENSE_FILE` env var | Path to license file | | 3 | `/etc/kafka-backup/license.key` | Default for containers and Kubernetes | | 4 | `~/.config/kafka-backup/license.key` | Default for CLI users | For Kubernetes deployments, mount the license as a secret: ``` apiVersion: v1 kind: Secret metadata: name: kafka-backup-license namespace: kafka-backup type: Opaque stringData: license.key: | -----BEGIN KAFKA-BACKUP LICENSE----- eyJsaWNlbnNlX2lkIjoiYWJjZGVmMTItMzQ1Ni03ODkwLWFiY2QtZWYxMjM0 ... -----END KAFKA-BACKUP LICENSE----- ``` Mount it in your deployment: ``` volumes: - name: license secret: secretName: kafka-backup-license volumeMounts: - name: license mountPath: /etc/kafka-backup readOnly: true ``` For CI/CD pipelines, pass the license as a base64-encoded environment variable: ``` export ENTERPRISE_LICENSE_KEY=$(base64 < license.key) kafka-backup backup --config backup.yaml ``` Enterprise features never block your backup. If a feature is configured but not licensed: - A warning is logged: `"Schema Registry backup is configured but not licensed. Skipping."` - The backup continues without the enterprise feature - Kafka data backup runs normally - Exit code is 0 (success) This means you can deploy the enterprise binary everywhere and activate features per-environment by providing licenses only where needed. License validation is fully offline: - The Ed25519 public key is embedded in the binary at compile time - No network calls, no license server, no DNS dependency - Works in air-gapped environments - Validation checks: signature integrity, expiry date, feature flags - The **private signing key** is held exclusively by OSO DevOps — never distributed - The **public verification key** is embedded in the enterprise binary - Licenses cannot be forged or modified without the private key (Ed25519 is 128-bit security) - License files contain customer information but no secrets — safe to store in version control or config management **Q: What happens when the license expires?** Enterprise features stop activating. The binary continues to work as an OSS-equivalent tool. No data loss, no downtime. **Q: Can I use the same license on multiple machines?** Yes. The licence is a site licence: it is not tied to a machine or MAC address and is not limited by the number of clusters or brokers. The `max_clusters` / `max_nodes` fields in the payload are informational and not enforced by the binary. **Q: Do I need internet access to validate the license?** No. Validation is 100% offline using Ed25519 signature verification. **Q: What if I lose my license file?** Contact [sales@oso.sh](mailto:sales@oso.sh) — we can reissue your license from our records. --- title: Role-Based Access Control (RBAC) description: Role-based access control for OSO Kafka Backup Enterprise source_url: html: https://kafkabackup.com/enterprise/rbac md: https://kafkabackup.com/enterprise/rbac.md --- # Role-Based Access Control (RBAC) OSO Kafka Backup Enterprise provides fine-grained access control to backup and restore operations. RBAC allows you to: - Control who can perform backup operations - Control who can perform restore operations - Restrict access to specific topics - Limit access to specific storage locations - Audit all access attempts ``` enterprise: rbac: enabled: true # Authentication method auth: method: oidc issuer: https://auth.company.com client_id: kafka-backup audience: kafka-backup-api # Role definitions roles: - name: backup-admin permissions: - "*" - name: backup-operator permissions: - backup:* - list:* - describe:* - validate:* - name: restore-operator permissions: - restore:* - list:* - describe:* - validate:* - name: readonly permissions: - list:* - describe:* ``` ``` enterprise: rbac: auth: method: oidc issuer: https://auth.company.com client_id: kafka-backup audience: kafka-backup-api scopes: - openid - profile - kafka-backup # Map OIDC claims to roles role_claim: roles # Or use groups # group_claim: groups ``` ``` enterprise: rbac: auth: method: ldap server: ldap://ldap.company.com:389 bind_dn: cn=service,dc=company,dc=com bind_password: ${LDAP_PASSWORD} user_base: ou=users,dc=company,dc=com user_filter: "(uid={0})" group_base: ou=groups,dc=company,dc=com group_filter: "(member={0})" ``` ``` enterprise: rbac: auth: method: static users: - username: admin password_hash: "$2b$12$..." # bcrypt hash roles: - backup-admin - username: operator password_hash: "$2b$12$..." roles: - backup-operator ``` ``` enterprise: rbac: auth: method: kubernetes # Uses Kubernetes RBAC # Roles mapped from ServiceAccount annotations ``` ``` :: ``` Examples: - `backup:*` - All backup operations - `restore:topic:orders` - Restore only orders topic - `list:*` - List all backups - `describe:backup:*` - Describe any backup | Operation | Description | | --- | --- | | `backup` | Create backups | | `restore` | Restore from backups | | `list` | List backups | | `describe` | View backup details | | `validate` | Validate backups | | `delete` | Delete backups | | `config` | Modify configuration | | `offset-reset` | Reset consumer offsets | | Resource | Description | | --- | --- | | `*` | All resources | | `topic:` | Specific topic | | `backup:` | Specific backup | | `storage:` | Storage location | | `group:` | Consumer group | ``` # Exact match - "restore:topic:orders" # Wildcard - "backup:topic:*" # Pattern - "restore:topic:orders-*" # Multiple patterns - "backup:topic:orders-*" - "backup:topic:payments-*" ``` ``` roles: # Full administrative access - name: admin permissions: - "*" # Can backup any topic - name: backup-operator permissions: - "backup:*" - "list:*" - "describe:*" - "validate:*" # Can restore any topic - name: restore-operator permissions: - "restore:*" - "list:*" - "describe:*" - "validate:*" - "offset-reset:*" # Read-only access - name: viewer permissions: - "list:*" - "describe:*" ``` ``` roles: # Team-specific role - name: orders-team permissions: - "backup:topic:orders" - "backup:topic:orders-*" - "restore:topic:orders" - "restore:topic:orders-*" - "list:*" - "describe:*" # Environment-specific role - name: prod-backup-only permissions: - "backup:storage:s3://prod-backups/*" - "list:storage:s3://prod-backups/*" - "describe:storage:s3://prod-backups/*" # Consumer offset management - name: offset-manager permissions: - "offset-reset:*" - "list:*" - "describe:*" ``` ``` enterprise: rbac: bindings: - user: alice@company.com roles: - backup-admin - user: bob@company.com roles: - backup-operator - restore-operator - user: charlie@company.com roles: - viewer ``` ``` enterprise: rbac: bindings: - group: platform-team roles: - backup-admin - group: dev-team roles: - backup-operator - group: support-team roles: - viewer ``` ``` enterprise: rbac: bindings: - serviceAccount: backup-cronjob namespace: kafka-backup roles: - backup-operator - serviceAccount: restore-service namespace: kafka-backup roles: - restore-operator ``` ``` apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: name: kafka-backup-operator namespace: kafka-backup rules: - apiGroups: ["kafka.oso.sh"] resources: ["kafkabackups"] verbs: ["get", "list", "watch", "create", "update"] - apiGroups: ["kafka.oso.sh"] resources: ["kafkarestores"] verbs: ["get", "list", "watch"] # No create - read only --- apiVersion: rbac.authorization.k8s.io/v1 kind: RoleBinding metadata: name: backup-operator-binding namespace: kafka-backup subjects: - kind: User name: alice@company.com apiGroup: rbac.authorization.k8s.io roleRef: kind: Role name: kafka-backup-operator apiGroup: rbac.authorization.k8s.io ``` ``` apiVersion: v1 kind: ServiceAccount metadata: name: backup-service namespace: kafka-backup annotations: kafka-backup.oso.sh/roles: "backup-operator,viewer" ``` ``` # OIDC login (opens browser) kafka-backup auth login # OIDC with device code (for headless) kafka-backup auth login --device-code # Username/password kafka-backup auth login --username alice --password-stdin ``` ``` # Show current user kafka-backup auth whoami # Output: # User: alice@company.com # Roles: backup-admin # Expires: 2024-12-02T10:00:00Z # Refresh token kafka-backup auth refresh # Logout kafka-backup auth logout ``` ``` # Use service account token export KAFKA_BACKUP_TOKEN="eyJ..." # Or via config kafka-backup --token-file /var/run/secrets/token backup ... ``` ``` Error: Access denied Operation: restore Resource: topic:production-orders User: bob@company.com Roles: backup-operator Required permission: restore:topic:production-orders User permissions: - backup:* - list:* - describe:* ``` ``` { "timestamp": "2024-12-01T10:00:00Z", "event": "access_denied", "user": "bob@company.com", "operation": "restore", "resource": "topic:production-orders", "required_permission": "restore:topic:production-orders", "user_roles": ["backup-operator"], "client_ip": "10.0.0.50" } ``` ``` roles: - name: prod-admin permissions: - "*:storage:s3://prod-backups/*" - name: staging-admin permissions: - "*:storage:s3://staging-backups/*" - name: dev-admin permissions: - "*:storage:s3://dev-backups/*" bindings: - group: prod-team roles: [prod-admin] - group: dev-team roles: [staging-admin, dev-admin] ``` ``` roles: - name: orders-team permissions: - "*:topic:orders*" - "list:*" - "describe:*" - name: payments-team permissions: - "*:topic:payments*" - "list:*" - "describe:*" - name: analytics-team permissions: - "backup:topic:analytics*" - "restore:topic:analytics*" - "list:*" - "describe:*" ``` ``` roles: # Automated hourly backup - minimal permissions - name: automated-backup permissions: - "backup:topic:*" - "list:backup:*" # To check existing backups # Automated validation - read only - name: automated-validation permissions: - "validate:*" - "list:*" - "describe:*" bindings: - serviceAccount: hourly-backup-job roles: [automated-backup] - serviceAccount: validation-job roles: [automated-validation] ``` ``` # Show detailed permission checks kafka-backup --rbac-debug backup --config backup.yaml ``` ``` kafka-backup auth permissions # Output: # User: alice@company.com # Roles: backup-operator, viewer # # Effective Permissions: # backup:* # list:* # describe:* # validate:* ``` ``` kafka-backup auth can-i restore topic:orders # Output: # yes (via role: restore-operator) kafka-backup auth can-i delete backup:my-backup # Output: # no (no matching permission) ``` 1. **Use groups over users** - Easier to manage 2. **Least privilege** - Grant minimum required permissions 3. **Separate environments** - Different roles per environment 4. **Regular audits** - Review role assignments 5. **Service accounts for automation** - Don't use personal credentials 6. **Document role purposes** - Clear descriptions - [Audit Logging](https://kafkabackup.com/enterprise/audit-logging.md) - Log all access attempts - [Security Setup](https://kafkabackup.com/guides/security-setup.md) - TLS and SASL configuration - [Kubernetes Operator](https://kafkabackup.com/operator.md) - Kubernetes RBAC integration --- title: Confluent CSFLE metadata backup description: Back up Confluent CSFLE keys, DEK Registry entries, encrypted subjects, and schema encryption rules with OSO Kafka Backup Enterprise source_url: html: https://kafkabackup.com/enterprise/encryption md: https://kafkabackup.com/enterprise/encryption.md --- # Confluent CSFLE metadata backup OSO Kafka Backup Enterprise backs up metadata from Confluent's Client-Side Field-Level Encryption (CSFLE) system. It captures Key Encryption Keys (KEKs), Data Encryption Keys (DEKs), encrypted subjects, and the Schema Registry rules that describe which fields are encrypted. > [!NOTE] > > [!NOTE] > > Scope > > [!NOTE] > > This feature preserves **existing Confluent CSFLE metadata**. It does not encrypt Kafka backup segment files, and it does not add encryption to plaintext Kafka records. Confluent CSFLE encrypts fields in the producer or serializer before the records reach Kafka; Kafka Backup Enterprise preserves the metadata needed to understand and recover that setup. > > > > The open source CLI does not provide backup encryption. The Strimzi Backup Operator does not support `spec.backup.encryption`. | Item | Captured data | | --- | --- | | KEK Registry | KEK names, KMS type, KMS key identifier, properties, and status | | DEK Registry | Subject, version, algorithm, and encrypted key material for each DEK | | Encrypted subjects | Subjects with Confluent `ENCRYPT` rules, their KEK, algorithm, and encrypted-field count | | Schema encryption rules | Raw Schema Registry versions containing encryption rules and metadata tags | | Encryption manifest | Counts, KMS types, backup timestamp, and per-subject summaries | The feature copies encrypted DEK material and KMS references. It does not export KMS master keys or other plaintext key material. - OSO Kafka Backup Enterprise v0.3.2 or later - A trial or license containing the `encryption` feature - Confluent Schema Registry with the DEK Registry API available - Existing Confluent CSFLE schemas and keys - `enterprise.schema_registry` in the same configuration file - A configured Kafka source and storage backend for the normal backup The DEK Registry is part of the Confluent Schema Registry API. The encryption backup therefore cannot run without `enterprise.schema_registry`, even when `dek_registry_url` is set explicitly. The smallest complete configuration looks like this: ``` mode: backup backup_id: daily-2026-07-21 source: bootstrap_servers: - kafka:9092 storage: backend: s3 bucket: kafka-backups region: eu-west-1 enterprise: schema_registry: enabled: true url: https://schema-registry:8081 auth: type: basic username: ${SR_USERNAME} password: ${SR_PASSWORD} encryption: enabled: true backup: include_dek_registry: true include_encryption_rules: true kek_filter: - "prod-*" exclude_keks: - "prod-test-*" ``` Authentication and TLS settings inherit from `enterprise.schema_registry` by default. Set `enterprise.encryption.auth` or `enterprise.encryption.tls` only when the DEK Registry needs different connection settings. | Field | Default | Description | | --- | --- | --- | | `enabled` | `true` | Runs the encryption metadata backup when the section is present | | `dek_registry_url` | Derived from Schema Registry | Overrides the DEK Registry base URL | | `auth` | Schema Registry auth | Optional DEK Registry authentication override | | `tls.ca_cert` | Schema Registry TLS | Optional path to a CA certificate for the DEK Registry | | `backup.include_dek_registry` | `true` | Exports KEKs and DEKs | | `backup.include_encryption_rules` | `true` | Exports schema versions containing encryption rules | | `backup.kek_filter` | `["*"]` | Glob patterns selecting KEK names | | `backup.exclude_keks` | `[]` | Glob patterns excluding KEK names | | `connection.timeout_ms` | `30000` | HTTP request timeout in milliseconds | | `connection.max_retries` | `3` | Maximum retry attempts | | `connection.retry_backoff_ms` | `1000` | Initial retry backoff in milliseconds | | `connection.rate_limit_rps` | `25` | Maximum DEK Registry requests per second | Environment-variable interpolation such as `${SR_USERNAME}` and `${SR_PASSWORD}` is supported in configuration values. ``` kafka-backup backup --config backup.yaml ``` A successful run logs a summary similar to: ``` Starting encryption metadata backup Encryption backup complete: 3 KEKs, 18 DEKs, 7 encrypted subjects ``` If encryption is configured but the current trial or license does not include the `encryption` feature, the enterprise binary logs a warning, skips the encryption metadata step, and continues the normal Kafka data backup. Artifacts are written below the normal backup ID: ``` daily-2026-07-21/ └── enterprise/ └── encryption/ ├── _manifest.json ├── kek-registry.json ├── dek-registry.json └── schemas/ └── / └── v.json ``` `_manifest.json` records the number of KEKs, DEKs, and encrypted subjects. The registry exports contain the full metadata returned by the DEK Registry. Schema files retain the raw encryption `ruleSet` and metadata tags. The supported v0.3.2 CLI workflow exports encryption metadata during backup. Automated encryption metadata import is not currently exposed by the `kafka-backup restore` command. Do not rely on `enterprise.encryption.restore` options for a production recovery workflow until CLI restore support is released and documented here. Keep the source KMS keys and access policies available for recovery. The backup contains encrypted DEKs and KMS references, not the KMS master keys themselves. These controls protect different parts of the backup path: | Control | Responsibility | | --- | --- | | Kafka TLS or mTLS | Encrypts traffic between Kafka and the backup process | | Storage-provider encryption | Encrypts complete backup objects at rest using S3 SSE/KMS, Azure CMK, GCS CMEK, or an equivalent bucket policy | | Confluent CSFLE | Encrypts selected message fields before records are written to Kafka | | Enterprise CSFLE metadata backup | Preserves the KEK, DEK, subject, and schema-rule metadata for an existing CSFLE deployment | Configure storage encryption in the cloud provider or bucket policy. The OSS `backup` configuration and Strimzi `KafkaBackup` CR do not provide a client-side segment-encryption setting. ``` Encryption backup requires enterprise.schema_registry to be configured ``` Add the Schema Registry URL and authentication settings to the same `enterprise` section. The DEK Registry is a sub-API of Schema Registry. ``` Encryption backup is configured but not licensed. Skipping encryption backup. ``` Check the active feature set: ``` kafka-backup license info ``` The output must include `encryption`, or an active trial must be in effect. The encryption client inherits Schema Registry authentication and TLS settings. Verify those credentials first. If the DEK Registry uses a separate endpoint or identity, configure `dek_registry_url` and the `auth` or `tls` override under `enterprise.encryption`. Confirm that the subjects contain Confluent Schema Registry `ENCRYPT` rules. The backup can still contain KEK and DEK Registry exports when no matching encrypted subjects are found. - [Enterprise installation](https://kafkabackup.com/enterprise/installation.md) - [Schema Registry backup](https://kafkabackup.com/enterprise/schema-registry.md) - [Licensing and feature flags](https://kafkabackup.com/enterprise/licensing.md) - [Security setup](https://kafkabackup.com/guides/security-setup.md) --- title: GDPR Erasure (restore-time key suppression) description: Restore-time key suppression — re-apply erasure requests when restoring a backup (Enterprise 0.4.0+) source_url: html: https://kafkabackup.com/enterprise/erasure md: https://kafkabackup.com/enterprise/erasure.md --- # GDPR Erasure (restore-time key suppression) **Enterprise 0.4.0+ · licence feature `erasure` · included in the 14-day built-in trial** A backup taken before an erasure request still holds the erased records. This feature re-applies the erasure while restoring: every record whose key is on your **erasure register** is produced as a tombstone or skipped, and the restore output records which register was applied (SHA-256) and how many records it affected. It is the "erasure on restore" half of the posture described in [Retention and Erasure](https://kafkabackup.com/guides/retention-and-erasure.md). restore.yaml ``` mode: restore backup_id: gdpr-demo target: bootstrap_servers: ["kafka:9092"] topics: include: ["customers"] storage: backend: s3 bucket: kafka-backups enterprise: erasure: suppression: keys_file: /etc/kafka-backup/erasure/suppressed-keys.txt topics: ["customers", "customer-*"] # optional; source-topic globs, default: all match: exact # exact | prefix | regex on_match: tombstone # tombstone | drop ``` | Key | Default | Description | | --- | --- | --- | | `keys_file` | required | An **absolute path** is read from the local filesystem; anything else is a key inside the configured backup storage (so the register can live next to the backups). | | `topics` | all | Source-topic names or globs the register applies to. | | `match` | `exact` | `exact` byte-for-byte; `prefix` — the key starts with the entry; `regex` — the entry is a regular expression evaluated over the raw key bytes. | | `on_match` | `tombstone` | What happens to a matching record (below). | One entry per line, UTF-8: ``` # Erasure register — comments and blank lines are ignored U2 # plain key base64:AAEC/w== # binary key ordersU9 # entry scoped to one source topic ``` Records with a **null key never match**. A malformed line (bad `base64:`, invalid regex) fails the restore and names the line. - **`tombstone`** (default) — the record is produced with its key and a null value. Use it for **compacted** topics: compaction retires every earlier copy of the key, and source→target offsets stay aligned one-to-one. - **`drop`** — the record is not produced. Use it for **non-compacted** topics where the record must not exist on the target. Consumer-group offset mapping stays exact: each dropped source offset maps to the next surviving record's target offset. Pairing them the other way round leaves data behind: `drop` on a compacted topic keeps older values of the key; `tombstone` on a plain topic leaves the earlier records intact. ``` Records suppressed: 2 (0 dropped, 2 tombstoned; keys_file=/etc/kafka-backup/erasure/suppressed-keys.txt, entries=1, sha256=d4f0f9e4…a9fa) ``` `three-phase-restore` prints the same line in its Phase 2 block. `validate-restore` prints an `Erasure Suppression (enterprise)` block (keys file, entries, SHA-256, match, on\_match, topics) and, with `--format json`, adds it under `enterprise.erasure_suppression`. The dry run loads and reports the register but does not evaluate records; counts appear on the real restore. The one thing this feature must never do is silently restore erased data. So: - If `enterprise.erasure.suppression` is configured and the `erasure` feature is **not licensed**, the command is refused: `Feature 'erasure' is not licensed … refused rather than run without it.` - If the keys file is **missing, unreadable or malformed**, the command is refused _before any engine is constructed_ — no topic is created, nothing is restored. `validate-restore` fails the same way, so a broken register is caught in the dry run. - The register's digest and the match configuration are folded into the restore checkpoint: resuming an interrupted restore with a changed register warns and restarts rather than mixing lists. The Enterprise binary ships a 14-day auto-trial that unlocks every feature, including `erasure` — no licence file needed. The demo [`cli/gdpr-erasure`](https://github.com/osodevops/kafka-backup-demos/tree/main/cli/gdpr-erasure) runs nine asserted proof points against Kafka + MinIO: backup before the request, tombstone restore, drop restore, the missing-register refusal, the licence-gate refusal (`KAFKA_BACKUP_NO_TRIAL=1`) and the validate-restore summary. In-place redaction of existing backup sets and crypto-shredding via per-set keys follow, designed so a set is declared _immutable_ (cyber-recovery) or _erasable_, never both — see the [enterprise erasure issue](https://github.com/osodevops/kafka-backup-enterprise/issues/16). --- title: Audit Logging description: Comprehensive audit trails for OSO Kafka Backup Enterprise source_url: html: https://kafkabackup.com/enterprise/audit-logging md: https://kafkabackup.com/enterprise/audit-logging.md --- # Audit Logging OSO Kafka Backup Enterprise provides comprehensive audit logging for compliance and security monitoring. Audit logging captures: - All backup operations - All restore operations - Configuration changes - Access attempts (successful and denied) - Administrative actions ``` enterprise: audit: enabled: true # Where to send audit logs destination: type: file path: /var/log/kafka-backup/audit.log # What to log events: - backup.started - backup.completed - backup.failed - restore.started - restore.completed - restore.failed - access.denied - config.changed ``` ``` enterprise: audit: destination: type: file path: /var/log/kafka-backup/audit.log rotation: max_size_mb: 100 max_files: 10 compress: true ``` ``` enterprise: audit: destination: type: s3 bucket: audit-logs prefix: kafka-backup/ region: us-west-2 # Logs are batched and uploaded periodically batch_interval_secs: 60 ``` ``` enterprise: audit: destination: type: cloudwatch log_group: /kafka-backup/audit log_stream: ${HOSTNAME} region: us-west-2 ``` ``` enterprise: audit: destination: type: kafka bootstrap_servers: - kafka:9092 topic: audit-events security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA256 sasl_username: audit-producer sasl_password: ${AUDIT_KAFKA_PASSWORD} ``` ``` enterprise: audit: destination: type: syslog server: syslog.company.com:514 protocol: tcp # or udp facility: local0 ``` ``` enterprise: audit: destination: type: webhook url: https://siem.company.com/ingest headers: Authorization: Bearer ${WEBHOOK_TOKEN} batch_size: 100 retry_attempts: 3 ``` ``` enterprise: audit: destinations: - type: file path: /var/log/kafka-backup/audit.log - type: s3 bucket: audit-archive prefix: kafka-backup/ - type: kafka bootstrap_servers: - kafka:9092 topic: audit-events ``` | Category | Events | | --- | --- | | **Backup** | started, completed, failed, cancelled | | **Restore** | started, completed, failed, cancelled | | **Validate** | started, completed, failed | | **Access** | granted, denied | | **Config** | changed, viewed | | **Auth** | login, logout, token\_refresh | | **Admin** | license\_applied, feature\_enabled | ``` enterprise: audit: events: # Backup events - backup.started - backup.completed - backup.failed - backup.checkpoint # Include progress checkpoints # Restore events - restore.started - restore.completed - restore.failed # Security events (always recommended) - access.denied - auth.login - auth.logout - auth.failed # Configuration events - config.changed # Exclude specific events exclude_events: - backup.checkpoint # If too verbose ``` Every audit event includes: ``` { "timestamp": "2024-12-01T10:15:30.123Z", "event_type": "backup.completed", "event_id": "evt_abc123def456", "version": "1.0", "actor": { "type": "user", "id": "alice@company.com", "roles": ["backup-operator"], "ip_address": "10.0.0.50", "user_agent": "kafka-backup-cli/1.0.0" }, "resource": { "type": "backup", "id": "production-backup-20241201", "path": "s3://kafka-backups/production/" }, "action": { "operation": "backup", "result": "success", "duration_ms": 45000 }, "context": { "cluster_id": "prod-us-west-2", "environment": "production", "correlation_id": "req_xyz789" } } ``` ``` { "event_type": "backup.completed", "details": { "backup_id": "production-backup-20241201", "topics": ["orders", "payments", "users"], "records_backed_up": 1500000, "bytes_backed_up": 524288000, "compressed_bytes": 104857600, "compression_ratio": 5.0, "duration_secs": 120, "checkpoints": 4 } } ``` ``` { "event_type": "restore.completed", "details": { "backup_id": "production-backup-20241201", "target_cluster": "dr-us-east-1", "topics_restored": ["orders", "payments"], "records_restored": 1000000, "pitr_enabled": true, "time_window_start": "2024-12-01T00:00:00Z", "time_window_end": "2024-12-01T12:00:00Z", "topic_mapping": { "orders": "restored-orders" } } } ``` ``` { "event_type": "access.denied", "details": { "operation": "restore", "resource": "topic:production-orders", "required_permission": "restore:topic:production-orders", "user_permissions": ["backup:*", "list:*"], "reason": "insufficient_permissions" } } ``` ``` { "event_type": "config.changed", "details": { "config_type": "backup", "changes": [ { "field": "compression_level", "old_value": "3", "new_value": "6" }, { "field": "topics.include", "old_value": ["orders"], "new_value": ["orders", "payments"] } ] } } ``` Ensure logs cannot be tampered with: ``` enterprise: audit: destination: type: s3 bucket: compliance-audit-logs object_lock: true # Use S3 Object Lock integrity: signing: enabled: true algorithm: HMAC-SHA256 key_env_var: AUDIT_SIGNING_KEY chain: enabled: true # Chain signatures for tamper detection ``` ``` { "event": { /* normal event */ }, "integrity": { "signature": "sha256:abc123...", "previous_signature": "sha256:xyz789...", "sequence_number": 12345 } } ``` ``` enterprise: audit: retention: min_days: 365 # Minimum retention (compliance) max_days: 2555 # Maximum retention (7 years) # S3 lifecycle lifecycle: - days: 90 storage_class: STANDARD_IA - days: 365 storage_class: GLACIER ``` All logs are JSON for easy parsing: ``` # Find all failed backups cat audit.log | jq 'select(.event_type == "backup.failed")' # Find access denied for user cat audit.log | jq 'select(.event_type == "access.denied" and .actor.id == "bob@company.com")' # Count events by type cat audit.log | jq '.event_type' | sort | uniq -c ``` ``` -- Failed operations in last 24 hours fields @timestamp, event_type, actor.id, details.error | filter event_type like /failed/ | sort @timestamp desc | limit 100 -- Access denied events fields @timestamp, actor.id, details.operation, details.resource | filter event_type = "access.denied" | sort @timestamp desc ``` ``` { "query": { "bool": { "must": [ { "match": { "event_type": "restore.completed" } }, { "range": { "timestamp": { "gte": "now-7d" } } } ] } }, "aggs": { "by_user": { "terms": { "field": "actor.id.keyword" } } } } ``` ``` enterprise: audit: alerts: - name: backup-failure condition: "event_type == 'backup.failed'" severity: high notify: - type: slack webhook: ${SLACK_WEBHOOK} - type: pagerduty routing_key: ${PAGERDUTY_KEY} - name: access-denied-spike condition: "event_type == 'access.denied'" threshold: 10 window: 5m severity: medium notify: - type: email to: security@company.com - name: unauthorized-restore condition: "event_type == 'restore.started' and context.environment == 'production'" severity: high notify: - type: slack channel: "#security-alerts" ``` ``` notify: - type: slack webhook: ${SLACK_WEBHOOK} channel: "#kafka-alerts" template: | :warning: *{{ .event_type }}* User: {{ .actor.id }} Resource: {{ .resource.id }} Time: {{ .timestamp }} ``` ``` notify: - type: pagerduty routing_key: ${PAGERDUTY_ROUTING_KEY} severity: "{{ .severity }}" ``` ``` notify: - type: email smtp_server: smtp.company.com:587 from: kafka-backup@company.com to: - ops@company.com - security@company.com ``` ``` apiVersion: v1 kind: ConfigMap metadata: name: audit-config namespace: kafka-backup data: audit.yaml: | enterprise: audit: enabled: true destination: type: s3 bucket: audit-logs events: - backup.* - restore.* - access.denied ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: audited-backup spec: enterprise: audit: enabled: true destination: type: file path: /var/log/audit/kafka-backup.log # Mount persistent volume for audit logs volumes: - name: audit-logs persistentVolumeClaim: claimName: audit-logs-pvc volumeMounts: - name: audit-logs mountPath: /var/log/audit ``` 1. **Log to multiple destinations** - Redundancy for compliance 2. **Enable immutable storage** - Prevent tampering 3. **Set appropriate retention** - Meet compliance requirements 4. **Alert on security events** - Real-time monitoring 5. **Regular log review** - Periodic audit reviews 6. **Secure log access** - Limit who can read logs 7. **Include correlation IDs** - Trace requests across systems Generate compliance reports from audit logs: ``` #!/bin/bash # Monthly compliance report START_DATE=$(date -d "last month" +%Y-%m-01) END_DATE=$(date -d "this month" +%Y-%m-01) echo "=== Kafka Backup Audit Report ===" echo "Period: $START_DATE to $END_DATE" echo "" echo "Backup Operations:" cat audit.log | jq -r "select(.timestamp >= \"$START_DATE\" and .timestamp < \"$END_DATE\" and .event_type | startswith(\"backup\")) | .event_type" | sort | uniq -c echo "" echo "Restore Operations:" cat audit.log | jq -r "select(.timestamp >= \"$START_DATE\" and .timestamp < \"$END_DATE\" and .event_type | startswith(\"restore\")) | .event_type" | sort | uniq -c echo "" echo "Access Denied Events:" cat audit.log | jq -r "select(.timestamp >= \"$START_DATE\" and .timestamp < \"$END_DATE\" and .event_type == \"access.denied\") | \"\(.actor.id) - \(.details.operation) - \(.details.resource)\"" ``` - [RBAC Configuration](https://kafkabackup.com/enterprise/rbac.md) - Access control - [Compliance Guide](https://kafkabackup.com/use-cases/compliance-audit.md) - Meeting regulations - [Security Setup](https://kafkabackup.com/guides/security-setup.md) - Secure configuration --- title: MSK ZooKeeper to KRaft Migration description: Migrate AWS MSK clusters from ZooKeeper to KRaft with a controlled cutover, offset continuity, and cryptographic evidence. AWS-tested enterprise migration pipeline. source_url: html: https://kafkabackup.com/enterprise/msk-kraft-migration md: https://kafkabackup.com/enterprise/msk-kraft-migration.md --- # MSK ZooKeeper to KRaft Migration > [!TIP] > > [!NOTE] > > Try it free — no license needed > > [!NOTE] > > `plan` and `precheck` are completely free. Run them against your production clusters today to see exactly what a migration looks like — generated runbook, cost estimate, IAM policies, and infrastructure readiness report. No signup, no trial activation. > > > > ``` > > kafka-backup migrate msk-kraft plan --config migration.yaml --format all --out-dir ./migration-plan > > kafka-backup migrate msk-kraft precheck --config migration.yaml > > ``` Migrate your AWS MSK clusters from ZooKeeper to KRaft mode with a **short coordinated producer freeze**, **validated offset continuity**, and a **cryptographically signed evidence bundle** that proves the migration succeeded. Consumers resume from translated target offsets so message continuity is preserved across the switch. The `0.3.2` enterprise release was production-qualified on May 8, 2026 from the public Homebrew tap against AWS MSK Provisioned clusters in `eu-west-2`. Apache Kafka 4.0 removes ZooKeeper entirely. KRaft (Kafka Raft) replaces ZooKeeper as the metadata management layer, bringing: - **Faster controller failover** — seconds instead of minutes - **Simplified operations** — one system to manage instead of two - **Better scalability** — millions of partitions per cluster - **Reduced infrastructure** — no ZooKeeper ensemble to provision, monitor, or patch AWS MSK supports KRaft from version 3.7.x onward. Apache Kafka 4.x removes ZooKeeper upstream, so ZooKeeper-mode MSK clusters need a migration plan before a Kafka 4.x upgrade. Yes. ZooKeeper was deprecated in Apache Kafka 3.5 (KIP-833) and removed in Kafka 4.0. AWS MSK's latest versions already support KRaft, and new clusters should be provisioned in KRaft mode. AWS MSK does not support in-place ZooKeeper-to-KRaft conversion. You must create a **new KRaft cluster** and move everything over: | What needs to migrate | What happens without tooling | | --- | --- | | Topic data (every partition, every record) | Manual MirrorMaker setup, ongoing maintenance | | Topic configurations (retention, compaction, replication) | Manual recreation, error-prone | | Consumer group offsets | **Lost** — consumers restart from earliest or latest | | ACL bindings | Manual recreation, security gaps during transition | | Proof that migration succeeded | Nothing — hope and prayer | The gap between "data is on the new cluster" and "consumers resume from the right place" is where migrations fail. A single incorrect offset means lost messages or reprocessed duplicates — silent data corruption that surfaces days later in downstream systems. | Capability | kafka-backup Enterprise | MirrorMaker 2 | Manual | | --- | --- | --- | --- | | Controlled cutover | Yes (coordinated producer freeze) | Partial | No | | Offset continuity (exact message resume) | Yes (offset-map translation) | No | No | | ACL migration with drift handling | Yes (merge/replace/refuse) | No | Manual | | Topic config preservation | Yes (automatic) | Partial | Manual | | Cryptographic evidence bundle | Yes (Ed25519-signed) | No | No | | 5-check automated validation incl. target offset-floor guard | Yes | No | No | | Rollback capability | Yes (pre-cutover) | No | No | | Resume after failure | Yes (journal-based) | Restart from scratch | Restart from scratch | | Cross-auth support (SCRAM → IAM) | Implemented | No | Manual | The migration runs through a deterministic journaled state machine. Every state transition is included in the final evidence bundle. ``` PLANNED → PRECHECK → TOPOLOGY_COPY → SEED → TAIL → DRAIN_READY ↓ FINALIZED ← VALIDATING ← AWAITING_CLIENT_SWITCH ← CUTOVER ``` | Phase | State | What happens | | --- | --- | --- | | Plan & Precheck | `planned` → `precheck` | Read-only analysis of both clusters. Detects blockers (incompatible versions, unreachable brokers, S3 permission issues) and warnings (cross-region egress, compacted topics, static members). | | Topology Copy | `topology_copy` | Creates missing topics on target with matching partition counts and configurations. Copies ACL bindings (filtering MSK internals like `User:ANONYMOUS`). | | Seed | `seed` | Bulk-copies all existing data through S3 — source → backup → S3 → restore → target. Builds the **offset map** that enables consumer group translation. | | Tail | `tail` | Continuously bridges the gap between seed and cutover. Replays new records as they arrive on source. Tracks per-partition lag. | | Drain Ready | `drain_ready` | All partitions within lag tolerance. Execution halts. Operator decides when to proceed. | | Cutover | `cutover` | Freezes producers (via webhook or manual), publishes sentinel records, drains final records, translates all consumer group offsets, commits translated offsets on target. | | Client Switch | `awaiting_client_switch` | Operator updates application configs to point to new KRaft cluster bootstrap servers. | | Validation | `validating` | Runs 5 automated checks: topic parity, counts and offsets with target offset-floor guard, spot-check record equality, sentinel presence, consumer group reconciliation. | | Finalize | `finalized` | Signs the evidence bundle with Ed25519 and uploads to S3. Migration complete. | At any point before cutover, you can **rollback** — the source cluster is never modified. kafka-backup implements every MSK authentication mode and cross-auth migration: > [!NOTE] > > [!NOTE] > > Release coverage > > [!NOTE] > > The May 8, 2026 AWS end-to-end release qualification proved SCRAM-SHA-512 source to SCRAM-SHA-512 target on MSK across four target KRaft clusters. IAM and cross-auth paths are implemented and covered by endpoint-selection, token, smoke, and access-map tests; rehearse them in staging before treating them as equivalent production evidence for your environment. | Source Auth | Target Auth | Implemented | AWS E2E coverage | Notes | | --- | --- | --- | --- | --- | | IAM | IAM | Yes | Unit/smoke | Most common MSK configuration | | SCRAM-SHA-512 | SCRAM-SHA-512 | Yes | May 8, 2026 full AWS E2E | Pre-provision SCRAM users on target | | SCRAM-SHA-512 | IAM | Yes | Unit/smoke | Auth modernization — ACLs emitted as `access-map.json` | | IAM | SCRAM-SHA-512 | Yes | Unit/smoke | | | mTLS | IAM | Yes | Unit/smoke | | | mTLS | mTLS | Yes | Unit/smoke | | | PLAINTEXT | Any | Yes | Local/dev E2E | Dev/test environments | Cross-auth migration (e.g., SCRAM source → IAM target) is a first-class feature. When the target uses IAM, Kafka ACLs don't apply — instead, the tool generates an `access-map.json` that maps each principal's permissions to the IAM policies you need to create. Before finalizing, the tool runs five independent validation checks: | Check | What it verifies | Pass criteria | | --- | --- | --- | | **Topic Parity** | Partition counts match between source and target | All topics match | | **Counts & Offsets** | Record counts within tolerance (default ±1 for sentinel) and target offset-floor safety | Per-partition span difference ≤ `count_tolerance`; `offset_floor_violations=0` | | **Spot-Check Records** | Sampled records byte-equal between source and target | All samples match (compacted topics allow warnings) | | **Sentinel Presence** | Cutover marker records landed on target | All sentinels found | | **Consumer Group Reconciliation** | Translated offsets committed correctly on target | All offsets match expected values | The overall outcome is **PASSED**, **WARNING**, or **FAILED**. A WARNING can be expected for compacted-topic drift or empty/zero-span partitions where no spot-check record exists; it is not a data mismatch when every comparable sample matched. Failed validation blocks finalization — you must investigate and remediate before the migration can complete. > [!NOTE] > > [!NOTE] > > Latest AWS release evidence > > [!NOTE] > > The `0.3.2` release binary was installed from the public Homebrew tap and exercised end-to-end in AWS on May 8, 2026 across four target KRaft clusters. All four migrations finalized with validation `PASSED`. The final target run, `release-v032-20260508T070819Z-r4`, validated 11 migrated topics, 68 partitions, 125,021 seeded records, 1,222 tail records, 3 consumer groups, 37 translated offset commits, 68 sentinels, and 204/204 spot-check records matched. Every migration produces an Ed25519-signed JSON evidence bundle uploaded to S3. This is your **auditable proof** that the migration succeeded: - Complete state transition journal (every phase, with timestamps) - Source and target cluster metadata snapshots - Topology diff (topics created, configs applied) - ACL plan (bindings copied, internals filtered) - Seed and tail statistics (records, bytes, partitions) - Full validation report with per-partition detail - Offset translation map - Cutover report (sentinel positions, freeze timing) The signature is verifiable offline with the Ed25519 public key. For regulated environments, the evidence bucket supports S3 Object Lock (COMPLIANCE mode) to prevent tampering. > [!NOTE] > > [!NOTE] > > For compliance teams > > [!NOTE] > > The evidence bundle answers: "Prove that every record made it to the new cluster and every consumer will resume from the right place." It's the difference between "we think it worked" and "here's the cryptographic proof." migration.yaml ``` enterprise: msk_kraft_migration: source: cluster_arn: arn:aws:kafka:us-east-1:123456789012:cluster/my-zk-cluster/abc-123 auth: mode: iam target: cluster_arn: arn:aws:kafka:us-east-1:123456789012:cluster/my-kraft-cluster/def-456 auth: mode: iam backup: s3_bucket: my-migration-segments s3_prefix: migrations/ evidence: s3_bucket: my-migration-evidence s3_prefix: evidence/ ``` ``` kafka-backup migrate msk-kraft plan \ --config migration.yaml \ --format all \ --out-dir ./migration-plan ``` This generates: - `plan.json` — machine-readable migration plan - `runbook.md` — step-by-step operator runbook - `aws-cli.sh` — AWS CLI commands for infrastructure setup - `iam-policy-templated.json` — IAM policy template - `iam-policy-concrete.json` — IAM policy with your ARNs filled in - `cost-estimate.json` — estimated S3 and data transfer costs ``` kafka-backup migrate msk-kraft precheck --config migration.yaml ``` Precheck analyzes both clusters and reports blockers, warnings, and informational findings. See the [Precheck Codes Reference](https://kafkabackup.com/enterprise/msk-kraft-precheck-codes.md) for remediation guidance. ``` kafka-backup migrate msk-kraft execute \ --config migration.yaml \ --journal-dir ./journal ``` See the [Production Migration Runbook](https://kafkabackup.com/guides/msk-kraft-migration-runbook.md) for the complete step-by-step process. MSK KRaft migration requires the `migrations:msk-kraft` feature in your enterprise license. A **14-day free trial** activates automatically on first run — no signup, no credit card. - `plan` and `precheck` are always free, even without a license - `execute`, `cutover`, `finalize`, and other mutation commands require an active license - Licenses are Ed25519-signed files validated offline — no license server, no phone-home [Learn more about licensing](https://kafkabackup.com/enterprise/licensing.md) | [Get a license](https://enterprise.kafkabackup.com) No. Apache Kafka 3.3+ supports KRaft mode (ZooKeeper-free). Kafka 4.0 removes ZooKeeper entirely. AWS MSK supports KRaft from version 3.7.x. Yes. KRaft mode replaces ZooKeeper with an internal Raft-based metadata quorum. New clusters should be provisioned in KRaft mode. ZooKeeper was deprecated in Kafka 3.5 and removed in Kafka 4.0. Existing ZooKeeper-mode clusters must migrate to KRaft before upgrading to Kafka 4.x. Yes. KRaft has been production-ready since Kafka 3.3 (KIP-833). AWS MSK supports KRaft in production from version 3.7.x. Major organizations have been running KRaft in production since 2024. KRaft (Kafka Raft) is the consensus protocol that replaces ZooKeeper for Kafka metadata management. It uses the Raft algorithm to elect a controller and replicate metadata across the cluster, eliminating the need for a separate ZooKeeper ensemble. See [KRaft Architecture](https://kafkabackup.com/architecture/msk-kraft-migration.md) for a deep dive. Migration time depends on data volume and network bandwidth. Rough estimates for a 3-broker cluster: | Data volume | Seed phase | Total (including tail + cutover) | | --- | --- | --- | | 10 GB | ~5 minutes | ~10 minutes | | 100 GB | ~30 minutes | ~45 minutes | | 1 TB | ~4 hours | ~5 hours | | 10 TB | ~36 hours | ~40 hours | The producer freeze window during cutover is typically under 60 seconds regardless of data volume. - [Production Migration Runbook](https://kafkabackup.com/guides/msk-kraft-migration-runbook.md) — step-by-step guide - [Configuration Reference](https://kafkabackup.com/enterprise/msk-kraft-config-reference.md) — every YAML field explained - [Precheck Codes Reference](https://kafkabackup.com/enterprise/msk-kraft-precheck-codes.md) — blocker and warning remediation - [CLI Reference](https://kafkabackup.com/enterprise/msk-kraft-cli-reference.md) — all 9 migration commands - [Architecture Deep Dive](https://kafkabackup.com/architecture/msk-kraft-migration.md) — how offset continuity works - [IAM-to-IAM Example](https://kafkabackup.com/examples/msk-kraft-migration-iam.md) — complete worked example - [Cross-Auth Example](https://kafkabackup.com/examples/msk-kraft-migration-cross-auth.md) — SCRAM to IAM migration --- title: MSK KRaft Migration Configuration Reference description: Complete YAML configuration reference for kafka-backup Enterprise MSK ZooKeeper to KRaft migration. source_url: html: https://kafkabackup.com/enterprise/msk-kraft-config-reference md: https://kafkabackup.com/enterprise/msk-kraft-config-reference.md --- # MSK KRaft Migration Configuration Reference All migration settings live under `enterprise.msk_kraft_migration` in your config YAML. The smallest valid config requires source, target, backup, and evidence: migration-minimal.yaml ``` enterprise: msk_kraft_migration: source: cluster_arn: arn:aws:kafka:us-east-1:123456789012:cluster/source-zk/abc-123 auth: mode: iam target: cluster_arn: arn:aws:kafka:us-east-1:123456789012:cluster/target-kraft/def-456 auth: mode: iam backup: s3_bucket: my-migration-segments s3_prefix: replay/ evidence: s3_bucket: my-migration-evidence s3_prefix: migrations/ ``` All other sections (`cutover`, `validation`, `seed`, `topology`, `acl`) have sensible defaults. migration-production.yaml ``` enterprise: msk_kraft_migration: enabled: true # default: true source: cluster_arn: arn:aws:kafka:us-east-1:123456789012:cluster/prod-zk/abc-123 auth: mode: iam target: cluster_arn: arn:aws:kafka:us-east-1:123456789012:cluster/prod-kraft/def-456 auth: mode: iam backup: s3_bucket: prod-migration-segments s3_prefix: zk-to-kraft/ kms_key_arn: arn:aws:kms:us-east-1:123456789012:key/mrk-abc123 # optional evidence: s3_bucket: prod-migration-evidence s3_prefix: migrations/ retention: 7y # S3 Object Lock retention signing_key_path: /etc/kafka-backup/evidence-signing.key # optional cutover: drain_timeout: 30m # max time waiting for lag to converge drain_max_partition_lag: 100 # records; all partitions must be below this drain_stable_window: 30s # lag must stay below threshold for this long max_producer_freeze: 60s # max time producers are frozen during cutover producer_freeze_webhook: https://api.example.com/kafka/freeze reverse_replication_enabled: false # not yet implemented validation: count_tolerance: 1 # ±1 record allowed (sentinel offset shift) spot_check_records_per_partition: 3 # records sampled per partition seed: max_concurrent_partitions: 4 # parallel partition transfers segment_max_bytes: 33554432 # 32 MB per S3 segment topology: on_config_drift: overwrite_with_source # overwrite_with_source | keep_target | refuse acl: on_drift: merge # merge | replace | refuse ``` * * * Each cluster reference describes how to connect and authenticate. | Field | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `cluster_arn` | string | Yes | — | AWS MSK cluster ARN. Region and account are extracted from this. | | `auth` | object | Yes | — | Authentication configuration (see below). | | `local_dev` | object | No | — | Dev-only override for Docker/local testing. Bypasses MSK API calls. | Authentication is configured with `mode` as a discriminator: ``` auth: mode: iam ``` Uses AWS IAM SASL/OAUTHBEARER. Credentials come from the standard AWS SDK chain (environment, instance profile, SSO, etc.). ``` auth: mode: scram-sha-512 username: admin password: ${KAFKA_PASSWORD} ``` | Field | Type | Required | Description | | --- | --- | --- | --- | | `username` | string | Yes | SASL username | | `password` | string | Yes | SASL password (supports `${ENV_VAR}` interpolation) | ``` auth: mode: mtls keystore: /path/to/client-keystore.jks keystore_password: ${KEYSTORE_PASS} truststore: /path/to/truststore.jks truststore_password: ${TRUSTSTORE_PASS} ``` | Field | Type | Required | Description | | --- | --- | --- | --- | | `keystore` | string | Yes | Path to JKS keystore | | `keystore_password` | string | Yes | Keystore password | | `truststore` | string | Yes | Path to JKS truststore | | `truststore_password` | string | Yes | Truststore password | ``` auth: mode: plain ``` PLAINTEXT connection with no authentication. For dev/test only. Bypasses MSK `DescribeCluster` and `GetBootstrapBrokers` API calls. Used for Docker-based testing. ``` local_dev: bootstrap: "localhost:9092,localhost:9093,localhost:9094" metadata_mode: "ZOOKEEPER" # or "KRAFT" kafka_version: "3.7.1" # default: 3.7.1 broker_count: 3 # default: 3 region: "local" # default: local disable_tls: false # default: false; set true for SASL_PLAINTEXT in Docker ``` | Field | Type | Default | Description | | --- | --- | --- | --- | | `bootstrap` | string | — | Comma-separated `host:port` list (required) | | `metadata_mode` | string | — | `"ZOOKEEPER"` or `"KRAFT"` (required) | | `kafka_version` | string | `"3.7.1"` | Kafka version string | | `broker_count` | integer | `3` | Number of brokers | | `region` | string | `"local"` | AWS region (used for S3 client) | | `disable_tls` | boolean | `false` | Downgrade SASL\_SSL to SASL\_PLAINTEXT | * * * S3 bucket used as the intermediate replication channel during seed and tail phases. | Field | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `s3_bucket` | string | Yes | — | S3 bucket name | | `s3_prefix` | string | Yes | — | S3 key prefix for migration data | | `kms_key_arn` | string | No | — | AWS KMS key ARN for server-side encryption (SSE-KMS) | * * * S3 bucket for the signed migration evidence bundle. | Field | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `s3_bucket` | string | Yes | — | S3 bucket name | | `s3_prefix` | string | No | `"migrations/"` | S3 key prefix | | `retention` | string | No | `"7y"` | S3 Object Lock retention period (e.g., `"7y"`, `"90d"`) | | `signing_key_path` | string | No | — | Path to Ed25519 private key for evidence signing | > [!NOTE] > > [!NOTE] > > Object Lock > > [!NOTE] > > If the evidence bucket has S3 Object Lock enabled, the evidence bundle is uploaded in COMPLIANCE mode with the configured retention. If Object Lock is not configured, the bundle is uploaded without retention and a warning is logged. * * * Controls the cutover phase behavior — drain convergence, producer freeze, and sentinel timing. | Field | Type | Default | Description | | --- | --- | --- | --- | | `drain_timeout` | duration | `10m` | Maximum time to wait for tail lag to converge before giving up. | | `drain_max_partition_lag` | integer | `1000` | Maximum per-partition lag (records) to consider "caught up". | | `drain_stable_window` | duration | `60s` | Lag must stay below threshold for this duration before drain is declared ready. | | `max_producer_freeze` | duration | `60s` | Maximum time producers remain frozen during cutover. Cutover aborts if exceeded. | | `producer_freeze_webhook` | string | — | URL to POST for producer freeze/unfreeze. If omitted, the tool prompts on TTY. | | `reverse_replication_grace` | duration | `1h` | Grace period for reverse replication (reserved for future use). | | `reverse_replication_enabled` | boolean | `false` | Enable reverse replication after cutover. **Not yet implemented** — setting to `true` is a blocker (B13). | Duration values use human-readable format: `30s`, `5m`, `1h`, `2h30m`. When configured, the tool sends HTTP POST requests: ``` POST Content-Type: application/json {"action": "freeze", "migration_id": "", "reason": "cutover"} ``` The webhook must return HTTP 2xx within `max_producer_freeze`. On failure or timeout, cutover aborts and the migration enters `failed` state. * * * Controls the 5-check validation suite that runs during finalize. The counts-and-offsets check also verifies the target offset floor so restored records have not been truncated before client switch. | Field | Type | Default | Description | | --- | --- | --- | --- | | `count_tolerance` | integer | `1` | Absolute difference in record count allowed per partition. Set to 1 to account for the sentinel record. | | `spot_check_records_per_partition` | integer | `3` | Number of records sampled per partition for byte-equality check. Minimum 1. | * * * Controls the bulk data transfer (seed) phase. | Field | Type | Default | Description | | --- | --- | --- | --- | | `max_concurrent_partitions` | integer | `4` | Number of partitions transferred in parallel. Higher values increase source cluster load. | | `segment_max_bytes` | integer | `33554432` (32 MB, migration profile; the core default is 128 MB) | Maximum S3 segment size in bytes. Larger segments reduce S3 API calls but use more memory. | > [!TIP] > > [!NOTE] > > Tuning for large clusters > > [!NOTE] > > For clusters with 500+ partitions, consider increasing `max_concurrent_partitions` to 8-16 to reduce seed phase duration. Monitor source cluster CPU and network during seed. * * * Controls how topic configuration drift is handled when the target already has pre-existing topics. | Field | Type | Default | Description | | --- | --- | --- | --- | | `on_config_drift` | enum | `overwrite_with_source` | How to handle config differences between source and target topics. | | Policy | Behavior | | --- | --- | | `overwrite_with_source` | Push source config values to target. Source keys absent on target are added. **Default.** | | `keep_target` | Keep pre-existing target values. Source keys absent on target are still added. | | `refuse` | Refuse to proceed if any config drift exists. Safest for regulated workloads. | * * * Controls how ACL binding drift is handled between source and target. | Field | Type | Default | Description | | --- | --- | --- | --- | | `on_drift` | enum | `merge` | How to handle ACL differences between source and target. | | Policy | Behavior | | --- | --- | | `merge` | Create missing bindings on target. Leave extra target-only bindings in place. **Default.** | | `replace` | Create missing bindings. Surface extra target-only bindings in the report. Deletion is deferred to the operator. | | `refuse` | Refuse to proceed if any target-only ACL binding exists. | > [!WARNING] > > [!NOTE] > > MSK internal ACLs > > [!NOTE] > > ACL bindings for MSK internal principals (e.g., `User:ANONYMOUS`, `__consumer_offsets`) are automatically filtered and never copied to the target. This filtering is reported as warning W09 in precheck. * * * | Variable | Description | | --- | --- | | `AWS_ACCESS_KEY_ID` | AWS credentials (standard SDK chain) | | `AWS_SECRET_ACCESS_KEY` | AWS credentials (standard SDK chain) | | `AWS_SESSION_TOKEN` | AWS session token for temporary credentials | | `AWS_REGION` | Default AWS region (overridden by cluster ARN region) | | `AWS_ENDPOINT_URL_S3` | Custom S3 endpoint (for LocalStack or S3-compatible storage) | | `RUST_LOG` | Log level control (e.g., `RUST_LOG=debug`) | All credential fields in YAML support `${ENV_VAR}` interpolation. - [Production Migration Runbook](https://kafkabackup.com/guides/msk-kraft-migration-runbook.md) — step-by-step guide - [Precheck Codes Reference](https://kafkabackup.com/enterprise/msk-kraft-precheck-codes.md) — every finding explained - [CLI Reference](https://kafkabackup.com/enterprise/msk-kraft-cli-reference.md) — command flags and options --- title: MSK KRaft Migration Precheck Codes description: Complete reference for all kafka-backup MSK KRaft migration precheck findings — blockers, warnings, and info codes with remediation. source_url: html: https://kafkabackup.com/enterprise/msk-kraft-precheck-codes md: https://kafkabackup.com/enterprise/msk-kraft-precheck-codes.md --- # MSK KRaft Migration Precheck Codes The `precheck` command performs read-only analysis of both clusters and reports findings at three severity levels. **Precheck is free — no license required.** ``` kafka-backup migrate msk-kraft precheck --config migration.yaml ``` | Severity | Code prefix | Effect | | --- | --- | --- | | **Blocker** | B## | Migration cannot proceed. Must be resolved first. | | **Warning** | W## | Migration can proceed, but review the finding. | | **Info** | I## | Informational. No action required. | * * * **Message:** `"{which} MSK Serverless; this migrator only operates on MSK Provisioned clusters"` **Cause:** One or both cluster ARNs point to MSK Serverless clusters. MSK Serverless is KRaft-only by construction — there is no ZK variant to migrate from. **Fix:** Use MSK Provisioned cluster ARNs. The source must be ZK-mode Provisioned, the target must be KRaft-mode Provisioned. * * * **Message:** `"source cluster metadata mode is {mode}, expected ZOOKEEPER"` **Cause:** The source cluster is already running in KRaft mode. There is nothing to migrate. **Fix:** Verify the source ARN points to a ZooKeeper-mode cluster. * * * **Message:** `"target cluster metadata mode is {mode}, expected KRAFT"` **Cause:** The target cluster is not running in KRaft mode. **Fix:** Provision the target MSK cluster with KRaft mode enabled (requires Kafka 3.7+). * * * **Message:** `"source and target ARNs are identical"` **Cause:** Both `source.cluster_arn` and `target.cluster_arn` point to the same cluster. In-place migration is not supported. **Fix:** Create a separate KRaft-mode MSK cluster for the target. * * * **Message:** `"target Kafka version {version} is below minimum 3.7 required for KRaft on MSK"` **Cause:** The target cluster is running a Kafka version that does not support KRaft on MSK. **Fix:** Upgrade the target MSK cluster to Kafka 3.7.x or later. * * * **Message:** `"backup S3 bucket '{bucket}' not reachable: {error}"` **Cause:** The `HeadBucket` API call failed on the backup S3 bucket. Either the bucket does not exist or the caller lacks permissions. **Fix:** 1. Create the bucket: `aws s3 mb s3:// --region ` 2. Ensure the migration runner's IAM role has `s3:HeadBucket`, `s3:GetObject`, `s3:PutObject`, `s3:ListBucket` on this bucket * * * **Message:** `"evidence S3 bucket '{bucket}' not reachable: {error}"` **Cause:** Same as B07, but for the evidence bucket. **Fix:** Same as B07. If using S3 Object Lock, ensure the bucket was created with Object Lock enabled (cannot be added retroactively). * * * **Message:** `"source Kafka protocol not reachable: {error}"` **Cause:** Cannot connect to the source cluster's bootstrap servers or fetch metadata. **Fix:** - Verify bootstrap servers are correct (check MSK console or `aws kafka get-bootstrap-brokers`) - Check security group ingress rules — the migration runner must reach the broker ports - Verify auth mode matches the cluster's authentication configuration - For SCRAM: verify the username/password are correct and the SCRAM secret exists in AWS Secrets Manager * * * **Message:** `"target Kafka protocol not reachable: {error}"` **Cause:** Same as B09, but for the target cluster. **Fix:** Same as B09. * * * **Message:** `"target broker message.max.bytes={value} is below the largest source topic's effective max.message.bytes={max} (topic '{topic}') — replay would fail with RecordTooLargeException"` **Cause:** The target cluster's `message.max.bytes` broker setting is smaller than the largest message size allowed by any source topic. During restore, oversized records would be rejected. **Fix:** Raise the target broker's `message.max.bytes` to at least match the source floor. Update the MSK cluster configuration: ``` aws kafka update-cluster-configuration \ --cluster-arn \ --configuration-info '{"Arn":"","Revision":}' \ --current-version ``` * * * **Message:** `"target broker replica.fetch.max.bytes={value} is below the largest source topic's effective max.message.bytes={max} (topic '{topic}') — replication would stall on oversized batches"` **Cause:** The target's inter-broker replication cannot handle the largest messages from source. **Fix:** Raise the target broker's `replica.fetch.max.bytes` alongside `message.max.bytes`. * * * **Message:** `"cutover.reverse_replication_enabled=true, but reverse replication is not implemented"` **Cause:** The config enables a feature that is not yet available. **Fix:** Set `cutover.reverse_replication_enabled: false` in your config. Post-cutover rollback to the source cluster is a manual procedure. * * * **Message:** `"target has {target} brokers but source has {source} — consider scaling up before seed"` **Cause:** The target cluster has fewer brokers than the source. Topics with replication factor equal to source broker count may not be replicable. **Action:** Consider scaling up the target cluster before migration. * * * **Message:** `"source region {source} ≠ target region {target} — seed + tail will incur egress bandwidth cost"` **Cause:** Source and target are in different AWS regions. Data transfer between regions incurs egress charges. **Action:** Review the cost estimate from `plan --format cost`. Consider whether the data transfer cost is acceptable. * * * **Message:** `"KMS key ARN set on backup channel — CMK access is not verified by this precheck phase; ensure the caller has kms:Encrypt/Decrypt/GenerateDataKey"` **Cause:** A custom KMS key is configured for S3 encryption. Precheck does not verify KMS permissions. **Action:** Ensure the migration runner's IAM role has `kms:Encrypt`, `kms:Decrypt`, and `kms:GenerateDataKey` on the specified KMS key ARN. * * * **Message:** `"could not verify target message-size floor ({reason}) — ensure target message.max.bytes and replica.fetch.max.bytes ≥ largest source topic's effective max.message.bytes"` **Cause:** The `DescribeConfigs` API call failed for source or target brokers, but the brokers are reachable. This is a fail-open scenario. **Action:** Manually verify that the target's `message.max.bytes` and `replica.fetch.max.bytes` are sufficient. * * * **Message:** `"source cluster has static consumer-group members ({summary}). Post-cutover, these consumers MUST restart against the target with the same group.instance.id values..."` **Cause:** Some consumer groups use static membership (`group.instance.id`). These consumers must reconnect to the target with identical instance IDs to avoid a full group rebalance. **Action:** Ensure application deployments preserve `group.instance.id` values when switching to the target cluster. * * * **Message:** `"source cluster has {total} transactional producer(s)..."` **Cause:** Transactional state (producer ID + epoch) does not migrate. Exactly-once guarantees do not span the cutover boundary. **Action:** Applications using transactions must call `initTransactions()` after reconnecting to the target. Active transactions should be drained on source before pressing cutover. * * * **Message:** `"{count} source topic(s) use cleanup.policy=compact..."` **Cause:** Compacted topics may have records deleted between seed and tail phases. The validation suite treats empty fetches and drift on compacted topics as warnings instead of failures. **Action:** If bit-for-bit parity is required for compacted topics, run your own diff after finalize. * * * **Message:** `"target uses SCRAM-SHA-512 — SCRAM user credentials cannot be read via the Kafka protocol..."` **Cause:** The target cluster uses SCRAM authentication. SCRAM user credentials (stored in AWS Secrets Manager) cannot be read or copied programmatically. If the same users don't exist on the target, copied ACLs will reference unauthenticated principals. **Action:** Pre-provision all SCRAM users on the target cluster before cutover. Use `aws kafka batch-associate-scram-secret` to associate the same Secrets Manager secrets. * * * **Message:** `"{count} source ACL binding(s) reference MSK/Kafka internal principals or resources and will be filtered during ACL copy"` **Cause:** Some ACL bindings on the source reference internal principals (e.g., `User:ANONYMOUS`) or internal resources (`__consumer_offsets`). These are managed by MSK automatically and should not be copied. **Action:** No action needed. The filtered bindings are logged for transparency. * * * **Message:** `"{count} source topic(s) use finite delete retention ({topic}({retention_ms}ms)...). SEED restores original CreateTime timestamps, so the target broker may advance log-start before cutover if old restored records become retention-eligible. Temporarily extend topic retention for the migration window, or rely on the cutover offset-floor guard to block the client switch if truncation occurs."` **Cause:** One or more source topics use `cleanup.policy=delete` with finite `retention.ms`. Kafka retention uses the record timestamp when `message.timestamp.type=CreateTime`, so restored historical records may become immediately retention-eligible on the target. **Action:** Temporarily extend retention for affected topics during the migration window, or set retention to `-1` until finalize completes. Keep the target offset-floor guard enabled; it verifies target log-start offsets before `READY_FOR_CLIENT_SWITCH` and blocks the switch if the target has already truncated copied data. Example topic configuration that triggers this warning: ``` cleanup.policy=delete message.timestamp.type=CreateTime retention.ms=604800000 segment.ms=604800000 ``` * * * **Message:** `"target is IAM-auth — ACLs will be emitted as access-map.json for customer IaC to translate to IAM policies (tool does not apply IAM)"` **Cause:** The target uses IAM authentication. Kafka ACLs don't apply on IAM-auth clusters. Instead, the tool generates an `access-map.json` that maps source ACL principals and permissions to the IAM policies you need to create. **Action:** After migration, apply the generated IAM policies using your infrastructure-as-code tooling (Terraform, CloudFormation, CDK). * * * - [Production Migration Runbook](https://kafkabackup.com/guides/msk-kraft-migration-runbook.md) — step-by-step guide - [Configuration Reference](https://kafkabackup.com/enterprise/msk-kraft-config-reference.md) — tune precheck-related settings - [Troubleshooting](https://kafkabackup.com/troubleshooting/msk-kraft-migration.md) — post-precheck error resolution --- title: MSK KRaft Migration CLI Reference description: Complete CLI reference for all kafka-backup migrate msk-kraft subcommands — plan, precheck, execute, cutover, finalize, and more. source_url: html: https://kafkabackup.com/enterprise/msk-kraft-cli-reference md: https://kafkabackup.com/enterprise/msk-kraft-cli-reference.md --- # MSK KRaft Migration CLI Reference All migration commands use the subcommand structure: ``` kafka-backup migrate msk-kraft [OPTIONS] ``` Generate a migration plan without making any changes. **No license required.** ``` kafka-backup migrate msk-kraft plan \ --config \ [--format ] \ [--out-dir ] \ [--migration-id ] \ [--force] ``` | Flag | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `--config` | PATH | Yes | — | Path to migration config YAML | | `--format` | enum | No | `json` | Output format: `json`, `runbook`, `aws-cli`, `iam-policy`, `cost`, `all` | | `--out-dir` | PATH | No | — | Directory for output artifacts. Required when `--format all`. | | `--migration-id` | string | No | auto-generated | Pin the migration ID for deterministic output | | `--force` | flag | No | `false` | Overwrite existing files in `--out-dir` | | File | Description | | --- | --- | | `plan.json` | Machine-readable migration plan | | `runbook.md` | Step-by-step operator runbook | | `aws-cli.sh` | AWS CLI commands for infrastructure setup | | `iam-policy-templated.json` | IAM policy template with placeholders | | `iam-policy-concrete.json` | IAM policy with your actual ARNs | | `cost-estimate.json` | Estimated S3 and data transfer costs | ``` kafka-backup migrate msk-kraft plan \ --config migration.yaml \ --format all \ --out-dir ./migration-plan ``` * * * Read-only infrastructure and compatibility checks. **No license required.** ``` kafka-backup migrate msk-kraft precheck --config ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `--config` | PATH | Yes | Path to migration config YAML | JSON-formatted report with findings categorized as Blocker, Warning, or Info. See [Precheck Codes Reference](https://kafkabackup.com/enterprise/msk-kraft-precheck-codes.md) for all codes. Example output: ``` W04 warn: could not verify target message-size floor (target broker DescribeConfigs returned no message.max.bytes or replica.fetch.max.bytes (dynamic-config only on this broker)) — ensure target `message.max.bytes` and `replica.fetch.max.bytes` ≥ largest source topic's effective max.message.bytes W03 info: KMS key ARN set on backup channel — CMK access is not verified by this precheck phase; ensure the caller has kms:Encrypt/Decrypt/GenerateDataKey I01 info: target is IAM-auth — ACLs will be emitted as access-map.json for customer IaC to translate to IAM policies (tool does not apply IAM) ``` | Code | Meaning | | --- | --- | | 0 | All checks passed (warnings are OK) | | 1 | One or more blockers detected | * * * Run the migration from `planned` through `drain_ready`. Blocks until the data is caught up, then returns. **License required.** ``` kafka-backup migrate msk-kraft execute \ --config \ [--migration-id ] \ [--journal-dir ] \ [--force-restart] ``` | Flag | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `--config` | PATH | Yes | — | Path to migration config YAML | | `--migration-id` | string | No | auto-generated | Resume a specific migration. If omitted, a new migration starts. | | `--journal-dir` | PATH | No | — | Local directory for the migration journal. If omitted, journal is stored in S3. | | `--force-restart` | flag | No | `false` | Allow re-running seed after a partial run. Accepts risk of duplicate records on target. | ``` planned → precheck → topology_copy → seed → tail → drain_ready ``` Example journal excerpt: ``` 2026-05-08T07:25:53.232357Z topology_copy -> seed 2026-05-08T07:30:06.232018Z seed -> tail 2026-05-08T07:31:23.055520Z tail -> drain_ready drain ready: max_partition_lag=0 records_replayed=1222 bytes_replayed=85708 ``` | Code | Meaning | | --- | --- | | 0 | Migration reached `drain_ready` successfully | | 1 | Migration failed (check logs and journal) | ``` kafka-backup migrate msk-kraft execute \ --config migration.yaml \ --journal-dir ./journal # With debug logging RUST_LOG=debug kafka-backup migrate msk-kraft execute \ --config migration.yaml \ --journal-dir ./journal ``` * * * Resume an in-flight migration from its last journal entry. **License required.** ``` kafka-backup migrate msk-kraft resume \ --config \ --migration-id \ [--journal-dir ] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `--config` | PATH | Yes | Path to migration config YAML | | `--migration-id` | string | Yes | ID of the migration to resume | | `--journal-dir` | PATH | No | Directory where the journal is stored | | Code | Meaning | | --- | --- | | 0 | Migration advanced or reached a terminal state | | 1 | Resume failed | | 10 | Migration is waiting on operator action (check stderr for `next_command`) | * * * Show current state and live lag for a migration. **License required.** ``` kafka-backup migrate msk-kraft status \ --config \ --migration-id \ [--journal-dir ] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `--config` | PATH | Yes | Path to migration config YAML | | `--migration-id` | string | Yes | ID of the migration to query | | `--journal-dir` | PATH | No | Directory where the journal is stored | Prints migration state, per-partition lag, and operational notes. * * * Transition from `drain_ready` through producer freeze, sentinel publishing, offset translation, to `awaiting_client_switch`. **License required.** ``` kafka-backup migrate msk-kraft cutover \ --config \ --migration-id \ [--journal-dir ] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `--config` | PATH | Yes | Path to migration config YAML | | `--migration-id` | string | Yes | ID of the migration | | `--journal-dir` | PATH | No | Directory where the journal is stored | 1. Freezes producers (via webhook or TTY prompt) 2. Publishes sentinel records to every partition on source 3. Drains final records from source to target 4. Snapshots all consumer group offsets from source 5. Translates offsets using the offset map 6. Commits translated offsets on target 7. Verifies target log-start offsets have not advanced past the copied data 8. Logs `READY_FOR_CLIENT_SWITCH` Example output: ``` 2026-05-08T07:32:13.480416Z cutover -> awaiting_client_switch READY_FOR_CLIENT_SWITCH: groups_translated=3 offsets_committed=37 warnings=0 ``` | Code | Meaning | | --- | --- | | 0 | Cutover completed, ready for client switch | | 1 | Cutover failed (producers are automatically unfrozen on failure) | * * * Acknowledge that clients have switched to the target cluster. Transitions from `awaiting_client_switch` to `validating`. **License required.** ``` kafka-backup migrate msk-kraft cutover-ack \ --config \ --migration-id \ [--journal-dir ] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `--config` | PATH | Yes | Path to migration config YAML | | `--migration-id` | string | Yes | ID of the migration | | `--journal-dir` | PATH | No | Directory where the journal is stored | * * * Abort the migration and mark it as rolled back. **Only available before cutover completes.** **License required.** ``` kafka-backup migrate msk-kraft rollback \ --config \ --migration-id \ [--journal-dir ] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `--config` | PATH | Yes | Path to migration config YAML | | `--migration-id` | string | Yes | ID of the migration | | `--journal-dir` | PATH | No | Directory where the journal is stored | `planned`, `precheck`, `topology_copy`, `seed`, `tail`, `drain_ready` `cutover`, `awaiting_client_switch`, `validating`, `finalized` - Unfreezes producers (if frozen) - Marks migration as `rolled_back` in the journal - Uploads rollback report to evidence bucket - Does **not** delete topics or data on the target (manual cleanup) * * * Run the 5-check validation suite and produce the signed evidence bundle. Transitions from `validating` to `finalized` when validation passes or completes with accepted warnings. If a previous validation attempt found a repairable issue, `finalize` can be retried from `failed` after the repair. **License required.** ``` kafka-backup migrate msk-kraft finalize \ --config \ --migration-id \ [--journal-dir ] ``` | Flag | Type | Required | Description | | --- | --- | --- | --- | | `--config` | PATH | Yes | Path to migration config YAML | | `--migration-id` | string | Yes | ID of the migration | | `--journal-dir` | PATH | No | Directory where the journal is stored | 1. **Topic parity** — partition counts match 2. **Counts & offsets** — record counts within tolerance and target offset-floor safety holds 3. **Spot-check records** — sampled records byte-equal 4. **Sentinel presence** — cutover markers landed 5. **Consumer group reconciliation** — translated offsets committed | Code | Meaning | | --- | --- | | 0 | Validation passed, evidence bundle uploaded | | 1 | Validation failed (review the validation report) | Example successful verification output: ``` overall=PASSED topic_parity=PASSED detail="11 topic(s) match on partition count" counts_and_offsets=PASSED detail="68 partition(s) within count_tolerance=1" offset_floor_violations=0 spot_check_records=PASSED detail="204 samples compared, all matched" sentinel_presence=PASSED detail="sentinel observed on 68 partition(s)" consumer_group_reconciliation=PASSED detail="3 group(s) reconciled across 37 (group,topic,partition) triples" ``` `WARNING` can still be acceptable when the warning is limited to empty or zero-span partitions with no record available for spot comparison and every comparable sample matched. * * * These flags appear across most subcommands: | Flag | Description | | --- | --- | | `--config ` | Path to migration configuration YAML (always required) | | `--migration-id ` | Migration identifier (auto-generated by `execute`, required by others) | | `--journal-dir ` | Local journal directory (if omitted, journal stored in S3) | - [Configuration Reference](https://kafkabackup.com/enterprise/msk-kraft-config-reference.md) — config YAML details - [Precheck Codes](https://kafkabackup.com/enterprise/msk-kraft-precheck-codes.md) — precheck finding reference - [Production Runbook](https://kafkabackup.com/guides/msk-kraft-migration-runbook.md) — step-by-step guide --- title: Architecture Overview description: Understanding the OSO Kafka Backup architecture and design principles source_url: html: https://kafkabackup.com/architecture/overview md: https://kafkabackup.com/architecture/overview.md --- # Architecture Overview OSO Kafka Backup is designed for high-performance, reliable Kafka data protection with minimal operational overhead. Built in Rust for maximum throughput: - **Zero-copy data paths** where possible - **Async I/O** throughout the pipeline - **Parallel processing** across partitions - **Streaming architecture** - no full dataset buffering Every operation ensures data integrity: - **Checksums** at multiple levels - **Atomic writes** to storage - **Transactional commits** with checkpoints - **Validation tools** for verification Designed for easy operation: - **Single binary** - no dependencies - **YAML configuration** - declarative and version-controllable - **Kubernetes-native** - CRDs and operators - **Comprehensive CLI** - all operations accessible ``` ┌─────────────────────────────────────────────────────────────────────────┐ │ OSO Kafka Backup │ ├─────────────────────────────────────────────────────────────────────────┤ │ │ │ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ │ │ Backup │ │ Restore │ │ Offset │ │ │ │ Engine │ │ Engine │ │ Manager │ │ │ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │ │ │ │ │ │ │ ▼ ▼ ▼ │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ Core Library │ │ │ │ ┌─────────┐ ┌─────────┐ ┌─────────┐ ┌─────────┐ │ │ │ │ │ Kafka │ │ Storage │ │ Compress│ │ Crypto │ │ │ │ │ │ Client │ │ Backend │ │ Pipeline│ │ Module │ │ │ │ │ └─────────┘ └─────────┘ └─────────┘ └─────────┘ │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ │ └─────────────────────────────────────────────────────────────────────────┘ │ │ │ ▼ ▼ ▼ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ Kafka │ │ Object │ │ Local │ │ Cluster │ │ Storage │ │ Storage │ └──────────┘ └──────────┘ └──────────┘ ``` Responsible for reading data from Kafka and writing to storage: ``` Kafka Partitions Backup Engine Storage ┌───────────────┐ ┌─────────────────────┐ ┌──────────────┐ │ Partition 0 │───▶│ Consumer Thread 0 │───▶│ │ ├───────────────┤ ├─────────────────────┤ │ │ │ Partition 1 │───▶│ Consumer Thread 1 │───▶│ Segment │ ├───────────────┤ ├─────────────────────┤ │ Files │ │ Partition 2 │───▶│ Consumer Thread 2 │───▶│ │ └───────────────┘ └─────────────────────┘ └──────────────┘ │ ▼ ┌─────────────────────┐ │ Checkpoint Manager │ │ (Progress Tracking)│ └─────────────────────┘ ``` **Key features:** - Parallel consumption across partitions - Configurable batch sizes - Checkpoint-based progress tracking - Graceful shutdown with state preservation Responsible for reading from storage and writing to Kafka: ``` Storage Restore Engine Kafka Partitions ┌──────────────┐ ┌─────────────────────┐ ┌───────────────┐ │ │───▶│ Reader Thread 0 │───▶│ Partition 0 │ │ │ ├─────────────────────┤ ├───────────────┤ │ Segment │───▶│ Reader Thread 1 │───▶│ Partition 1 │ │ Files │ ├─────────────────────┤ ├───────────────┤ │ │───▶│ Reader Thread 2 │───▶│ Partition 2 │ └──────────────┘ └─────────────────────┘ └───────────────┘ │ ▼ ┌─────────────────────┐ │ Offset Header │ │ Injection │ └─────────────────────┘ ``` **Key features:** - PITR filtering by timestamp - Topic and partition remapping - Original offset preservation in headers - Idempotent writes support Handles consumer group offset translation: ``` ┌─────────────────────────────────────────────────────────────┐ │ Offset Manager │ ├─────────────────────────────────────────────────────────────┤ │ │ │ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ │ │ Mapping │───▶│ Strategy │───▶│ Reset │ │ │ │ Builder │ │ Resolver │ │ Executor │ │ │ └─────────────┘ └─────────────┘ └─────────────┘ │ │ │ │ │ │ ▼ ▼ │ │ ┌─────────────┐ ┌─────────────┐ │ │ │ Header │ │ Kafka │ │ │ │ Scanner │ │ Admin │ │ │ └─────────────┘ └─────────────┘ │ │ │ └─────────────────────────────────────────────────────────────┘ ``` **Key features:** - Multiple reset strategies - Parallel bulk operations - Snapshot and rollback support - Verification tools Abstraction layer for different storage systems: ``` ┌─────────────────────────────────────────────────────────────┐ │ Storage Backend Interface │ ├─────────────────────────────────────────────────────────────┤ │ pub trait StorageBackend { │ │ fn put(&self, key: &str, data: &[u8]) -> Result<()>; │ │ fn get(&self, key: &str) -> Result>; │ │ fn list(&self, prefix: &str) -> Result>; │ │ fn delete(&self, key: &str) -> Result<()>; │ │ } │ └─────────────────────────────────────────────────────────────┘ │ │ │ │ ▼ ▼ ▼ ▼ ┌─────────┐ ┌─────────┐ ┌─────────┐ ┌─────────┐ │ S3 │ │ Azure │ │ GCS │ │ Local │ │ Backend │ │ Backend │ │ Backend │ │ Backend │ └─────────┘ └─────────┘ └─────────┘ └─────────┘ ``` ``` 1. Initialize ├── Parse configuration ├── Connect to Kafka ├── Discover partitions └── Load checkpoint (if resuming) 2. Consume ├── Fetch records from Kafka ├── Batch records └── Apply compression 3. Store ├── Write segment file ├── Update manifest └── Save checkpoint 4. Finalize ├── Flush pending data ├── Write final manifest └── Close connections ``` ``` 1. Initialize ├── Parse configuration ├── Load backup manifest ├── Connect to target Kafka └── Create topics (if needed) 2. Filter ├── Apply time window (PITR) ├── Apply topic selection └── Apply partition mapping 3. Produce ├── Read segment files ├── Decompress records ├── Inject offset headers └── Write to Kafka 4. Finalize ├── Flush producers ├── Generate offset mapping └── Close connections ``` ``` Main Thread │ ├──▶ Partition Consumer 0 ──▶ Compression ──▶ Storage Writer │ ├──▶ Partition Consumer 1 ──▶ Compression ──▶ Storage Writer │ ├──▶ Partition Consumer 2 ──▶ Compression ──▶ Storage Writer │ └──▶ Checkpoint Thread (periodic) ``` - Each partition has dedicated consumer - Compression happens in parallel - Storage writes are batched and pipelined - Checkpoint updates are asynchronous ``` Main Thread │ ├──▶ Segment Reader 0 ──▶ Decompression ──▶ Producer Pool │ ├──▶ Segment Reader 1 ──▶ Decompression ──▶ Producer Pool │ └──▶ Progress Reporter (periodic) ``` - Multiple segment readers in parallel - Shared producer pool for Kafka writes - Backpressure handling to prevent memory overflow ``` # Memory constraints backup: batch_size: 10000 # Max records per batch max_batch_bytes: 104857600 # 100 MB max batch size storage: buffer_size: 16777216 # 16 MB write buffer ``` Data flows through the system without full buffering: ``` ┌─────────┐ ┌─────────┐ ┌─────────┐ ┌─────────┐ │ Read │───▶│ Process │───▶│ Compress│───▶│ Write │ │ (Kafka) │ │ (Memory)│ │ (Stream)│ │(Storage)│ └─────────┘ └─────────┘ └─────────┘ └─────────┘ │ ▼ Bounded Buffer (configurable) ``` ``` ┌─────────────────────────────────────────────────────┐ │ Error Classification │ ├─────────────────────────────────────────────────────┤ │ Transient │ Retry with backoff │ │ - Network timeout │ - Exponential backoff │ │ - Rate limited │ - Max retries configurable │ ├────────────────────┼────────────────────────────────┤ │ Recoverable │ Checkpoint and resume │ │ - Partial failure │ - Resume from checkpoint │ │ - Process restart │ - No data loss │ ├────────────────────┼────────────────────────────────┤ │ Fatal │ Fail with clear error │ │ - Auth failure │ - Detailed error message │ │ - Invalid config │ - Exit with error code │ └────────────────────┴────────────────────────────────┘ ``` For handling sustained failures: ``` backup: circuit_breaker: failure_threshold: 5 # Failures before opening reset_timeout_secs: 60 # Time before retry half_open_requests: 3 # Test requests when half-open ``` ``` ┌─────────────────────────────────────────────────────────────┐ │ Security Layer │ ├─────────────────────────────────────────────────────────────┤ │ │ │ Kafka Authentication │ │ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ │ │ SASL │ │ TLS │ │ mTLS │ │ │ │ (Various) │ │ (Encryption)│ │ (Mutual) │ │ │ └─────────────┘ └─────────────┘ └─────────────┘ │ │ │ │ Storage Authentication │ │ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ │ │ IAM Role │ │ Service │ │ Static │ │ │ │ (AWS/GCP) │ │ Principal │ │ Keys │ │ │ └─────────────┘ └─────────────┘ └─────────────┘ │ │ │ └─────────────────────────────────────────────────────────────┘ ``` - **In Transit**: TLS 1.2+ for all connections - **At Rest**: Storage-native encryption (SSE-S3, SSE-KMS, Azure encryption) - **Field-Level**: Enterprise feature for sensitive data ``` ┌─────────────────────────────────────────────────────────────┐ │ Metrics Collection │ ├─────────────────────────────────────────────────────────────┤ │ │ │ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ │ │ Counter │ │ Gauge │ │ Histogram │ │ │ │ (Records) │ │ (Progress) │ │ (Latency) │ │ │ └──────┬──────┘ └──────┬──────┘ └──────┬──────┘ │ │ │ │ │ │ │ └──────────────────┼──────────────────┘ │ │ │ │ │ ▼ │ │ ┌─────────────────┐ │ │ │ Prometheus │ │ │ │ Exporter │ │ │ └────────┬────────┘ │ │ │ │ └────────────────────────────┼────────────────────────────────┘ │ ▼ ┌─────────────────┐ │ Prometheus │ │ /metrics │ └─────────────────┘ ``` - [Offset Translation](https://kafkabackup.com/architecture/offset-translation.md) - How offset mapping works - [PITR Implementation](https://kafkabackup.com/architecture/pitr-implementation.md) - Point-in-time recovery details - [Compression](https://kafkabackup.com/architecture/compression.md) - Compression algorithms and tuning - [Zero-Copy Optimization](https://kafkabackup.com/architecture/zero-copy-optimization.md) - Performance optimizations --- title: Offset Translation description: Understanding how OSO Kafka Backup handles offset mapping between clusters source_url: html: https://kafkabackup.com/architecture/offset-translation md: https://kafkabackup.com/architecture/offset-translation.md --- # Offset Translation One of the most challenging aspects of Kafka backup and restore is maintaining consumer position across clusters. OSO Kafka Backup solves this through offset translation. Kafka offsets are cluster-specific and cannot be directly transferred: ``` Source Cluster Target Cluster ┌────────────────────────┐ ┌────────────────────────┐ │ Topic: orders │ │ Topic: orders │ │ ┌────────────────────┐ │ │ ┌────────────────────┐ │ │ │ Offset 0: Order A │ │ ───▶ │ │ Offset 0: Order A │ │ │ │ Offset 1: Order B │ │ │ │ Offset 1: Order B │ │ │ │ Offset 2: Order C │ │ │ │ Offset 2: Order C │ │ │ │ ... │ │ │ │ ... │ │ │ │ Offset 999: Last │ │ │ │ Offset 999: Last │ │ │ └────────────────────┘ │ └────────────────────────┘ └────────────────────────┘ ❌ Offsets may differ! ``` Offsets can differ because: 1. **Compacted topics** - Different compaction states 2. **Partial restore** - PITR excludes some messages 3. **Topic recreation** - Fresh topic starts at 0 4. **Partition changes** - Different partition count OSO Kafka Backup stores original offset information in message headers: ``` Original Message (Source) Restored Message (Target) ┌─────────────────────────┐ ┌─────────────────────────┐ │ Offset: 12345 │ │ Offset: 100 │ │ Key: order-123 │ ───▶ │ Key: order-123 │ │ Value: {...} │ │ Value: {...} │ │ Headers: [] │ │ Headers: │ │ │ │ x-original-offset │ │ │ │ : 12345 (i64 LE) │ │ │ │ x-original-timestamp │ │ │ │ : 1701234567890 │ │ │ │ x-source-cluster │ │ │ │ : source-cluster │ └─────────────────────────┘ └─────────────────────────┘ ``` All numeric values are **binary little-endian**, not decimal strings. | Header | Added by | Value | Example | | --- | --- | --- | --- | | `x-original-offset` | backup (default on) and restore | source offset, `i64` LE | `12345` | | `x-original-timestamp` | backup (default on) and restore | source timestamp, `i64` LE (epoch ms) | `1701234567890` | | `x-source-cluster` | backup, when `source_cluster_id` is set | UTF-8 | `prod-us-east-1` | | `x-source-partition` | restore only | source partition, `i32` LE | `0` | Offset headers are archived **by default** (`backup.include_offset_headers: true`); set `source_cluster_id` to also record the cluster: ``` backup: include_offset_headers: true # default source_cluster_id: "prod-us-east-1" ``` Restore adds its own set only when asked (default `false`; implied by `consumer_group_strategy: header-based`): ``` restore: include_original_offset_header: true ``` For records that are header-for-header identical to the source, drop the archived headers at restore time — offset mapping still works because the source offset is stored natively in every segment record: ``` restore: strip_offset_headers: true # v0.19.0+ include_original_offset_header: false ``` See [Offset-tracking headers](https://kafkabackup.com/reference/config-yaml.md#offset-tracking-headers). After restore, scan headers to build offset mapping: ``` ┌─────────────────────────────────────────────────────────────┐ │ Offset Mapping Table │ ├────────────────┬────────────────┬────────────────┬──────────┤ │ Topic │ Partition │ Source Offset │ Target │ ├────────────────┼────────────────┼────────────────┼──────────┤ │ orders │ 0 │ 12345 │ 100 │ │ orders │ 0 │ 12346 │ 101 │ │ orders │ 0 │ 12347 │ 102 │ │ orders │ 1 │ 5000 │ 50 │ │ orders │ 1 │ 5001 │ 51 │ │ payments │ 0 │ 8000 │ 200 │ └────────────────┴────────────────┴────────────────┴──────────┘ ``` ``` kafka-backup show-offset-mapping \ --bootstrap-servers target-kafka:9092 \ --topic orders \ --source-cluster "prod-us-east-1" ``` Output: ``` { "topic": "orders", "source_cluster": "prod-us-east-1", "mappings": [ { "partition": 0, "source_range": {"start": 12345, "end": 15000}, "target_range": {"start": 100, "end": 2755} }, { "partition": 1, "source_range": {"start": 5000, "end": 7500}, "target_range": {"start": 50, "end": 2550} } ] } ``` OSO Kafka Backup supports multiple strategies for resetting consumer offsets: Find target offset by scanning for original offset in headers: ``` offset_reset: strategy: header-based source_cluster: "prod-us-east-1" ``` ``` Consumer Group: order-processor Source Offset: 12345 (topic: orders, partition: 0) Scan target topic for x-original-offset: 12345 ↓ Found at target offset: 100 ↓ Reset consumer to offset: 100 ``` Find target offset by matching record timestamp: ``` offset_reset: strategy: timestamp timestamp: 1701234567890 ``` ``` Consumer Group: order-processor Source Timestamp: 1701234567890 Find first offset >= timestamp ↓ Target offset: 100 ↓ Reset consumer to offset: 100 ``` Simple reset to beginning or end: ``` offset_reset: strategy: earliest # or: latest ``` Reset to known offset: ``` offset_reset: strategy: offset offset: 100 ``` Use pre-generated mapping: ``` offset_reset: strategy: from-mapping mapping_file: /path/to/offset-mapping.json ``` ``` # Generate plan kafka-backup offset-reset plan \ --config offset-reset.yaml \ --output reset-plan.json # Review plan cat reset-plan.json # Execute kafka-backup offset-reset execute \ --plan reset-plan.json \ --bootstrap-servers target-kafka:9092 ``` For complete restoration including consumer offsets: ``` Phase 1: Data Restore ┌─────────────────────────────────────────┐ │ Restore messages to target cluster │ │ Include offset headers │ └─────────────────────────────────────────┘ │ ▼ Phase 2: Build Mapping ┌─────────────────────────────────────────┐ │ Scan restored topics │ │ Build source→target offset mapping │ └─────────────────────────────────────────┘ │ ▼ Phase 3: Reset Offsets ┌─────────────────────────────────────────┐ │ For each consumer group: │ │ Look up source offset │ │ Find corresponding target offset │ │ Reset consumer group │ └─────────────────────────────────────────┘ ``` Execute with single command: ``` kafka-backup three-phase-restore --config restore.yaml ``` When source offsets have gaps (compaction, deletion): ``` Source: 100, 101, 105, 106, 110 (gaps at 102-104, 107-109) Target: 0, 1, 2, 3, 4 (contiguous) Mapping: Source 100 → Target 0 Source 101 → Target 1 Source 105 → Target 2 (gap handled) Source 106 → Target 3 Source 110 → Target 4 ``` The mapping correctly handles gaps by using the actual records present. For compacted topics where messages are removed: ``` backup: include_offset_headers: true # Headers survive compaction if key remains ``` Consumer reset uses the most recent offset for each key: ``` If consumer was at offset 100 (key: A) But key A was compacted to offset 200 in source Then find the record with key A in target Reset to that offset ``` When target has different partition count: ``` Source: 3 partitions Target: 6 partitions Original message in partition 1, offset 500 May land in different partition due to repartitioning Solution: Include partition in header Scan ALL partitions in target for matching header ``` Configuration: ``` restore: # Records may be in different partitions include_original_offset_header: true offset_reset: strategy: header-based scan_all_partitions: true # Required for partition changes ``` When topics are renamed during restore: ``` restore: topic_mapping: orders: production-orders payments: production-payments ``` The offset mapping tracks both: ``` { "source_topic": "orders", "target_topic": "production-orders", "source_cluster": "prod-us-east-1", "mappings": [...] } ``` Scanning headers for offset mapping can be slow for large topics: ``` offset_reset: strategy: header-based # Performance tuning parallel_consumers: 10 # Parallel partition scanning sample_rate: 1.0 # 1.0 = scan all, 0.1 = 10% sample timeout_secs: 3600 # Max scan time ``` For many consumer groups, use bulk reset: ``` kafka-backup offset-reset-bulk \ --config offset-reset.yaml \ --groups-file consumer-groups.txt \ --parallelism 50 ``` This provides up to 50x speedup for large numbers of consumer groups. Offset mapping can be cached for repeated use: ``` # Generate and save mapping kafka-backup show-offset-mapping \ --bootstrap-servers kafka:9092 \ --all-topics \ --source-cluster "prod" \ --output mapping.json # Use cached mapping for multiple resets kafka-backup offset-reset execute \ --strategy from-mapping \ --mapping-file mapping.json \ --groups group1,group2,group3 ``` After resetting offsets: ``` # Check consumer group positions kafka-consumer-groups \ --bootstrap-server target-kafka:9092 \ --group order-processor \ --describe ``` Ensure consumers will process correct messages: ``` # Show what message consumer will receive next kafka-console-consumer \ --bootstrap-server target-kafka:9092 \ --topic orders \ --partition 0 \ --offset 100 \ --max-messages 1 \ --property print.headers=true ``` Check the `x-original-offset` header (8-byte little-endian `i64`) matches the expected source offset. 1. **Always enable offset headers** during backup 2. **Use three-phase restore** for complete migrations 3. **Test offset reset** in non-production first 4. **Take offset snapshots** before resetting (rollback capability) 5. **Verify consumer positions** after reset - [PITR Implementation](https://kafkabackup.com/architecture/pitr-implementation.md) - How time filtering affects offsets - [Offset Management Guide](https://kafkabackup.com/guides/offset-management.md) - Practical offset operations - [Three-Phase Restore](https://kafkabackup.com/reference/cli-reference.md#three-phase-restore) - CLI reference --- title: Point-in-Time Recovery Implementation description: How point-in-time recovery works in OSO Kafka Backup source_url: html: https://kafkabackup.com/architecture/pitr-implementation md: https://kafkabackup.com/architecture/pitr-implementation.md --- # Point-in-Time Recovery Implementation OSO Kafka Backup provides millisecond-precision point-in-time recovery (PITR), allowing you to restore Kafka data to any specific moment. Every Kafka record has a timestamp. PITR uses these timestamps to filter which records are restored: ``` Backup Contains Records: ┌──────────────────────────────────────────────────────────────┐ │ T=1000 T=1001 T=1002 T=1003 T=1004 T=1005 T=1006 │ │ │ │ │ │ │ │ │ │ │ ▼ ▼ ▼ ▼ ▼ ▼ ▼ │ │ ┌───┐ ┌───┐ ┌───┐ ┌───┐ ┌───┐ ┌───┐ ┌───┐ │ │ │ A │ │ B │ │ C │ │ D │ │ E │ │ F │ │ G │ │ │ └───┘ └───┘ └───┘ └───┘ └───┘ └───┘ └───┘ │ └──────────────────────────────────────────────────────────────┘ PITR Request: time_window_start=1002, time_window_end=1005 Restored Records: ┌──────────────────────────────────────────────────────────────┐ │ T=1002 T=1003 T=1004 T=1005 │ │ │ │ │ │ │ │ ▼ ▼ ▼ ▼ │ │ ┌───┐ ┌───┐ ┌───┐ ┌───┐ │ │ │ C │ │ D │ │ E │ │ F │ │ │ └───┘ └───┘ └───┘ └───┘ │ └──────────────────────────────────────────────────────────────┘ ``` Kafka records have two timestamp types: | Type | Description | When Set | | --- | --- | --- | | `CreateTime` | When producer created the record | Producer-side | | `LogAppendTime` | When broker received the record | Broker-side | PITR filters based on the record's actual timestamp, regardless of type. ``` mode: restore backup_id: "production-backup-20241201" restore: # Unix timestamp in milliseconds time_window_start: 1701388800000 # Dec 1, 2024 00:00:00 UTC time_window_end: 1701475199000 # Dec 1, 2024 23:59:59 UTC ``` The CLI accepts human-readable timestamps: ``` kafka-backup restore \ --config restore.yaml \ --time-start "2024-12-01T00:00:00Z" \ --time-end "2024-12-01T23:59:59Z" ``` Restore everything up to a specific moment: ``` restore: time_window_end: 1701450000000 # Stop at this point # Omit time_window_start to include all earlier records ``` Restore everything after a specific moment: ``` restore: time_window_start: 1701450000000 # Start from this point # Omit time_window_end to include all later records ``` Backups are stored in segment files with time ranges: ``` backup/ ├── manifest.json └── segments/ ├── segment-0000.dat # T: 1000-2000 ├── segment-0001.dat # T: 2001-3000 ├── segment-0002.dat # T: 3001-4000 └── segment-0003.dat # T: 4001-5000 PITR Request: T: 2500-3500 Segments to read: ✗ segment-0000.dat (T: 1000-2000) - Skip entirely ✓ segment-0001.dat (T: 2001-3000) - Read, filter records ✓ segment-0002.dat (T: 3001-4000) - Read, filter records ✗ segment-0003.dat (T: 4001-5000) - Skip entirely ``` This segment-level filtering significantly improves performance. Within each segment, individual records are filtered: ``` // Pseudocode for record filtering for record in segment.records() { let timestamp = record.timestamp(); if let Some(start) = time_window_start { if timestamp < start { continue; // Skip record before window } } if let Some(end) = time_window_end { if timestamp > end { continue; // Skip record after window } } // Record is within window output.write(record); } ``` The backup manifest tracks time ranges for efficient filtering: ``` { "backup_id": "production-backup-20241201", "topics": { "orders": { "partitions": { "0": { "segment_files": [ { "file": "orders-0-0000.dat", "start_offset": 0, "end_offset": 9999, "start_timestamp": 1701388800000, "end_timestamp": 1701392400000, "record_count": 10000 }, { "file": "orders-0-0001.dat", "start_offset": 10000, "end_offset": 19999, "start_timestamp": 1701392400001, "end_timestamp": 1701396000000, "record_count": 10000 } ] } } } } } ``` Different topics can have different recovery points: ``` restore: topics: - name: orders time_window_start: 1701388800000 time_window_end: 1701475199000 - name: payments time_window_start: 1701400000000 # Different start time_window_end: 1701475199000 - name: audit-log # No time window - restore everything ``` Bad data published at 14:30:00, restore to 14:29:59: ``` restore: # Restore everything before the corruption time_window_end: 1701437399000 # 14:29:59 ``` Investigate data between 10:00 and 11:00: ``` restore: time_window_start: 1701421200000 # 10:00:00 time_window_end: 1701424800000 # 11:00:00 # Restore to investigation cluster topic_mapping: orders: investigation-orders ``` Provide data as it existed on a specific date: ``` restore: # Full day - Dec 1, 2024 time_window_start: 1701388800000 # 00:00:00 UTC time_window_end: 1701475199999 # 23:59:59.999 UTC ``` Restore data to state before deployment: ``` restore: # Deployment started at 15:00 time_window_end: 1701442800000 # 14:59:59 # Only restore affected topics topics: - user-events - notifications ``` For accurate PITR, ensure producers set correct timestamps: ``` // Java producer example ProducerRecord record = new ProducerRecord<>( "orders", null, // partition System.currentTimeMillis(), // timestamp - set explicitly key, value ); ``` If producer clocks are skewed, PITR may not work as expected: ``` Producer A (clock +5 minutes): Record timestamp = 10:05 Producer B (clock correct): Record timestamp = 10:00 Actual time: 10:00 PITR restore to 10:02 would: ✓ Include record from Producer B (10:00) ✗ Exclude record from Producer A (10:05) - even though it was "really" at 10:00 ``` **Mitigation:** - Use NTP synchronization on all producers - Consider using `LogAppendTime` for broker-controlled timestamps Configure topics to use broker timestamp: ``` kafka-configs --bootstrap-server kafka:9092 \ --entity-type topics --entity-name orders \ --alter --add-config message.timestamp.type=LogAppendTime ``` PITR avoids reading unnecessary segments: ``` Backup: 100 segments (10 GB each = 1 TB total) PITR Window: 1 hour Segments in window: 4 segments (40 GB) Segments skipped: 96 segments (960 GB) I/O saved: 96% ``` Records are filtered during streaming, not post-processing: ``` Storage → Read → Filter → Decompress → Write to Kafka ↓ Skip early (no decompression overhead) ``` Each partition is processed in parallel with independent time filtering: ``` Partition 0: Filter T=1000-2000 ─────────────┐ Partition 1: Filter T=1000-2000 ─────────────┼──▶ Target Cluster Partition 2: Filter T=1000-2000 ─────────────┘ ``` After PITR restore, verify the time range: ``` # Check earliest message kafka-console-consumer \ --bootstrap-server kafka:9092 \ --topic orders \ --from-beginning \ --max-messages 1 \ --property print.timestamp=true # Check latest message kafka-run-class kafka.tools.GetOffsetShell \ --broker-list kafka:9092 \ --topic orders \ --time -1 ``` ``` kafka-backup describe \ --path s3://bucket/backups \ --backup-id production-20241201 \ --format json | jq '.time_range' ``` Output: ``` { "earliest_timestamp": 1701388800000, "latest_timestamp": 1701475199999, "earliest_timestamp_iso": "2024-12-01T00:00:00Z", "latest_timestamp_iso": "2024-12-01T23:59:59.999Z" } ``` - Kafka timestamps are millisecond precision - PITR is also millisecond precision - Sub-millisecond ordering is not guaranteed PITR filters by timestamp, not transaction boundaries: ``` Transaction: Records A, B, C (all at T=1000) PITR end: T=999 Result: Transaction partially restored (none of A, B, C) ``` If transaction integrity is required, ensure time windows align with transaction boundaries. For compacted topics, PITR works on the backup data: ``` Backup contains: All records at backup time PITR restores: Subset by timestamp Note: Compaction state differs from source ``` 1. **Use consistent time sources** - NTP synchronization 2. **Document time zones** - Use UTC for clarity 3. **Test PITR regularly** - Verify recovery works 4. **Include buffer time** - Add a few seconds margin 5. **Verify after restore** - Check time ranges match - [Offset Translation](https://kafkabackup.com/architecture/offset-translation.md) - Consumer offset handling with PITR - [Restore PITR Guide](https://kafkabackup.com/guides/restore-pitr.md) - Practical PITR operations - [CLI Reference](https://kafkabackup.com/reference/cli-reference.md#restore) - Restore command options --- title: Compression description: How OSO Kafka Backup compresses data for efficient storage source_url: html: https://kafkabackup.com/architecture/compression md: https://kafkabackup.com/architecture/compression.md --- # Compression OSO Kafka Backup supports multiple compression algorithms to reduce storage costs and improve transfer speeds. | Algorithm | Description | Best For | | --- | --- | --- | | **Zstd** | Zstandard - high ratio, fast | General use (default) | | **LZ4** | Very fast, moderate ratio | Speed-critical workloads | | **None** | No compression | Pre-compressed data | | Metric | Zstd (level 3) | Zstd (level 9) | LZ4 | None | | --- | --- | --- | --- | --- | | Compression ratio | 4-6x | 6-10x | 2-3x | 1x | | Compression speed | ~400 MB/s | ~100 MB/s | ~700 MB/s | N/A | | Decompression speed | ~1000 MB/s | ~900 MB/s | ~2000 MB/s | N/A | | CPU usage | Medium | High | Low | None | | Memory usage | Medium | High | Low | None | For 100 GB of JSON Kafka messages: | Algorithm | Compressed Size | Backup Time | Restore Time | | --- | --- | --- | --- | | Zstd-3 | ~20 GB | 5 min | 2 min | | Zstd-9 | ~12 GB | 15 min | 2 min | | LZ4 | ~40 GB | 3 min | 1 min | | None | 100 GB | 2 min | 2 min | ``` backup: compression: zstd # Algorithm: zstd, lz4, none compression_level: 3 # Level: 1-22 for zstd, 1-12 for lz4 ``` ``` backup: compression: zstd compression_level: 3 # Default, good balance # Level guidelines: # 1-3: Fast compression, good ratio # 4-6: Balanced (recommended) # 7-12: Slower, better ratio # 13-22: Very slow, best ratio (archival) ``` ``` backup: compression: lz4 compression_level: 1 # LZ4 levels have less impact ``` ``` backup: compression: none # Use when: # - Data is already compressed (images, video) # - Speed is critical and storage is cheap # - Debugging/inspection needed ``` ``` Kafka Records → Batch → Compress → Write to Storage ↓ ↓ ↓ ↓ Raw data Group Apply Segment (1 MB) records algorithm file (10 MB) (2 MB) (.zst) ``` Detailed flow: ``` ┌─────────────────────────────────────────────────────────────────────┐ │ Compression Pipeline │ ├─────────────────────────────────────────────────────────────────────┤ │ │ │ 1. Batch Records │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ Record 1 │ Record 2 │ Record 3 │ ... │ Record N │ │ │ │ (100 B) │ (200 B) │ (150 B) │ │ (180 B) │ │ │ └─────────────────────────────────────────────────────────────┘ │ │ │ │ │ ▼ │ │ 2. Serialize Batch │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ Header │ Record 1 │ Record 2 │ ... │ Record N │ Checksum │ │ │ │ (32 B) │ │ │ │ │ (4 B) │ │ │ └─────────────────────────────────────────────────────────────┘ │ │ │ │ │ ▼ │ │ 3. Compress │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ Zstd Frame │ │ │ │ ┌───────────────────────────────────────────────────────┐ │ │ │ │ │ Magic │ Frame Header │ Compressed Blocks │ Checksum │ │ │ │ │ └───────────────────────────────────────────────────────┘ │ │ │ └─────────────────────────────────────────────────────────────┘ │ │ │ │ │ ▼ │ │ 4. Write Segment │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ segment-0001.dat.zst │ │ │ └─────────────────────────────────────────────────────────────┘ │ │ │ └─────────────────────────────────────────────────────────────────────┘ ``` ``` Storage → Read → Decompress → Parse → Write to Kafka ↓ ↓ ↓ ↓ ↓ Segment Stream Apply Extract Produce file read algorithm records records (.zst) (2 MB) (10 MB) (array) (1 at a time) ``` OSO Kafka Backup uses streaming compression to minimize memory usage: ``` Traditional (Buffer All): ┌──────────────────────────────────────────────────────────────┐ │ Read all records → Buffer 10 GB → Compress → Write │ │ Memory usage: 10 GB+ │ └──────────────────────────────────────────────────────────────┘ Streaming (OSO Kafka Backup): ┌──────────────────────────────────────────────────────────────┐ │ Read batch → Compress batch → Write batch → (repeat) │ │ Memory usage: ~100 MB (configurable) │ └──────────────────────────────────────────────────────────────┘ ``` ``` backup: compression: zstd compression_level: 3 # Batch size controls memory vs efficiency batch_size: 10000 # Records per batch max_batch_bytes: 104857600 # 100 MB max batch ``` ``` Level 1-3: Fast mode ├── Speed: Very fast ├── Ratio: 3-4x └── Use case: Real-time backup, bandwidth constrained Level 4-6: Default mode ├── Speed: Fast ├── Ratio: 4-6x └── Use case: Daily backups, general use Level 7-12: High compression ├── Speed: Moderate ├── Ratio: 5-8x └── Use case: Weekly backups, archival Level 13-22: Ultra compression ├── Speed: Slow ├── Ratio: 6-10x └── Use case: Long-term archival, cold storage ``` ``` # Speed priority (CI/CD, real-time) backup: compression: zstd compression_level: 1 # Balanced (daily backups) backup: compression: zstd compression_level: 3 # Storage priority (archival) backup: compression: zstd compression_level: 9 # Maximum compression (cold storage, rare access) backup: compression: zstd compression_level: 19 ``` Different data types compress differently: | Data Type | Typical Ratio | Recommendation | | --- | --- | --- | | JSON | 6-10x | Zstd level 3-6 | | Avro | 4-6x | Zstd level 3 | | Protobuf | 3-5x | Zstd level 3 | | Plain text | 5-8x | Zstd level 3-6 | | Binary (random) | 1-1.5x | None or LZ4 | | Pre-compressed | 0.9-1.1x | None | Different topics may benefit from different settings: ``` # Global default backup: compression: zstd compression_level: 3 # Topic-specific overrides (if supported) # Note: Currently applies globally ``` Kafka itself supports compression (gzip, snappy, lz4, zstd). OSO Kafka Backup compresses at the batch level: ``` Kafka Message (may be compressed) ↓ OSO Backup reads (decompressed by Kafka client) ↓ OSO Backup compresses batch ↓ Storage Result: Backup compression works on decompressed data ``` If Kafka topics use compression: ``` # Kafka topic has gzip compression # OSO Backup still compresses (on decompressed data) backup: compression: zstd compression_level: 3 # This is efficient because: # 1. Kafka client decompresses automatically # 2. Zstd often achieves better ratio than gzip # 3. Zstd decompression is faster for restore ``` ``` High CPU, Low Storage: compression_level: 9-12 Result: Slower backup, smaller files, fast restore Balanced: compression_level: 3-6 Result: Fast backup, good compression, fast restore Low CPU, Higher Storage: compression_level: 1-2 Result: Very fast backup, larger files, fast restore ``` Zstd supports multi-threaded compression (used automatically): ``` CPU Cores: 8 Partitions: 16 Thread allocation: - 8 parallel partition consumers - Each with dedicated compression context - Effective throughput: 8x single-thread ``` ``` backup: compression: zstd compression_level: 3 # Higher level = more memory # Level 3: ~100 MB per compression context # Level 9: ~500 MB per compression context # Level 19: ~1 GB per compression context ``` ``` # Compression ratio kafka_backup_compression_ratio # Compression throughput (MB/s) rate(kafka_backup_bytes_compressed_total[5m]) / 1048576 # Time spent compressing kafka_backup_compression_duration_seconds ``` ``` kafka-backup describe \ --path s3://bucket/backups \ --backup-id my-backup \ --format json | jq '.compression' ``` Output: ``` { "algorithm": "zstd", "level": 3, "original_size_bytes": 10737418240, "compressed_size_bytes": 2147483648, "ratio": 5.0, "compression_time_secs": 120 } ``` 1. **Start with Zstd level 3** - good default for most cases 2. **Use LZ4 for speed-critical** - when backup window is tight 3. **Use higher levels for archival** - level 9+ for cold storage 4. **Disable for pre-compressed data** - images, video, encrypted ``` # Tier 1: Hot backups (hourly, 7-day retention) backup: compression: zstd compression_level: 3 # Fast backup, reasonable size # Tier 2: Warm backups (daily, 30-day retention) backup: compression: zstd compression_level: 6 # Balanced for daily use # Tier 3: Cold backups (weekly, 1-year retention) backup: compression: zstd compression_level: 12 # Maximum compression for long-term storage ``` For bandwidth-constrained environments: ``` backup: compression: zstd compression_level: 9 # Higher compression = less transfer # Trade-off: # - Slower backup (more CPU time) # - Less network transfer # - Smaller storage ``` ``` # Reduce compression level backup: compression: zstd compression_level: 1 # Fastest # Or switch to LZ4 backup: compression: lz4 ``` Check data type: ``` # Sample topic data kafka-console-consumer \ --bootstrap-server kafka:9092 \ --topic my-topic \ --max-messages 100 > sample.txt # Check compressibility zstd -3 sample.txt ls -la sample.txt* ``` ``` # Reduce batch size backup: compression: zstd compression_level: 3 batch_size: 1000 # Smaller batches max_batch_bytes: 10485760 # 10 MB max ``` - [Zero-Copy Optimization](https://kafkabackup.com/architecture/zero-copy-optimization.md) - Additional performance techniques - [Performance Tuning Guide](https://kafkabackup.com/guides/performance-tuning.md) - Optimization strategies - [Storage Format](https://kafkabackup.com/reference/storage-format.md) - How compressed data is stored --- title: Zero-Copy Optimization description: Performance optimizations in OSO Kafka Backup source_url: html: https://kafkabackup.com/architecture/zero-copy-optimization md: https://kafkabackup.com/architecture/zero-copy-optimization.md --- # Zero-Copy Optimization OSO Kafka Backup is built in Rust for maximum performance, employing various zero-copy and optimization techniques. Zero-copy refers to techniques that minimize or eliminate data copying between memory locations: ``` Traditional Copy Path: ┌─────────────────────────────────────────────────────────────────────┐ │ Network Buffer → Kernel Buffer → User Buffer → Process Buffer │ │ Copy 1 Copy 2 Copy 3 │ │ Total: 3 copies per record │ └─────────────────────────────────────────────────────────────────────┘ Zero-Copy Path: ┌─────────────────────────────────────────────────────────────────────┐ │ Network Buffer → User Buffer (mapped) → Process (reference) │ │ Copy 1 No copy No copy │ │ Total: 1 copy per record │ └─────────────────────────────────────────────────────────────────────┘ ``` ``` // Rust: Zero-cost abstractions // No garbage collection pauses // Predictable memory usage // Example: Processing Kafka records fn process_records(records: &[Record]) -> Result<()> { for record in records { // Borrow, don't copy let key = record.key(); // Reference, not copy let value = record.value(); // Reference, not copy // Process without allocation write_to_storage(key, value)?; } Ok(()) } ``` | Aspect | Rust (OSO Kafka Backup) | Java (Typical) | | --- | --- | --- | | GC Pauses | None | Yes (can be 100ms+) | | Memory overhead | ~0% | 30-50% (objects, GC) | | Startup time | Instant | Seconds (JVM warmup) | | Peak memory | Predictable | Variable | Reuse buffers instead of allocating new ones: ``` Without Pooling: ┌────────────────────────────────────────────────────────────────────┐ │ Record 1: allocate buffer → process → deallocate │ │ Record 2: allocate buffer → process → deallocate │ │ Record 3: allocate buffer → process → deallocate │ │ ... │ │ 1 million records = 1 million allocations │ └────────────────────────────────────────────────────────────────────┘ With Pooling: ┌────────────────────────────────────────────────────────────────────┐ │ Get buffer from pool → process Record 1 → return to pool │ │ Get buffer from pool → process Record 2 → return to pool │ │ Get buffer from pool → process Record 3 → return to pool │ │ ... │ │ 1 million records = ~10 allocations (pool size) │ └────────────────────────────────────────────────────────────────────┘ ``` Process data as streams without buffering entire datasets: ``` Buffered Approach (Memory-Heavy): ┌────────────────────────────────────────────────────────────────────┐ │ │ │ Read all records Store in memory Process all Write all │ │ (10 GB read) → (10 GB RAM) → (process) → (10 GB) │ │ │ │ Memory usage: 10+ GB │ └────────────────────────────────────────────────────────────────────┘ Streaming Approach (OSO Kafka Backup): ┌────────────────────────────────────────────────────────────────────┐ │ │ │ Read batch → Process batch → Write batch → (repeat) │ │ (100 MB) (100 MB) (100 MB) │ │ │ │ Memory usage: ~100 MB │ └────────────────────────────────────────────────────────────────────┘ ``` Non-blocking I/O for maximum throughput: ``` Synchronous (Blocking): ┌────────────────────────────────────────────────────────────────────┐ │ Thread 1: Read ████░░░░░░░░░ Write ████░░░░░░░░░ Read ████ │ │ (Idle while waiting) │ │ │ │ Throughput: Limited by sequential operations │ └────────────────────────────────────────────────────────────────────┘ Asynchronous (Non-Blocking): ┌────────────────────────────────────────────────────────────────────┐ │ Task 1: Read ████████████████████████████████████████ │ │ Task 2: ░░░░Write ████████████████████████████████████ │ │ Task 3: ░░░░░░░░░Read ████████████████████████████████ │ │ │ │ Throughput: Limited by I/O bandwidth │ └────────────────────────────────────────────────────────────────────┘ ``` Rust async implementation: ``` // Concurrent partition processing async fn backup_partitions(partitions: Vec) -> Result<()> { let futures: Vec<_> = partitions .into_iter() .map(|p| backup_partition(p)) .collect(); // Process all partitions concurrently join_all(futures).await?; Ok(()) } ``` Single Instruction, Multiple Data for compression: ``` Scalar Processing: ┌────────────────────────────────────────────────────────────────────┐ │ Process byte 0 │ │ Process byte 1 │ │ Process byte 2 │ │ Process byte 3 │ │ ... (one at a time) │ └────────────────────────────────────────────────────────────────────┘ SIMD Processing: ┌────────────────────────────────────────────────────────────────────┐ │ Process bytes 0-31 simultaneously (256-bit registers) │ │ Process bytes 32-63 simultaneously │ │ ... (32 at a time with AVX2) │ └────────────────────────────────────────────────────────────────────┘ ``` Zstd uses SIMD automatically when available. ``` ┌─────────────────────────────────────────────────────────────────────┐ │ Optimized Backup Path │ ├─────────────────────────────────────────────────────────────────────┤ │ │ │ Kafka Consumer │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ Fetch batch (zero-copy from network buffer) │ │ │ └──────────────────────────┬──────────────────────────────────┘ │ │ │ │ │ ▼ │ │ Record Processor │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ Iterate records (no copy, references only) │ │ │ │ Inject headers (minimal allocation) │ │ │ │ Serialize to batch format (streaming) │ │ │ └──────────────────────────┬──────────────────────────────────┘ │ │ │ │ │ ▼ │ │ Compression │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ Stream compress (SIMD-accelerated) │ │ │ │ Output directly to storage buffer │ │ │ └──────────────────────────┬──────────────────────────────────┘ │ │ │ │ │ ▼ │ │ Storage Writer │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ Async write (non-blocking) │ │ │ │ Multipart upload for large files │ │ │ └─────────────────────────────────────────────────────────────┘ │ │ │ └─────────────────────────────────────────────────────────────────────┘ ``` ``` ┌─────────────────────────────────────────────────────────────────────┐ │ Optimized Restore Path │ ├─────────────────────────────────────────────────────────────────────┤ │ │ │ Storage Reader │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ Async read (prefetch next segment) │ │ │ │ Range requests for PITR (skip unnecessary data) │ │ │ └──────────────────────────┬──────────────────────────────────┘ │ │ │ │ │ ▼ │ │ Decompression │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ Stream decompress (SIMD-accelerated) │ │ │ │ Decompress faster than network can deliver │ │ │ └──────────────────────────┬──────────────────────────────────┘ │ │ │ │ │ ▼ │ │ Record Parser │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ Parse records (zero-copy views into buffer) │ │ │ │ Apply PITR filter (skip without full parse) │ │ │ │ Apply topic remapping (header modification only) │ │ │ └──────────────────────────┬──────────────────────────────────┘ │ │ │ │ │ ▼ │ │ Kafka Producer │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ Batch produce (linger.ms optimization) │ │ │ │ Async send with backpressure │ │ │ └─────────────────────────────────────────────────────────────┘ │ │ │ └─────────────────────────────────────────────────────────────────────┘ ``` ``` // Cache-efficient: Contiguous memory struct RecordBatch { offsets: Vec, // Contiguous array timestamps: Vec, // Contiguous array keys: Vec, // Contiguous references values: Vec, // Contiguous references } // Iteration is cache-friendly for i in 0..batch.len() { process(batch.offsets[i], batch.values[i]); } ``` ``` // Bad: Allocation per record fn process_bad(records: &[Record]) -> Vec { records.iter() .map(|r| ProcessedRecord::new(r)) // Allocation! .collect() } // Good: In-place processing fn process_good(records: &mut [Record]) { for record in records { record.process_in_place(); // No allocation } } ``` ``` # Optimal TCP settings for high-throughput source: kafka_config: socket.receive.buffer.bytes: 1048576 # 1 MB receive buffer fetch.max.bytes: 52428800 # 50 MB max fetch fetch.min.bytes: 1048576 # 1 MB min (reduce round trips) fetch.max.wait.ms: 500 # Wait for batches target: kafka_config: socket.send.buffer.bytes: 1048576 # 1 MB send buffer batch.size: 1048576 # 1 MB batches linger.ms: 100 # Wait for batches ``` ``` Without Pipelining: ┌────────────────────────────────────────────────────────────────────┐ │ Request 1 → Wait → Response 1 → Request 2 → Wait → Response 2 │ │ ████ ████ │ │ Idle time: High │ └────────────────────────────────────────────────────────────────────┘ With Pipelining: ┌────────────────────────────────────────────────────────────────────┐ │ Request 1 → Request 2 → Request 3 → Response 1 → Response 2 → ... │ │ │ │ Idle time: Minimal │ └────────────────────────────────────────────────────────────────────┘ ``` Large files are uploaded in parallel parts: ``` Single Upload: ┌────────────────────────────────────────────────────────────────────┐ │ Upload 1 GB file: ████████████████████████████████ 60s │ └────────────────────────────────────────────────────────────────────┘ Multipart Upload (10 parts): ┌────────────────────────────────────────────────────────────────────┐ │ Part 1: ████████ 6s │ │ Part 2: ████████ 6s (parallel) │ │ Part 3: ████████ 6s (parallel) │ │ ... │ │ Total: ~10s (6x faster) │ └────────────────────────────────────────────────────────────────────┘ ``` Configuration: ``` storage: backend: s3 multipart_threshold: 104857600 # 100 MB multipart_part_size: 10485760 # 10 MB parts max_concurrent_uploads: 10 ``` Read-ahead for sequential access: ``` Without Prefetch: ┌────────────────────────────────────────────────────────────────────┐ │ Read segment 1 → Process → Read segment 2 → Process → ... │ │ (wait) (wait) │ └────────────────────────────────────────────────────────────────────┘ With Prefetch: ┌────────────────────────────────────────────────────────────────────┐ │ Read 1 → Process 1 → Process 2 → Process 3 → ... │ │ Read 2 ──────────────┘ │ │ │ Read 3 ───────────────────┘ │ │ (overlap read with process) │ └────────────────────────────────────────────────────────────────────┘ ``` Testing on AWS (c5.4xlarge, GP3 storage): | Operation | OSO Kafka Backup | Typical Java Tool | | --- | --- | --- | | Backup (1 partition) | 150 MB/s | 40 MB/s | | Backup (10 partitions) | 1.2 GB/s | 300 MB/s | | Restore (1 partition) | 200 MB/s | 60 MB/s | | Restore (10 partitions) | 1.5 GB/s | 400 MB/s | | Dataset Size | OSO Kafka Backup | Typical Java Tool | | --- | --- | --- | | 1 GB backup | 100 MB | 1.5 GB | | 10 GB backup | 150 MB | 4 GB | | 100 GB backup | 200 MB | 16 GB+ | Backup operation latency (per batch): | Percentile | OSO Kafka Backup | Typical Java Tool | | --- | --- | --- | | p50 | 2 ms | 10 ms | | p99 | 8 ms | 50 ms | | p99.9 | 15 ms | 200 ms | ``` source: kafka_config: fetch.max.bytes: 104857600 # 100 MB max.partition.fetch.bytes: 10485760 # 10 MB per partition fetch.min.bytes: 1048576 # 1 MB minimum backup: batch_size: 100000 # Large batches max_batch_bytes: 104857600 # 100 MB max compression: zstd compression_level: 1 # Fast compression checkpoint_interval_secs: 60 # Less frequent checkpoints storage: multipart_threshold: 52428800 # 50 MB multipart_part_size: 10485760 # 10 MB max_concurrent_uploads: 20 ``` ``` source: kafka_config: fetch.max.wait.ms: 100 # Don't wait too long backup: batch_size: 10000 # Smaller batches max_batch_bytes: 10485760 # 10 MB max compression: lz4 # Fastest compression checkpoint_interval_secs: 10 # Frequent checkpoints ``` ``` # Records per second rate(kafka_backup_records_total[5m]) # Bytes per second rate(kafka_backup_bytes_total[5m]) # Batch processing time histogram_quantile(0.99, kafka_backup_batch_duration_seconds_bucket) # Memory usage process_resident_memory_bytes ``` ``` If records/sec is low but CPU is low: → Bottleneck is I/O (network or storage) → Increase batch sizes, parallelism If records/sec is low and CPU is high: → Bottleneck is compression → Reduce compression level or use LZ4 If memory is growing: → Backpressure not working correctly → Reduce batch sizes ``` - [Performance Tuning Guide](https://kafkabackup.com/guides/performance-tuning.md) - Practical optimization - [Compression](https://kafkabackup.com/architecture/compression.md) - Algorithm tuning - [Metrics Reference](https://kafkabackup.com/reference/metrics.md) - All performance metrics --- title: MSK KRaft Migration Architecture description: Deep dive into the kafka-backup Enterprise MSK ZooKeeper to KRaft migration architecture — state machine, offset continuity, data flow, and evidence bundle design. source_url: html: https://kafkabackup.com/architecture/msk-kraft-migration md: https://kafkabackup.com/architecture/msk-kraft-migration.md --- # MSK KRaft Migration Architecture This page covers the internal architecture of the MSK ZK→KRaft migration pipeline — how data flows, how offsets are translated, and how the evidence bundle is constructed. Three principles guide the architecture: 1. **No in-place migration.** S3 is the replication channel between source and target. The source cluster is never modified — it remains a safe rollback target until you decommission it. 2. **Deterministic state machine.** Every phase transition is journaled. The migration can be resumed from any point, rolled back before cutover, or audited post-facto from the evidence bundle. 3. **Offset continuity over offset identity.** Target offsets will differ from source offsets (compaction, replication timing). What matters is that consumers resume from the **same message** — the offset map translates between the two. The migration uses 12 persisted states with deterministic transitions: ``` +--------------------------------------------------------------------+ | FORWARD PATH | +--------------------------------------------------------------------+ PLANNED -> PRECHECK -> TOPOLOGY_COPY -> SEED -> TAIL -> DRAIN_READY | v FINALIZED <- VALIDATING <- AWAITING_CLIENT_SWITCH <- CUTOVER <------+ ROLLBACK (before cutover only): PLANNED / PRECHECK / TOPOLOGY_COPY / SEED / TAIL / DRAIN_READY | v ROLLED_BACK FAILED <- (reachable from any state) ``` | State | Serde name | Description | | --- | --- | --- | | Planned | `planned` | Fresh migration, no mutations yet | | Precheck | `precheck` | Read-only cluster analysis | | TopologyCopy | `topology_copy` | Topics and ACLs created on target | | Seed | `seed` | Bulk data transfer via S3 | | Tail | `tail` | Continuous replication bridging seed to cutover | | DrainReady | `drain_ready` | Lag converged, awaiting operator | | Cutover | `cutover` | Producer freeze, sentinel, offset translation | | AwaitingClientSwitch | `awaiting_client_switch` | Operator switching applications | | Validating | `validating` | 5-check validation running, including target offset-floor safety inside counts and offsets | | Finalized | `finalized` | Evidence signed and uploaded (terminal) | | RolledBack | `rolled_back` | Migration aborted (terminal) | | Failed | `failed` | Phase failed (resumable) | On resume, the system reads the journal and dispatches based on the tip state: - **Planned through Tail**: Re-enter execute from the last non-failed state - **DrainReady / AwaitingClientSwitch**: Return "waiting for operator" (exit code 10) - **Failed**: Walk backward (max 4 hops) to find the last non-failed state; re-enter from there - **Finalized / RolledBack**: No-op (terminal) A **resume fingerprint** (SHA256 of source ARN, target ARN, bucket names) is embedded in the first journal entry. On resume, the fingerprint is verified — this detects silent config changes between crash and resume. The seed phase performs a bulk copy of all source data through S3: ``` Source Cluster S3 Target Cluster ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ │ │ │ │ │ │ Topic A:0 │──────► │ segments/ │──────► │ Topic A:0 │ │ Topic A:1 │ read │ A/0/... │ write │ Topic A:1 │ │ Topic B:0 │ │ A/1/... │ │ Topic B:0 │ │ ... │ │ B/0/... │ │ ... │ │ │ │ │ │ │ └──────────────┘ │ offset-map │ └──────────────┘ │ .json │ └──────────────┘ ``` 1. **BackupEngine** reads source partitions and writes S3 segments (up to `segment_max_bytes` each) 2. **RestoreEngine** reads segments from S3 and produces to target partitions 3. **Offset map sidecar** records the mapping: for partition P, source offsets `[src_first, src_last]` landed at target offsets `[tgt_first, tgt_last]` The seed phase runs `max_concurrent_partitions` transfers in parallel (default 4). The current artifact layout stores the migration offset map as: ``` s3://///offset-map.json ``` After seed, the tail phase bridges the gap between the seed snapshot and real-time: ``` +------------------+ +------------------+ +------------------+ | Source Cluster | | Tail Bridge | | Target Cluster | | | | | | | | Consumer |----->| Producer |----->| New records | | (from seed | | (to target) | | appended | | end offset) | | | | | +------------------+ | Updates | +------------------+ | offset-map | +------------------+ ``` Tail is a direct consumer→producer bridge: - Consumes from source starting at the seed end offsets - Produces to target, updating the offset map continuously - Tracks per-partition lag (source LEO - consumer position) - Declares `drain_ready` when all partitions are within `drain_max_partition_lag` for `drain_stable_window` Tail progress is checkpointed to S3 — on resume, it continues from the checkpoint without re-streaming from seed offsets. Cutover is the critical phase that ensures offset continuity: Producers on the source cluster are frozen (via webhook or manual confirmation). This creates a stable "end of source" boundary. A sentinel record is published to every partition on the source: ``` Header: x-kbe-cutover-sentinel = Key: __kbe_sentinel__ Value: {"migration_id": "", "partition":

, "timestamp": ""} ``` The sentinel marks the exact boundary between "migrated data" and "nothing else." Tail continues until every sentinel is replicated to the target. At this point, the target has all source data up to and including the sentinel. All consumer group committed offsets are fetched from the source cluster. For each consumer group's committed offset on each partition: ``` target_offset = target_first + (source_committed - source_first) ``` Where `source_first`/`target_first` come from the offset map sidecar. **Edge cases:** - `source_committed < source_first` → **ResetToStart**: consumer was behind the earliest retained offset. Translated to `target_first`. - `source_committed > source_last` → **Clamped**: consumer was ahead of the migration boundary. Translated to `target_last`. - `target_first` missing → **Skip**: partition has no data on target. Logged as warning. Translated offsets are committed to the target cluster's consumer group coordinators via the Kafka `OffsetCommit` API. Before the tool logs `READY_FOR_CLIENT_SWITCH`, it fetches target earliest and latest offsets for every migrated partition. It blocks the client switch if the target log-start has advanced past the first copied offset or if the target end offset is behind the expected copied range. This catches retention or `DeleteRecords` truncation that a latest-offset-only drain check would miss. This is the core differentiator. After cutover: ``` Source (ZK cluster): Target (KRaft cluster): +------------------------------+ +------------------------------+ | Partition 0: | | Partition 0: | | [0] msg-A | | [0] msg-A | | [1] msg-B | | [1] msg-B | | [2] msg-C <- committed | | [2] msg-C <- committed | | [3] msg-D <- next msg | | [3] msg-D <- next msg | | [4] msg-E | | [4] msg-E | | [5] sentinel | | [5] sentinel | +------------------------------+ +------------------------------+ ``` The consumer group's committed offset points to the same **message content** on both clusters. When the consumer reconnects to the target, it reads `msg-D` next — exactly where it left off on the source. The offset **numbers** may differ between source and target (due to compaction or replication timing), but the offset map ensures the translation is correct. All ACL bindings are fetched from the source via the `DescribeAcls` API. MSK internal bindings are filtered: - `User:ANONYMOUS` (MSK internal) - Bindings on `__consumer_offsets`, `__transaction_state`, and other internal topics - `kafka-cluster:ClusterAction` on cluster resource (MSK auto-manages these) | Policy | Source-only ACLs | Target-only ACLs | | --- | --- | --- | | `merge` | Created on target | Left in place | | `replace` | Created on target | Reported (not deleted) | | `refuse` | Error — migration stops | Error — migration stops | When the target uses IAM auth, Kafka ACLs don't apply. Instead, the tool generates `access-map.json` — a mapping of each principal's permissions to the equivalent IAM actions. The operator applies these via their IAM tooling. The evidence bundle is a JSON document signed with Ed25519: ``` { "bundle_json": "", "signature_b64": "", "public_key_b64": "" } ``` The `bundle_json` payload contains: | Section | Contents | | --- | --- | | `migration_id` | Unique migration identifier | | `tool_version` | kafka-backup version | | `signed_at` | UTC timestamp | | `config_fingerprint` | Hash of the migration config | | `journal` | Complete state transition history | | `source` / `target` | Cluster metadata snapshots | | `plan` | Full migration plan | | `topology` | Topics created/updated, configs applied | | `acls` | ACL bindings copied, internals filtered | | `seed` | Records, bytes, partitions transferred | | `tail` | Records replayed during tail phase | | `drain_final` | Per-partition lag at finalize time | | `cutover` | Sentinel positions, freeze timing, offset translations | | `validation` | All 5 check outcomes with per-partition detail; `counts_and_offsets.data.offset_floor_violations` records target offset-floor safety | Evidence is uploaded to two keys: 1. Attempt-scoped immutable key: `s3://///evidence-attempts/--.json` 2. Latest alias for the migration: `s3://///evidence.json` Each upload first tries S3 PutObject with COMPLIANCE-mode Object Lock retention derived from `evidence.retention`. If the bucket lacks Object Lock, the tool uploads without retention and logs a warning. Every state transition is appended to `journal.jsonl`: ``` {"migration_id":"...","from":"seed","to":"tail","at":"2026-04-24T10:40:00Z","reason":"seed complete: 900 records, 4 partitions"} {"migration_id":"...","from":"tail","to":"drain_ready","at":"2026-04-24T10:41:00Z","reason":"all partitions within lag tolerance"} ``` On resume, the journal is loaded from S3 (or local directory), the tip state is determined, and execution re-enters from the appropriate phase. If the tip state is `failed`, the system walks backward through the journal to find the last non-failed state (capped at 4 hops to prevent oscillation). It then re-enters from that state. Available from: `planned`, `precheck`, `topology_copy`, `seed`, `tail`, `drain_ready` Rollback: 1. Verifies the resume fingerprint 2. Best-effort producer unfreeze (if frozen) 3. Uploads `rollback-report.json` to evidence bucket 4. Appends `→ rolled_back` to journal 5. Does **not** delete topics/data on target (manual cleanup) - [MSK KRaft Migration Overview](https://kafkabackup.com/enterprise/msk-kraft-migration.md) — feature summary and quick start - [Configuration Reference](https://kafkabackup.com/enterprise/msk-kraft-config-reference.md) — every config field - [Production Runbook](https://kafkabackup.com/guides/msk-kraft-migration-runbook.md) — step-by-step guide --- title: Index source_url: html: https://kafkabackup.com/benchmarks/index md: https://kafkabackup.com/benchmarks/index.md --- --- title: Index source_url: html: https://kafkabackup.com/well-architected/index md: https://kafkabackup.com/well-architected/index.md --- --- title: Operational Excellence description: Design principles and best practices for running and monitoring Kafka backup workloads effectively with OSO Kafka Backup source_url: html: https://kafkabackup.com/well-architected/operational-excellence md: https://kafkabackup.com/well-architected/operational-excellence.md --- # Operational Excellence > "The ability to run and monitor Kafka backup workloads effectively, gain insight into their operations, and continuously improve supporting processes and procedures to deliver reliable data protection." The Operational Excellence pillar focuses on ensuring your backup and restore operations are reliable, observable, and continuously improving. It encompasses how your team organises around backup responsibilities, automates routine tasks, monitors health, responds to incidents, and evolves practices over time. 1. **Perform operations as code** -- Define backup schedules, retention policies, and restore procedures as version-controlled configuration. Eliminate manual, ad-hoc CLI invocations for production workloads. 2. **Make frequent, small, reversible changes** -- Roll out configuration changes incrementally (e.g., one topic pattern at a time). Use canary deployments for operator upgrades and validate each change before proceeding. 3. **Refine operations procedures frequently** -- Review runbooks and automation after every incident and at regular intervals. Update them to reflect current topology, tooling versions, and lessons learned. 4. **Anticipate failure** -- Design backup pipelines assuming that Kafka brokers, storage backends, and network connectivity will fail. Build pre-mortems into planning and run regular disaster-recovery drills. 5. **Learn from all operational events** -- Treat successful restores, slow backups, and outright failures equally as sources of insight. Conduct blameless post-incident reviews and feed findings back into automation and monitoring. 6. **Use managed services where possible** -- Leverage managed object storage (S3, GCS, Azure Blob), managed Kubernetes, and managed Kafka where appropriate to reduce undifferentiated operational burden and let your team focus on backup-specific concerns. * * * Establish clear ownership, on-call procedures, and team competency for Kafka backup operations. Backup systems that lack a clear owner tend to drift into a neglected state. When a restore is needed during an outage, confusion over who is responsible and what steps to follow turns a recoverable incident into a prolonged one. - **Assign a backup operations owner** -- A named individual or team accountable for backup health, capacity planning, and restore readiness. - **Define on-call procedures** -- Include backup/restore responsibilities in your existing on-call rotation. Ensure on-call engineers have the necessary access and credentials. - **Maintain a disaster-recovery playbook** -- Document exact `kafka-backup` CLI commands for every recovery scenario. Store the playbook alongside your infrastructure code, not in a separate wiki. - **Run quarterly DR drills** -- Execute full and partial restores against a staging environment. Record time-to-restore and compare against your RTO targets. - **Train all platform engineers** -- Every engineer on the team should be able to execute a restore independently. Avoid single points of knowledge. - **Define escalation paths** -- Document when to escalate from on-call to the backup owner, and from the backup owner to OSO support (for Enterprise customers). > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Store your DR playbook in the same Git repository as your backup configuration. This ensures the playbook is always versioned alongside the config it references. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **No designated owner** -- Backup is "everyone's responsibility", which means it is no one's responsibility. > > - **Untested playbook** -- A restore procedure that has never been executed is not a procedure; it is a hope. > > - **Single point of knowledge** -- Only one engineer knows how to operate `kafka-backup`. When they are unavailable, the team is blocked. * * * Define and automate the full lifecycle of backups: scheduling, validation, retention, and deletion. Without lifecycle automation, storage costs grow unchecked, stale backups give a false sense of security, and teams discover validation failures only when a restore is attempted during an incident. - **Define schedules with cron** -- Use Kubernetes CronJobs or the operator's built-in scheduling to run backups at predictable intervals. - **Automate validation** -- Run `kafka-backup validate --deep` after every backup completes. Deep validation checks segment integrity, offset continuity, and header consistency. - **Set retention policies per environment:** - **Development:** 7 days - **Production:** 90 days - **Compliance/Audit:** 7 years (with immutable storage locks) - **Automate deletion deliberately** -- Use `KafkaBackup.spec.retention` for operator-managed pruning on PVC/local, S3/S3-compatible, and Azure Blob Storage. Use storage lifecycle policies when you need backend-native legal hold, object lock, cross-account enforcement, or GCS support. Never rely on ad hoc manual cleanup. - **Tag backups** -- Apply metadata labels (environment, team, compliance tier) to every backup for filtering, reporting, and cost allocation. ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: production-nightly namespace: kafka-backup labels: environment: production team: platform compliance-tier: standard spec: schedule: "0 0 2 * * * *" stopAtCurrentOffsets: true kafkaCluster: bootstrapServers: - kafka-0.kafka:9092 - kafka-1.kafka:9092 - kafka-2.kafka:9092 topics: - "orders-*" - "payments-*" storage: storageType: s3 s3: bucket: acme-kafka-backups-prod region: eu-west-1 prefix: nightly/ credentialsSecret: name: s3-credentials compression: zstd retention: enabled: true maxAgeDays: 90 keepLast: 7 dryRun: true ``` This example starts in retention dry-run mode. Review the retention status fields and switch `dryRun` to `false` only after the policy has been approved. If retention must be enforced by the storage backend instead, pair the backup prefix with a lifecycle rule: ``` { "Rules": [ { "ID": "expire-nightly-backups-after-90-days", "Status": "Enabled", "Filter": { "Prefix": "nightly/" }, "Expiration": { "Days": 90 } } ] } ``` > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Always enable `validation.deep: true` for production backups. Shallow validation only checks that files exist; it does not verify data integrity. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **No automated deletion** -- Storage costs grow linearly and old backups become a liability rather than an asset. > > - **No post-backup validation** -- You discover corrupt backups at the worst possible time: during a restore. > > - **Same retention everywhere** -- Applying production retention to development wastes storage; applying development retention to compliance data violates regulations. * * * Instrument backup and restore operations with metrics, dashboards, and alerts to maintain full visibility into pipeline health. Backups are background processes. Without observability, failures go unnoticed until a restore is needed. By then, your most recent valid backup may be hours or days old -- far outside your RPO. - **Enable Prometheus metrics** on port 8080 for all `kafka-backup` instances. - **Monitor key metrics:** - `kafka_backup_lag_records` -- Consumer lag per partition. Rising lag indicates the backup cannot keep pace with production throughput. - `kafka_backup_records_total` -- Total records backed up. Use the rate to track throughput. - `kafka_backup_compression_ratio` -- Compression efficiency. A sudden change may indicate a shift in message format. - `kafka_backup_storage_write_latency_seconds` -- Storage backend latency. Elevated latency degrades backup performance and may indicate storage issues. - **Build Grafana dashboards** for: - Overall backup health (active jobs, success/failure rates) - Per-topic backup status and lag - Storage growth trends and cost projection - Restore operation tracking and duration - **Configure alerts** for: - Backup job failure (any job that does not complete successfully) - Consumer lag exceeding RPO threshold - Storage write errors or elevated latency - Checkpoint staleness (no checkpoint update within expected interval) ``` # kafka-backup config metrics: enabled: true port: 8080 bind_address: "0.0.0.0" path: "/metrics" ``` ``` # Prometheus scrape config scrape_configs: - job_name: kafka-backup kubernetes_sd_configs: - role: pod relabel_configs: - source_labels: [__meta_kubernetes_pod_label_app] regex: kafka-backup action: keep - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port] action: replace target_label: __address__ regex: (.+) replacement: ${1}:8080 ``` > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Set up a dedicated "Backup Health" Grafana dashboard and include it in your team's daily standup review. Catching a slow backup trend early is far cheaper than discovering a gap during an incident. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **No monitoring at all** -- You have no idea whether backups are running, succeeding, or falling behind. > > - **Monitoring backup jobs but not storage** -- A backup that completes but fails to write to storage is worse than a visible failure; it is a silent one. > > - **No alerting** -- Dashboards that no one watches provide no value. Alerts ensure the right people are notified at the right time. * * * Create detailed, executable runbooks for every backup and restore scenario. Automate routine operations and keep manual procedures as copy-paste-ready CLI commands. During an outage, engineers are under pressure. Runbooks that contain exact commands -- not prose descriptions -- dramatically reduce mean time to recovery (MTTR). Automation eliminates human error from repetitive tasks. - **Maintain runbooks for:** - Full cluster restore from backup - Single topic point-in-time recovery (PITR) - Consumer offset recovery - Backup failure investigation and remediation - Storage backend failover - Configuration change rollout - **Each runbook must include:** - Prerequisites (access, credentials, tooling versions) - Step-by-step CLI commands (copy-paste ready) - Validation steps after each action - Rollback procedure - Estimated time to complete ``` # Step 1: List available backups for the target topic kafka-backup list \ --storage s3 \ --bucket acme-kafka-backups-prod \ --prefix nightly/ \ --topic orders-events \ --from "2026-03-20T00:00:00Z" \ --to "2026-03-23T23:59:59Z" ``` ``` # Step 2: Describe the selected backup to confirm contents kafka-backup describe \ --storage s3 \ --bucket acme-kafka-backups-prod \ --backup-id backup-20260322-020000 ``` ``` # Step 3: Create restore configuration (restore-orders.yaml) restore: kafka: bootstrapServers: "kafka-0.kafka:9092,kafka-1.kafka:9092,kafka-2.kafka:9092" storage: type: s3 s3: bucket: acme-kafka-backups-prod region: eu-west-1 topics: include: - "orders-events" pointInTime: "2026-03-22T14:30:00Z" targetTopic: "orders-events-restored" restoreOffsets: true ``` ``` # Step 4: Execute the restore kafka-backup restore --config restore-orders.yaml # Step 5: Validate the restored topic kafka-backup validate \ --deep \ --topic orders-events-restored \ --kafka-bootstrap "kafka-0.kafka:9092" ``` > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Always restore to a separate target topic (e.g., `orders-events-restored`) first. Validate the data before swapping consumers to the restored topic. Never overwrite a production topic directly. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **Tribal knowledge** -- Restore procedures exist only in one engineer's head. They are effectively unavailable at 3 a.m. on a Sunday. > > - **Prose without commands** -- "Connect to the cluster and restore the topic" is not a runbook. Exact commands with exact flags are a runbook. > > - **Referencing an external wiki during an outage** -- If your Confluence page is behind an SSO that depends on the infrastructure you are trying to recover, your runbook is inaccessible when you need it most. * * * Establish feedback loops that drive ongoing improvement to backup operations, configuration, and tooling. Kafka topologies evolve, throughput changes, compliance requirements shift, and new `kafka-backup` releases bring performance improvements. A backup strategy that is never revisited will silently fall behind operational needs. - **Conduct post-incident reviews** after every backup or restore incident. Document root cause, timeline, impact, and action items. Track action item completion. - **Run monthly metrics reviews** covering: - Backup duration trends (are backups taking longer as data volume grows?) - Storage growth rate and cost trajectory - Restore success rate and time-to-restore - RPO/RTO compliance percentage - **Update configurations** when the environment changes: - New topics or topic patterns added to Kafka - Significant throughput increases - New compliance or regulatory requirements - Infrastructure changes (new regions, storage tiers) - **Benchmark against new releases** -- Test new `kafka-backup` versions in staging. Measure throughput, compression ratio, and resource usage against your current version before upgrading production. > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Add a recurring calendar event for a monthly "Backup Operations Review". Use it to walk through metrics dashboards, review open action items, and assess whether current configurations still meet requirements. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **Set-and-forget** -- Deploying a backup configuration once and never reviewing it. Environments change; configurations must follow. > > - **No post-incident reviews** -- Repeating the same failure because the team never analysed the first occurrence. > > - **Annual-only review** -- Reviewing backup strategy once a year guarantees that it is out of date for eleven months. * * * Use these questions to assess your operational maturity. For each question, rate your current state as **None**, **Basic**, **Advanced**, or **Expert**. 1. Is there a designated owner (individual or team) accountable for Kafka backup operations? 2. Are backup schedules, retention policies, and validation steps defined as version-controlled configuration? 3. Do you run automated deep validation (`kafka-backup validate --deep`) after every backup? 4. Are Prometheus metrics enabled and scraped for all `kafka-backup` instances? 5. Do you have Grafana dashboards (or equivalent) providing visibility into backup health, lag, and storage growth? 6. Are alerts configured for backup failures, RPO threshold breaches, and storage errors? 7. Do you maintain copy-paste-ready runbooks for every restore scenario (full cluster, single topic PITR, offset recovery)? 8. Have you executed a full disaster-recovery drill in the last quarter? 9. Do you conduct post-incident reviews after every backup or restore incident? 10. Do you review backup metrics and configurations at least monthly to ensure they match current requirements? * * * - [Deployment Guide](https://kafkabackup.com/deployment.md) -- Infrastructure setup for all supported platforms - [CLI Reference](https://kafkabackup.com/reference/cli-reference.md) -- Complete `kafka-backup` command reference - [Kubernetes Operator](https://kafkabackup.com/operator.md) -- Operator installation, CRDs, and guides - [Metrics Reference](https://kafkabackup.com/reference/metrics.md) -- Full list of Prometheus metrics - [Monitoring Setup Guide](https://kafkabackup.com/guides/monitoring-setup.md) -- Step-by-step Prometheus and Grafana configuration --- title: Security description: Protecting backup data, configurations, and operations through identity management, encryption, access controls, and audit logging for OSO Kafka Backup source_url: html: https://kafkabackup.com/well-architected/security md: https://kafkabackup.com/well-architected/security.md --- # Security > Protecting backup data, configurations, and operations through identity management, encryption, access controls, and audit logging — with the same rigour applied to production Kafka data. Backups contain a complete copy of your Kafka data. A compromised backup is as damaging as a compromised production cluster. The Security pillar ensures that every layer of your backup architecture — from credentials and network paths to storage buckets and audit trails — is hardened, monitored, and regularly reviewed. 1. **Implement strong identity foundation** — Apply the principle of least privilege to every service account, IAM role, and operator that interacts with backup infrastructure. 2. **Enable traceability** — Audit-log every backup, restore, and configuration change so you can answer _who did what, when, where, and with what outcome_. 3. **Apply security at all layers** — Secure data in transit, at rest, at the access layer, and inside configuration files. No single control should be the only line of defence. 4. **Automate security best practices** — Rotate credentials on a schedule, scan configurations for drift, and enforce policies via code rather than manual checklists. 5. **Protect data at rest with encryption** — Encrypt all backup data using customer-managed keys where compliance demands it, and verify encryption status continuously. 6. **Prepare for security events** — Maintain runbooks for credential revocation, backup quarantine, and forensic investigation so the team can respond quickly to incidents. * * * Effective IAM starts with dedicated service accounts scoped to the minimum permissions each operation requires. | Role | Storage Access | Kafka Access | | --- | --- | --- | | **Backup** | Write-only to storage | Read-only from source Kafka | | **Restore** | Read-only from storage | Write to target Kafka | | **Validate** | Read-only from storage | None | > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Use IAM roles (AWS), managed identities (Azure), or workload identity (GCP) instead of static credentials. These provide automatic credential rotation and eliminate the risk of leaked long-lived keys. Enterprise editions support RBAC for multi-team environments, allowing administrators to grant granular permissions per team, topic pattern, or environment. **Backup role** — write-only storage with read-only Kafka: ``` { "Version": "2012-10-17", "Statement": [ { "Sid": "BackupWriteOnly", "Effect": "Allow", "Action": [ "s3:PutObject", "s3:PutObjectAcl", "s3:ListBucket" ], "Resource": [ "arn:aws:s3:::my-kafka-backups", "arn:aws:s3:::my-kafka-backups/*" ] } ] } ``` **Restore role** — read-only storage: ``` { "Version": "2012-10-17", "Statement": [ { "Sid": "RestoreReadOnly", "Effect": "Allow", "Action": [ "s3:GetObject", "s3:ListBucket" ], "Resource": [ "arn:aws:s3:::my-kafka-backups", "arn:aws:s3:::my-kafka-backups/*" ] } ] } ``` > [!CAUTION] > > [!NOTE] > > Anti-Patterns > > [!NOTE] > > - **Admin credentials for backup operations** — Over-privileged accounts increase blast radius if compromised. > > - **Same IAM role for backup and restore** — Violates least privilege; a compromised backup process could overwrite or delete data. > > - **Static access keys** — Long-lived keys are a top cause of cloud security incidents. > > - **No environment separation** — Production and staging sharing credentials or storage accounts. * * * All backup data must be encrypted at rest, with key management appropriate to your compliance requirements. | Cloud | Default | Customer-Managed Key | | --- | --- | --- | | **AWS S3** | SSE-S3 (AES-256) | SSE-KMS / SSE-C | | **Azure Blob** | Microsoft-managed keys | Customer-managed keys (CMK) | | **GCS** | Google-managed keys | Customer-managed encryption keys (CMEK) | Enterprise preserves metadata from an existing Confluent CSFLE deployment: KEKs, encrypted DEKs, encrypted subjects, and Schema Registry encryption rules. It does not encrypt Kafka backup segment files. See [Confluent CSFLE Metadata Backup](https://kafkabackup.com/enterprise/encryption.md) for the supported scope and configuration. > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > For regulated industries (finance, healthcare), use customer-managed keys (CMK/CMEK) with automatic key rotation. Combine with S3 Object Lock or Azure immutable storage to protect against tampering and ransomware. Configure encryption at rest through the storage provider: an S3 default encryption setting or bucket policy, Azure Storage encryption with a customer-managed key, GCS CMEK, or the equivalent control for your backend. Kafka Backup does not expose a client-side segment-encryption key in its backup YAML. > [!CAUTION] > > [!NOTE] > > Anti-Patterns > > [!NOTE] > > - **Unencrypted storage buckets** — Backup data readable by anyone with network access to the storage layer. > > - **Default S3 encryption without key management** — SSE-S3 encrypts data but Amazon manages the keys; insufficient for regulatory requirements that mandate customer-controlled keys. > > - **Public buckets** — S3 buckets or Azure containers with public access enabled, exposing backup data to the internet. * * * Encrypt all data moving between Kafka, the backup process, and cloud storage. - Use TLS for all Kafka connections; **mTLS is preferred** for mutual authentication. - Minimum **TLS 1.2**; prefer **TLS 1.3** where supported. - Validate broker certificates against a trusted CA. HTTPS is the default for S3, Azure Blob, and GCS. Ensure custom or on-premises storage endpoints also enforce TLS. | Cloud | Service | Benefit | | --- | --- | --- | | **AWS** | S3 Gateway Endpoint | Traffic stays within VPC | | **Azure** | Private Endpoint | Private IP for storage account | | **GCP** | Private Google Access | No public IP required | > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > VPC endpoints and private links prevent backup data from traversing the public internet, reducing exposure to man-in-the-middle attacks and data exfiltration. ``` source: bootstrap_servers: - kafka:9093 security: security_protocol: SASL_SSL ssl_ca_location: /certs/ca.crt ssl_certificate_location: /certs/client.crt ssl_key_location: /certs/client.key ``` kafka-backup supports the following security properties: | Property | Description | | --- | --- | | `security_protocol` | `SASL_SSL`, `SSL`, `SASL_PLAINTEXT`, `PLAINTEXT` | | `ssl_ca_location` | Path to CA certificate | | `ssl_certificate_location` | Path to client certificate | | `ssl_key_location` | Path to client private key | > [!CAUTION] > > [!NOTE] > > Anti-Patterns > > [!NOTE] > > - **PLAINTEXT connections** — Data and credentials sent in the clear on the network. > > - **Backup traffic over the public internet** — Without VPC endpoints, data traverses public networks unnecessarily. > > - **Self-signed certificates without a trust chain** — Disabling certificate verification removes protection against MITM attacks. > > - **TLS 1.0 / 1.1** — Deprecated protocols with known vulnerabilities. * * * Credentials must never appear in plain text inside configuration files or source control. | Backend | Use Case | | --- | --- | | **AWS Secrets Manager / SSM Parameter Store** | AWS-native workloads | | **Azure Key Vault** | Azure-native workloads | | **HashiCorp Vault** | Multi-cloud or on-premises | | **Kubernetes Secrets (encrypted with KMS)** | Kubernetes-native deployments | > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Use short-lived credentials wherever possible — AWS STS tokens, federated identities, or OIDC-based workload identity. These expire automatically, limiting the window of exposure if intercepted. kafka-backup supports `${VAR_NAME}` environment variable substitution in configuration files, allowing secrets to be injected at runtime from any secrets backend: ``` source: bootstrap_servers: - ${KAFKA_BOOTSTRAP_SERVERS} security: security_protocol: SASL_SSL sasl_mechanism: ${SASL_MECHANISM} sasl_username: ${SASL_USERNAME} sasl_password: ${SASL_PASSWORD} ``` ``` apiVersion: v1 kind: Secret metadata: name: kafka-backup-credentials namespace: kafka-backup type: Opaque stringData: SASL_USERNAME: backup-service-account SASL_PASSWORD: changeme-use-sealed-secrets AWS_ACCESS_KEY_ID: AKIAIOSFODNN7EXAMPLE AWS_SECRET_ACCESS_KEY: wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY ``` > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > The example above uses `stringData` for readability. In production, use **SealedSecrets**, **External Secrets Operator**, or **Vault Agent** to inject secrets securely rather than storing them directly in Kubernetes manifests committed to Git. > [!CAUTION] > > [!NOTE] > > Anti-Patterns > > [!NOTE] > > - **Plaintext credentials in YAML committed to Git** — Secrets in version control are exposed to anyone with repository access and persist in Git history. > > - **Long-lived static access keys** — Keys that never rotate accumulate risk over time. > > - **Shared credentials across environments** — A compromise in staging exposes production. * * * Maintain a tamper-evident record of every backup operation to support compliance audits and incident investigations. Enterprise editions emit structured audit events capturing: - **Who** — service account or user identity - **What** — operation performed (backup, restore, validate, configure) - **When** — timestamp with timezone - **Where** — source cluster, target storage, topic(s) - **Outcome** — success, failure, partial completion | Cloud | Service | Captures | | --- | --- | --- | | **AWS** | S3 Server Access Logging + CloudTrail | Object-level reads/writes, API calls | | **Azure** | Storage Analytics Logging | Read/write/delete operations | | **GCP** | Cloud Audit Logs | Data access and admin activity | > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Align log retention periods with your compliance framework (SOC 2, PCI-DSS, HIPAA). Configure alerts for anomalous operations such as unexpected restore activity, bulk deletes, or access from unfamiliar IP ranges. ``` audit: enabled: true destination: type: datadog api_key: ${DATADOG_API_KEY} site: datadoghq.eu events: - backup.started - backup.completed - backup.failed - restore.started - restore.completed - restore.failed - config.changed - credentials.rotated retention: days: 365 ``` > [!CAUTION] > > [!NOTE] > > Anti-Patterns > > [!NOTE] > > - **No audit trail for restore operations** — Restores modify production data; without logging, you cannot trace who restored what or diagnose issues. > > - **Audit logs co-located with backups** — An attacker who compromises backup storage can also tamper with logs. Store logs in a separate, append-only destination. > > - **No log retention policy** — Logs that expire before an audit window closes leave gaps in compliance evidence. * * * Backup data often contains personally identifiable information (PII) subject to privacy regulations such as GDPR, CCPA, and HIPAA. Enterprise editions support schema-aware, field-level data masking that integrates with Schema Registry to identify and redact sensitive fields before data is written to storage. Capabilities include: - **Right to be forgotten** — Delete or mask specific records to satisfy GDPR erasure requests. - **Masking policies per topic and field** — Define rules such as _mask `email` in `user-events`_ or _hash `ssn` in `customer-records`_. - **Schema Registry integration** — Automatically detect new fields and apply default masking policies. Separate backup data into tiers based on sensitivity: | Tier | Access | Retention | Content | | --- | --- | --- | --- | | **Full (unmasked)** | Restricted — security/compliance team only | Short (regulatory minimum) | Complete data for disaster recovery | | **Masked** | Broader — development, analytics teams | Longer | PII redacted for safe use in non-production | > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Masking must be applied at backup time, not at restore time. Once unmasked PII is written to storage, it is subject to the same data protection obligations as the original Kafka topic. > [!CAUTION] > > [!NOTE] > > Anti-Patterns > > [!NOTE] > > - **Backing up PII without masking** — Creates a shadow copy of sensitive data outside the controls applied to production systems. > > - **No GDPR deletion process for backups** — Failing to honour erasure requests in backup data is a compliance violation. > > - **Assuming backups are exempt from privacy regulations** — Regulators consider backups as part of the data lifecycle; all obligations apply. * * * Use the following questions during architecture reviews to assess your backup security posture: 1. Are dedicated, least-privilege service accounts used for backup, restore, and validate operations? 2. Is IAM role assumption or workload identity used instead of static access keys? 3. Is all backup data encrypted at rest with appropriate key management (SSE-KMS, CMK, CMEK)? 4. Are all Kafka connections secured with TLS 1.2+ or mTLS? 5. Does backup traffic stay within private network paths (VPC endpoints, private links)? 6. Are secrets managed through a dedicated secrets backend rather than stored in configuration files? 7. Are credentials short-lived and rotated automatically? 8. Is there an audit trail for every backup, restore, and configuration change? 9. Are audit logs stored separately from backup data with appropriate retention policies? 10. Is PII identified, masked, or encrypted before being written to backup storage? * * * - [Security Setup Guide](https://kafkabackup.com/guides/security-setup.md) — Step-by-step TLS, SASL, and encryption configuration - [Confluent CSFLE Metadata Backup (Enterprise)](https://kafkabackup.com/enterprise/encryption.md) — Preserve KEKs, DEKs, subjects, and schema encryption rules - [RBAC (Enterprise)](https://kafkabackup.com/enterprise/rbac.md) — Role-based access control for multi-team environments - [Audit Logging (Enterprise)](https://kafkabackup.com/enterprise/audit-logging.md) — Structured audit events and compliance reporting - [Schema Registry (Enterprise)](https://kafkabackup.com/enterprise/schema-registry.md) — Schema-aware masking and data governance --- title: Reliability description: Ensuring consistent backup and restore operations, graceful failure recovery, and meeting defined RPO and RTO targets under all conditions with OSO Kafka Backup source_url: html: https://kafkabackup.com/well-architected/reliability md: https://kafkabackup.com/well-architected/reliability.md --- # Reliability > The ability to consistently perform backup and restore operations correctly, recover from failures gracefully, and meet defined RPO and RTO targets under all conditions. Reliability is the foundation of any backup system. A backup that cannot be restored is worse than no backup at all — it creates a false sense of safety. The Reliability pillar ensures that every backup is validated, every restore is rehearsed, and every failure scenario has a documented, tested recovery path. 1. **Automatically recover from failure** — Use checkpoints and incremental resume so that a failed backup picks up where it left off rather than starting from scratch. 2. **Test recovery procedures, not just backup creation** — A backup is only as good as its last successful restore. Validate regularly. 3. **Scale horizontally to handle partition growth** — As topics gain partitions or new topics are added, the backup infrastructure must scale without manual intervention. 4. **Manage change through automation** — Use GitOps workflows and Kubernetes CRDs to version, review, and roll out configuration changes predictably. 5. **Design for zero data loss** — Understand and configure your RPO targets explicitly; do not rely on defaults to meet business requirements. 6. **Implement fault isolation** — One partition failure should not affect the backup of other partitions. Isolate failure domains so blast radius is minimised. * * * Every backup must be validated to confirm it can be used for a successful restore. Silent corruption, incomplete segments, or missing offsets can render a backup useless when it is needed most. | Check | Description | | --- | --- | | **Manifest completeness** | All expected topics and partitions are present in the backup manifest | | **Segment integrity** | Checksums match and compressed segments can be decompressed without errors | | **Offset continuity** | No gaps in offset sequences within each partition | | **Time window coverage** | Backup spans the expected time range without holes | Backup operations can report success (exit code 0) while producing incomplete or corrupt output. Network interruptions, storage throttling, or transient Kafka errors can result in partial writes that are only detectable through explicit validation. - Run `kafka-backup validate --deep` after every backup operation. - Automate validation in your CI/CD pipeline or as a post-backup Kubernetes Job. - Periodically perform a full restore-to-temporary-cluster validation to confirm end-to-end recoverability. > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Schedule a weekly automated restore to a temporary cluster. This catches issues that static validation cannot detect, such as schema compatibility problems or consumer group restoration failures. **Deep validation of a backup:** ``` kafka-backup validate \ --path s3://my-kafka-backups/production \ --backup-id 2026-03-24T00-00-00Z \ --deep ``` **Describe backup metadata for programmatic checks:** ``` kafka-backup describe \ --path s3://my-kafka-backups/production \ --backup-id 2026-03-24T00-00-00Z \ --format json ``` > [!CAUTION] > > [!NOTE] > > Anti-Patterns > > [!NOTE] > > - **Assuming success from exit code alone** — A zero exit code means the process completed, not that the backup is complete or uncorrupted. > > - **Never validating until a real restore is needed** — Discovering corruption during a production incident turns a recoverable situation into a crisis. > > - **Validating only the latest backup** — Older backups may have degraded in storage; periodic re-validation catches bit rot and storage issues. * * * Point-in-time recovery (PITR) allows you to restore data to a specific moment, not just the latest backup. This is critical for recovering from data corruption, accidental deletes, or application bugs that produced bad data. Before implementing PITR, answer these questions: - **What granularity is needed?** — Can you tolerate restoring to the nearest hour, or do you need minute-level precision? - **What is the maximum acceptable backup window?** — The gap between the last backup and the failure determines potential data loss. - **Which topics need PITR?** — Not every topic requires the same recovery granularity. | Backup Frequency | Achievable RPO | Use Case | | --- | --- | --- | | **Continuous** | [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Epoch timestamps must be in **milliseconds**, not seconds. A common mistake is using a 10-digit Unix timestamp (seconds) instead of a 13-digit epoch millisecond value, resulting in restoring data from 1970. **PITR restore configuration:** ``` restore: source: path: s3://my-kafka-backups/production backup_id: 2026-03-24T00-00-00Z target: bootstrap_servers: - kafka-restore:9092 options: time_window_start: 1711234800000 # 2026-03-23T15:00:00Z time_window_end: 1711238400000 # 2026-03-23T16:00:00Z topic_mapping: - source: orders target: orders-restored - source: payments target: payments-restored ``` > [!CAUTION] > > [!NOTE] > > Anti-Patterns > > [!NOTE] > > - **Daily backups when RPO is 1 hour** — The backup frequency must match or exceed the RPO requirement. A daily backup cannot deliver an hourly RPO. > > - **Not understanding epoch timestamps** — Misconfigured time windows restore the wrong data range, wasting time during an incident. > > - **No topic-level RPO classification** — Treating all topics with the same backup frequency wastes resources on low-value data and under-protects high-value data. * * * Restoring messages is only half the battle. If consumer offsets are not recovered correctly, applications will either reprocess data (duplicates) or skip data (loss). Offset recovery must be planned as part of every restore operation. kafka-backup supports multiple offset recovery strategies: | Strategy | Description | Best For | | --- | --- | --- | | **Timestamp-based** | Reset offsets to a specific timestamp | PITR restores | | **Offset-based (header)** | Use original offset headers embedded in backup | Exact replay | | **Group-based** | Restore committed consumer group offsets | Resuming existing consumers | | **Cluster-scan** | Scan target cluster to determine appropriate offsets | Cross-cluster migration | | **Manual** | Specify exact offsets per partition | Surgical recovery | Incorrect offset recovery is the most common cause of post-restore issues. Applications may appear healthy but silently skip records or reprocess hours of data, causing downstream inconsistencies. Always use the two-phase approach: **plan** first, then **execute**. 1. Generate a plan to review before applying changes. 2. Review the plan to confirm offsets are correct. 3. Execute the plan to apply offset resets. > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > The plan phase is non-destructive and produces a reviewable output. Always inspect the plan before executing, especially during incident recovery when mistakes are costly. **Generate an offset reset plan:** ``` kafka-backup offset-reset plan \ --path s3://kafka-backups/production \ --backup-id production-20260719 \ --groups order-service,payment-service \ --bootstrap-servers kafka-dr:9092 \ --format json > offset-plan.json ``` **Review the plan:** ``` cat offset-plan.json | jq '.partitions[] | {topic, partition, current_offset, new_offset}' ``` **Execute the offset reset:** ``` kafka-backup offset-reset execute \ --path s3://kafka-backups/production \ --backup-id production-20260719 \ --groups order-service,payment-service \ --bootstrap-servers kafka-dr:9092 ``` > [!CAUTION] > > [!NOTE] > > Anti-Patterns > > [!NOTE] > > - **Restoring data without considering consumer offsets** — Messages are restored but consumers either skip them entirely or reprocess old data. > > - **Blindly resetting to earliest or latest** — `earliest` causes full reprocessing; `latest` skips all restored data. Neither is appropriate without understanding the restore context. > > - **Skipping the plan phase** — Executing offset resets without review during a high-pressure incident leads to compounding errors. * * * A disaster recovery plan defines how you will restore Kafka data when the worst happens. Without explicit RPO and RTO targets per data tier, recovery is ad-hoc and unpredictable. Classify topics into tiers based on business impact and assign RPO/RTO targets: | Tier | Examples | RPO | RTO | | --- | --- | --- | --- | | **Tier 1 — Critical** | Payments, orders, financial transactions | [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > DR documentation must be accessible during an outage. If your runbooks are stored on the same infrastructure that has failed, you cannot access them when you need them most. Keep copies in at least two independent locations. > [!CAUTION] > > [!NOTE] > > Anti-Patterns > > [!NOTE] > > - **No defined RPO/RTO targets** — Without targets, there is no way to measure whether your backup strategy is adequate or whether recovery was successful. > > - **DR plan exists but has never been tested** — An untested plan is an assumption, not a plan. > > - **DR documentation stored only on the failing system** — Wiki on the same cloud region, runbooks in the same Kubernetes cluster, or playbooks in the same Git hosting provider. > > - **Single-person DR knowledge** — If only one engineer knows how to restore, your RTO depends on their availability. * * * Backup infrastructure must be isolated from the systems it protects. A failure that takes down your Kafka cluster should not also take down your ability to restore from backup. | Isolation Domain | Recommendation | | --- | --- | | **Storage location** | Different account, region, or cloud provider from the source cluster | | **Compute** | Dedicated backup infrastructure, not co-located on broker nodes | | **Network** | Separate failure domain; backup process reachable even if source network is degraded | | **Partition-level** | Per-partition fault isolation — one partition failure does not block others | If backups are stored in the same region and account as the source cluster, a single cloud incident (region outage, account compromise, IAM misconfiguration) can destroy both production data and all backups simultaneously. - Enable **S3 cross-region replication** to maintain backup copies in a separate region. - Run backup processes on **dedicated infrastructure**, not on Kafka broker nodes. - Use **per-partition checkpointing** so a failure in one partition allows others to continue. - Implement **checkpoint-based resume** so interrupted backups pick up where they left off. - For ransomware protection, maintain **air-gapped backups** using S3 Object Lock or equivalent immutable storage. > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Use a separate AWS account for backup storage. Even if the production account is compromised, the backup account remains isolated. Cross-account IAM roles provide secure access without shared credentials. **Cross-region S3 storage with replication:** ``` storage: type: s3 s3: bucket: my-kafka-backups-dr region: us-west-2 # Different region from source cluster endpoint: "" force_path_style: false backup: checkpoint_interval: 30s # Frequent checkpoints for resume per_partition_isolation: true ``` > [!CAUTION] > > [!NOTE] > > Anti-Patterns > > [!NOTE] > > - **Backups in the same region and account as the source cluster** — A region-wide outage or account compromise destroys both production data and backups. > > - **Running backup processes on Kafka broker nodes** — Broker failure takes down both the data source and the backup process simultaneously. > > - **No redundancy for backup storage** — A single storage location is a single point of failure. > > - **Single point of failure in backup infrastructure** — One backup server, one storage bucket, one network path. * * * A disaster recovery plan is only reliable if it is tested regularly. Chaos engineering validates that your backup infrastructure handles real-world failure scenarios, not just ideal conditions. | Test Type | Frequency | Scope | | --- | --- | --- | | **Tabletop exercise** | Monthly | Walk through DR scenarios with the team; identify gaps in runbooks | | **Single topic restore** | Weekly (automated) | Restore one topic to a temporary cluster and validate integrity | | **Full cluster restore** | Quarterly | Restore all tiered topics and measure actual RTO | | **Failover drill** | Semi-annually | Simulate primary cluster loss and execute full DR procedure | | **Chaos test** | Quarterly | Inject failures into backup infrastructure and observe behaviour | Without regular testing, your DR plan degrades over time. Infrastructure changes, new topics, configuration drift, and team turnover all erode recovery capabilities. Testing keeps them current. Design chaos scenarios that exercise your failure modes: | Scenario | What It Tests | | --- | --- | | **Kill backup process mid-run** | Checkpoint resume — does backup continue from last checkpoint? | | **Storage outage (revoke S3 access)** | Error handling and retry logic | | **Network partition (block Kafka port)** | Graceful degradation and reconnection | | **Corrupt a backup segment** | Validation detection — does `validate --deep` catch it? | | **Full region failure** | Cross-region restore from replicated backup | > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Always run chaos tests in a controlled environment with clear rollback procedures. Document the blast radius of each test before executing. Never run destructive chaos tests against production backups without an isolated copy. Every DR test must produce a written report including: - **Actual RTO achieved** vs target RTO - **Actual RPO achieved** vs target RPO - **Issues encountered** during the test - **Action items** with owners and deadlines - **Pass/fail determination** against defined success criteria > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Track RTO and RPO measurements over time. Trending data reveals whether your recovery capabilities are improving or degrading, and provides evidence for compliance audits. > [!CAUTION] > > [!NOTE] > > Anti-Patterns > > [!NOTE] > > - **DR drills only after incidents** — Reactive testing means you discover problems during real outages, not before them. > > - **Testing in non-production-like environments** — A DR test against a cluster with 10 partitions does not validate recovery of a production cluster with 10,000 partitions. > > - **Undocumented test results** — Without written records, lessons are lost, the same issues recur, and compliance auditors have no evidence of testing. > > - **Never testing full-scale restore** — Single-topic restores build confidence but do not validate that your infrastructure can handle a complete cluster recovery within the RTO window. * * * Use the following questions during architecture reviews to assess the reliability of your backup strategy: 1. Are all backups validated automatically after completion using `kafka-backup validate --deep`? 2. Is a full restore-to-temporary-cluster test performed at least quarterly? 3. Are RPO and RTO targets defined and documented for every topic tier? 4. Does the backup frequency match or exceed the RPO requirement for each tier? 5. Is consumer offset recovery planned and tested as part of every restore procedure? 6. Are backups stored in a different failure domain (region, account, or cloud provider) from the source cluster? 7. Does the backup process use per-partition fault isolation and checkpoint-based resume? 8. Is there a documented, tested DR procedure with clear escalation and communication steps? 9. Are DR drills conducted at least quarterly, with results documented and action items tracked? 10. Are chaos engineering scenarios used to validate backup infrastructure resilience? * * * - [Point-in-Time Recovery Guide](https://kafkabackup.com/guides/restore-pitr.md) — Step-by-step PITR restore procedures - [Offset Management Guide](https://kafkabackup.com/guides/offset-management.md) — Consumer offset recovery strategies and workflows - [CLI Reference](https://kafkabackup.com/reference/cli-reference.md) — Complete command reference for backup, restore, and validate operations - [PITR Architecture](https://kafkabackup.com/architecture/pitr-implementation.md) — Technical deep-dive into point-in-time recovery implementation - [Multi-Cluster DR Examples](https://kafkabackup.com/examples/multi-cluster-dr.md) — Example configurations for cross-region and multi-cluster disaster recovery --- title: Performance Efficiency description: Efficient use of compute, storage, and network resources to meet backup and restore throughput requirements with OSO Kafka Backup source_url: html: https://kafkabackup.com/well-architected/performance-efficiency md: https://kafkabackup.com/well-architected/performance-efficiency.md --- # Performance Efficiency > "Using compute, storage, and network resources efficiently to meet backup and restore throughput requirements, and maintaining that efficiency as data volumes grow." Kafka backup workloads must keep pace with production data rates while consuming the minimum necessary resources. The Performance Efficiency pillar ensures your backup architecture is right-sized, properly tuned, and continuously benchmarked — so that growing data volumes never outstrip your ability to protect them. 1. **Right-size for throughput requirements** — Match compute, memory, and network capacity to measured data rates rather than guessing. Over-provisioning wastes budget; under-provisioning causes backup lag. 2. **Use compression to reduce storage and network overhead** — Compress backup segments to shrink storage footprint and reduce network transfer time, choosing an algorithm that balances ratio against CPU cost. 3. **Benchmark before deploying to production** — Validate throughput, latency, and resource consumption under realistic load before going live. Assumptions about performance are not a substitute for measurement. 4. **Monitor performance continuously and act on trends** — Track throughput, compression ratios, and resource utilisation over time. Identify degradation early and address it before it becomes an incident. 5. **Go serverless / managed where possible** — Prefer managed object storage (S3, GCS, Azure Blob) over self-managed block storage. Managed services scale automatically and eliminate storage infrastructure overhead. 6. **Experiment with configuration in staging before production** — Test tuning changes (segment sizes, compression levels, concurrency) in a staging environment that mirrors production topology and data characteristics. * * * Establish clear targets for your environment and validate them through benchmarking: | Metric | Target | Notes | | --- | --- | --- | | **Throughput per partition** | 100+ MB/s | Depends on network, storage backend, and message size | | **Checkpoint latency (p99)** | [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Start with `segment_max_bytes: 134217728` (128 MB) and increase to 256 MB only after confirming that your storage backend handles larger PUT requests without timeout issues. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **Using default configuration for high-throughput topics** — Defaults are conservative. Production workloads almost always benefit from tuning segment size, fetch size, and concurrency. > > - **Deploying backup in a different region from Kafka** — Cross-region data transfer adds latency, cost, and fragility. > > - **Using very small segments ( > - **No benchmarking before production** — Performance assumptions based on development data volumes are unreliable. * * * | Algorithm | Compression Ratio | Speed | CPU Usage | Best For | | --- | --- | --- | --- | --- | | **zstd** | 3--5x | Fast | Moderate | General-purpose default | | **lz4** | 2--3x | Very fast | Low | Latency-sensitive workloads | | **none** | 1x | Fastest | None | Pre-compressed or encrypted data | | Data Format | Typical Ratio (zstd) | Notes | | --- | --- | --- | | **JSON** | 5--8x | Highly repetitive structure compresses well | | **Avro** | 2--4x | Binary format, moderate compressibility | | **Protobuf** | 2--3x | Compact binary, less headroom for compression | | **Already compressed** | [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Levels 1--3 offer the best throughput-to-ratio trade-off for real-time backup. Only increase beyond 3 if you have measured a meaningful improvement in your specific data and can tolerate the additional CPU overhead. Monitor the `kafka_backup_compression_ratio` metric to detect changes in data compressibility over time. A sudden drop in ratio may indicate a change in message format or the introduction of pre-compressed payloads. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **Compressing already-compressed data** — Wastes CPU cycles for negligible size reduction. Disable compression for topics containing gzip, snappy, or encrypted payloads. > > - **Using maximum compression levels for latency-sensitive workloads** — High zstd levels (10+) can add significant CPU time per segment, increasing flush latency. > > - **Never benchmarking compression** — Different data formats yield vastly different ratios. Benchmark your actual payloads to choose the right algorithm and level. * * * | Backend | Throughput | Durability | Relative Cost | Notes | | --- | --- | --- | --- | --- | | **Amazon S3** | High | 11 9's | $$ | Most mature, widest tooling ecosystem | | **Azure Blob Storage** | High | 17 9's (GRS) | $$ | Native integration with Azure workloads | | **Google Cloud Storage** | High | 11 9's | $$ | Strong multi-region replication | | **MinIO** | High | Depends on deployment | $ | S3-compatible, self-hosted | | **Filesystem** | Very high | Depends on underlying storage | Free | No network overhead, limited durability | - **Use object storage** — S3, Azure Blob, or GCS provide the durability, scalability, and lifecycle management required for production backup data. - **Enable versioning** — Protect against accidental overwrites or deletions. Versioning also supports compliance requirements for immutable backups. - **Use Transfer Acceleration for cross-region** — When restoring data to a different region, enable S3 Transfer Acceleration or equivalent to reduce transfer time. - **Filesystem or MinIO for local development** — Avoids cloud costs and network dependencies during development and testing. - **Mirror production configuration** — Use the same segment sizes, compression settings, and directory structures so that performance characteristics transfer to production. Deploy VPC endpoints (AWS), private endpoints (Azure), or Private Service Connect (GCP) to keep backup traffic on the cloud provider's backbone network. This eliminates internet gateway latency and reduces data transfer costs. **Amazon S3:** ``` storage: type: s3 bucket: my-kafka-backups region: eu-west-1 # Use VPC endpoint for private network access endpoint: https://s3.eu-west-1.amazonaws.com ``` **Azure Blob Storage:** ``` storage: type: azure container: my-kafka-backups account_name: mystorageaccount ``` **Filesystem (development):** ``` storage: type: filesystem base_path: /var/lib/kafka-backup/data ``` > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **Using filesystem storage for production** — Filesystem storage lacks the durability guarantees, lifecycle management, and cross-region replication of object storage. > > - **No versioning on backup buckets** — A single accidental deletion or overwrite can destroy your recovery capability. > > - **Storing backups on a different continent from Kafka** — Intercontinental data transfer adds significant latency and cost to every backup segment write. > > - **No VPC endpoints** — Routing backup traffic through the public internet adds latency, increases cost, and exposes data to unnecessary network hops. * * * | Partition Count | CPU | Memory | Network | | --- | --- | --- | --- | | **1--4** | 1 vCPU | 512 MB | 1 Gbps | | **5--16** | 2 vCPU | 1 GB | 5 Gbps | | **17--64** | 4 vCPU | 2 GB | 10 Gbps | | **65+** | Scale horizontally | — | — | > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > OSO Kafka Backup is written in Rust and is designed to use less than 500 MB of memory for 4 partitions. The sizing table above includes headroom for bursts and garbage-free operation. Set both requests and limits to ensure predictable scheduling and prevent noisy-neighbour effects: ``` apiVersion: apps/v1 kind: Deployment metadata: name: kafka-backup spec: replicas: 1 selector: matchLabels: app: kafka-backup template: metadata: labels: app: kafka-backup spec: containers: - name: kafka-backup image: ghcr.io/osodevops/kafka-backup:latest resources: requests: cpu: "1" memory: "512Mi" limits: cpu: "2" memory: "1Gi" volumeMounts: - name: config mountPath: /etc/kafka-backup volumes: - name: config configMap: name: kafka-backup-config ``` Scale based on backup lag to ensure partitions are covered as throughput increases: ``` apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: kafka-backup-hpa spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: kafka-backup minReplicas: 1 maxReplicas: 8 metrics: - type: Pods pods: metric: name: kafka_backup_lag_bytes target: type: AverageValue averageValue: "104857600" # 100 MB average lag ``` > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **No resource limits** — Without limits, a misbehaving backup instance can consume all node resources and affect other workloads. > > - **No resource requests** — Without requests, the Kubernetes scheduler cannot make informed placement decisions, leading to resource contention. > > - **Vertical scaling only** — A single large instance has a failure blast radius that affects all partitions. Horizontal scaling isolates failures. > > - **Running backup on Kafka broker nodes** — Backup I/O competes with broker I/O for disk and network bandwidth, degrading both. * * * Deploy backup instances in the **same availability zone** as the Kafka brokers they read from. This provides the lowest latency and eliminates cross-AZ data transfer charges. > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Cross-AZ data transfer in AWS costs $0.01/GB in each direction. For a topic producing 1 TB/day, this adds approximately $600/month in transfer costs alone. - **VPC endpoints** — Use gateway or interface VPC endpoints for S3 and other storage services to keep traffic off the public internet. - **Rack awareness** — Configure `client.rack` so that Kafka directs fetch requests to the closest replica, reducing cross-rack and cross-AZ traffic. Increase `fetch_max_bytes` to reduce the number of fetch round trips: ``` kafka: fetch_max_bytes: 52428800 # 50 MB fetch_max_wait_ms: 500 ``` > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Larger fetch sizes amortise the per-request overhead but increase memory usage. Monitor memory consumption when tuning fetch sizes upward. When restoring data to a different region: - Enable **S3 Transfer Acceleration** or equivalent service to optimise long-distance transfers. - Consider pre-staging backup data to the target region using storage replication before initiating the restore. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **Backup in a different region from Kafka** — Adds latency to every fetch request and segment upload, reducing throughput and increasing costs. > > - **Small fetch sizes with high-throughput topics** — Causes excessive round trips to brokers, wasting network capacity on request overhead. > > - **Shared network without QoS** — Backup traffic can saturate shared links, affecting production Kafka clients. > > - **No consideration for cross-AZ costs** — Multi-AZ deployments are resilient but expensive. Understand the cost trade-off and optimise placement accordingly. * * * Use the [kafka-backup-demos](https://github.com/osodevops/kafka-backup-demos) benchmark suite to measure performance under controlled conditions. ``` # Clone the benchmark suite git clone https://github.com/osodevops/kafka-backup-demos.git cd kafka-backup-demos # Run the throughput benchmark ./benchmarks/run-throughput-test.sh --partitions 4 --message-size 1024 --duration 300 ``` Run the following scenarios to build a complete performance profile: | Scenario | Purpose | | --- | --- | | **Max throughput per partition** | Establish single-partition ceiling | | **Multi-partition scaling** | Validate linear scaling with partition count | | **Large messages (> 1 MB)** | Identify buffer and timeout issues | | **Restore speed** | Measure time-to-recovery for capacity planning | | **WAN latency simulation** | Understand cross-region performance impact | For each scenario, record: - **Throughput** — MB/s sustained over the test duration - **Duration** — Total time to back up or restore the test dataset - **Resource utilisation** — CPU, memory, network, and disk I/O during the test Store baselines in version control alongside your backup configuration so they can be compared over time. - After **version upgrades** of `kafka-backup` - After **configuration changes** to segment size, compression, or concurrency - After **infrastructure changes** such as instance type, network topology, or storage backend - **Quarterly** as a routine check against baseline drift > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Benchmarks run on small data volumes ( [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **No benchmarking before production** — Deploying without performance validation is deploying blind. Every environment has unique characteristics that affect throughput. > > - **Benchmarking with unrealistically small data volumes** — Small datasets fit entirely in OS page cache, producing misleadingly high throughput numbers. > > - **Benchmarking backup only, not restore** — Restore performance is equally critical and often has different bottlenecks (e.g., Kafka producer throughput to the target cluster). > > - **Benchmarking on different hardware than production** — Results from a developer laptop do not predict production performance. * * * Use the following questions during architecture reviews to assess performance efficiency: 1. Have you established throughput targets (MB/s per partition) and validated them through benchmarking? 2. Is compression enabled, and have you chosen the algorithm and level based on your data format? 3. Is the storage backend appropriate for your durability, throughput, and cost requirements? 4. Are resource requests and limits set for all backup workloads running in Kubernetes? 5. Are backup instances co-located with Kafka brokers in the same availability zone? 6. Are VPC endpoints or private network paths used for storage access? 7. Have you benchmarked restore performance in addition to backup performance? 8. Do you re-run benchmarks after configuration changes, version upgrades, and on a quarterly schedule? 9. Is horizontal scaling used instead of vertical scaling for high partition counts? 10. Are you monitoring `kafka_backup_compression_ratio`, throughput, and resource utilisation continuously? * * * - [Performance Tuning Guide](https://kafkabackup.com/guides/performance-tuning.md) — Step-by-step tuning for throughput and latency - [Configuration Reference](https://kafkabackup.com/reference/config-yaml.md) — Complete YAML configuration options - [Metrics Reference](https://kafkabackup.com/reference/metrics.md) — Full list of Prometheus metrics for performance monitoring - [Compression Architecture](https://kafkabackup.com/architecture/compression.md) — How compression is implemented in the backup pipeline - [Zero-Copy Optimisation](https://kafkabackup.com/architecture/zero-copy-optimization.md) — Rust zero-copy design for minimal overhead --- title: Cost Optimisation description: Managing and reducing the total cost of Kafka backup operations through efficient resource utilisation, smart storage tiering, and cost-aware architecture decisions source_url: html: https://kafkabackup.com/well-architected/cost-optimisation md: https://kafkabackup.com/well-architected/cost-optimisation.md --- # Cost Optimisation > "Managing and reducing the total cost of Kafka backup operations through efficient resource utilisation, smart storage tiering, and cost-aware architecture decisions." Backup infrastructure can become a significant and often invisible cost centre. Without active cost management, storage grows unbounded, oversized compute runs around the clock, and network transfer charges accumulate unnoticed. The Cost Optimisation pillar ensures you achieve reliable data protection at the lowest reasonable cost by continuously measuring, right-sizing, and governing every component of your backup architecture. 1. **Implement cloud financial management** — Tag every backup resource, track spending against budgets, and make cost data visible to the teams that control it. 2. **Adopt a consumption model** — Pay only for what you use. Scale compute to match backup windows, use lifecycle policies to tier storage, and avoid paying for idle capacity. 3. **Measure overall efficiency** — Define a unit cost metric such as cost per GB backed up. Track it over time and use it to evaluate architecture changes. 4. **Stop spending on undifferentiated heavy lifting** — Use managed object storage (S3, GCS, Azure Blob) rather than self-hosted storage. Let the cloud provider handle durability, availability, and scaling. 5. **Analyse and attribute expenditure** — Break down backup costs by team, environment, and topic. Attribution drives accountability and surfaces optimisation opportunities. 6. **Right-size retention to actual business need** — Not all data needs the same retention period. Match retention to compliance, operational, and business requirements rather than applying a single blanket policy. * * * Minimise storage costs through lifecycle policies, compression, deduplication, and continuous monitoring — without compromising data durability or restore capability. Storage is typically the largest component of backup cost. A single unmanaged S3 bucket can grow from manageable to expensive within months. Lifecycle policies alone can reduce long-term storage costs by 95% compared to keeping everything in the default storage class. **AWS S3 Lifecycle Tiers** | Age | Storage Class | Approx. Cost (USD/GB/mo) | Use Case | | --- | --- | --- | --- | | 0–30 days | S3 Standard | $0.023 | Active backups, frequent restores | | 31–90 days | S3 Standard-IA | $0.0125 | Infrequent access, still fast retrieval | | 91–365 days | S3 Glacier Instant Retrieval | $0.004 | Archival with millisecond access | | 1+ years | S3 Glacier Deep Archive | $0.00099 | Long-term compliance, rare access | **Azure Blob Storage Tiers** | Age | Access Tier | Use Case | | --- | --- | --- | | 0–30 days | Hot | Active backups | | 31–90 days | Cool | Infrequent access | | 91–365 days | Cold | Archival with moderate retrieval time | | 1+ years | Archive | Long-term compliance | **GCS Storage Classes** | Age | Storage Class | Use Case | | --- | --- | --- | | 0–30 days | Standard | Active backups | | 31–90 days | Nearline | Monthly access pattern | | 91–365 days | Coldline | Quarterly access pattern | | 1+ years | Archive | Annual access or compliance | **Compression and deduplication:** - Enable compression in `kafka-backup` for a typical **3–5x reduction** in stored data size - Use deduplication to eliminate redundant segments across incremental backups - Automate deletion of backups beyond the retention policy with operator-managed retention or storage lifecycle rules **Understand total cost components:** - Storage at rest (the largest component) - API calls (PUT, GET, LIST operations) - Data transfer (retrieval and cross-region) - Compute (the backup process itself) > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Enable S3 Intelligent-Tiering for buckets where access patterns are unpredictable. It automatically moves objects between tiers based on access frequency, with no retrieval fees for the frequent and infrequent access tiers. **S3 Lifecycle Policy:** ``` { "Rules": [ { "ID": "kafka-backup-lifecycle", "Status": "Enabled", "Filter": { "Prefix": "backups/" }, "Transitions": [ { "Days": 30, "StorageClass": "STANDARD_IA" }, { "Days": 90, "StorageClass": "GLACIER_IR" }, { "Days": 365, "StorageClass": "DEEP_ARCHIVE" } ], "Expiration": { "Days": 2555 } } ] } ``` **kafka-backup compression configuration:** ``` backup: compression: enabled: true algorithm: zstd level: 3 storage: s3: bucket: my-kafka-backups region: eu-west-1 storage-class: STANDARD ``` > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **All Standard, forever** — Storing every backup in S3 Standard with no lifecycle policy. A 10 TB dataset costs ~$230/month in Standard vs ~$10/month in Deep Archive. > > - **No lifecycle policies** — Relying on manual cleanup that never happens. Storage grows linearly and silently. > > - **No cost monitoring** — Discovering a $5,000/month storage bill during quarterly budget review instead of catching it at $500. > > - **No compression** — Storing uncompressed Kafka segments when 3–5x compression is available with minimal CPU overhead. * * * Match compute resources to actual backup workload requirements, scaling up for backup windows and down during idle periods. Backup workloads are inherently bursty. Running large instances 24/7 for a workload that peaks during a two-hour backup window wastes 90% of the compute spend. - **Start with PE-04 sizing guidance** — Use the performance efficiency pillar's sizing recommendations as a baseline, then refine based on observed utilisation. - **Monitor actual utilisation** — Track CPU, memory, and network usage during backup runs. If peak utilisation is below 50%, you are over-provisioned. - **Right-size with 20–30% headroom** — Allow enough capacity to handle spikes without throttling, but no more. **Kubernetes requests and limits:** ``` resources: requests: cpu: "500m" memory: "512Mi" limits: cpu: "2000m" memory: "2Gi" ``` - **Use Spot/Preemptible instances for non-critical workloads** — Validation jobs and development backups tolerate interruption. Enable checkpointing so interrupted jobs resume rather than restart. ``` # Kubernetes node affinity for Spot instances affinity: nodeAffinity: preferredDuringSchedulingIgnoredDuringExecution: - weight: 80 preference: matchExpressions: - key: node.kubernetes.io/instance-type operator: In values: - spot ``` - **Schedule scale-up during backup windows** — Use Kubernetes CronJobs or cluster autoscaler profiles to add capacity before backup runs and release it afterwards. - **Leverage Rust efficiency** — `kafka-backup` is written in Rust, which delivers significantly lower CPU and memory consumption compared to JVM-based alternatives. This translates directly into smaller instance sizes and lower compute costs. > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Combine Spot instances with checkpointing for development and validation workloads. If a Spot instance is reclaimed, the job resumes from the last checkpoint rather than restarting from scratch. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **Large instances 24/7 for a 2-hour daily backup** — Running an `m5.4xlarge` around the clock when a `m5.xlarge` during the backup window would suffice. > > - **No utilisation monitoring** — Provisioning based on initial estimates and never revisiting. Workloads change; compute allocation should too. > > - **Ignoring Spot/Preemptible** — Paying full on-demand price for fault-tolerant workloads that can run on Spot at 60–90% discount. * * * Minimise data transfer charges by co-locating backup components, using private endpoints, and compressing data before transmission. Cloud providers charge for data that crosses availability zone, region, or internet boundaries. For high-throughput Kafka clusters, transfer costs can rival or exceed storage costs if not managed carefully. **Typical transfer pricing (AWS):** | Path | Approx. Cost (USD/GB) | | --- | --- | | Same AZ (private IP) | Free | | Cross-AZ | $0.01 | | Cross-region | $0.02–$0.09 | | Internet egress | $0.05–$0.12 | - **Co-locate in the same AZ** — Run `kafka-backup` in the same availability zone as your Kafka brokers. This eliminates cross-AZ charges for the data read path. ``` # Pod topology constraint for same-AZ placement topologySpreadConstraints: - maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: DoNotSchedule labelSelector: matchLabels: app: kafka-backup ``` - **Use VPC endpoints** — Access S3, GCS, or Azure Blob via private endpoints rather than public internet. This eliminates NAT gateway charges and reduces latency. ``` # Create S3 VPC Gateway Endpoint aws ec2 create-vpc-endpoint \ --vpc-id vpc-abc123 \ --service-name com.amazonaws.eu-west-1.s3 \ --route-table-ids rtb-abc123 ``` - **Backup locally, then replicate** — Write backups to a storage bucket in the same region as Kafka. Use storage-native cross-region replication (e.g., S3 CRR) for DR copies — this is cheaper and more reliable than backing up directly to a remote region. - **Compression reduces transfer cost proportionally** — A 4x compression ratio means 75% less data traverses the network, reducing transfer charges by the same proportion. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **Direct cross-region backup** — Streaming backup data from `eu-west-1` Kafka to a `us-east-1` bucket, paying cross-region transfer on every byte. > > - **No VPC endpoints** — Routing S3 traffic through a NAT gateway, paying $0.045/GB for NAT processing on top of transfer charges. > > - **No transfer cost accounting** — Tracking storage costs but ignoring the transfer charges that can exceed them. * * * Define retention periods per topic or data classification, automating the deletion of backups that are no longer needed. Keeping all backups indefinitely is the single largest driver of runaway storage costs. A well-designed retention policy reduces storage by 60–80% while still meeting every compliance and operational requirement. **Retention tiers by data classification:** | Tier | Example Topics | Retention | Justification | | --- | --- | --- | --- | | Compliance | `financial-transactions`, `audit-log` | 7 years | Regulatory requirement (e.g., SOX, MiFID II) | | Critical | `orders`, `payments`, `customer-updates` | 90 days | Operational recovery and dispute resolution | | Standard | `user-events`, `page-views`, `click-stream` | 30 days | Analytics replay, short-term debugging | | Ephemeral | `logs`, `metrics`, `health-checks` | 7 days | Troubleshooting only, easily regenerated | **Automate with operator retention or storage lifecycle policies:** For Kubernetes operator deployments, `KafkaBackup.spec.retention` can prune complete backup sets after successful backup runs. It is disabled by default and must be enabled per backup: ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: production-nightly spec: schedule: "0 0 2 * * * *" stopAtCurrentOffsets: true storage: storageType: s3 s3: bucket: kafka-backups region: eu-west-1 prefix: production/nightly credentialsSecret: name: s3-credentials retention: enabled: true maxAgeDays: 90 keepLast: 7 dryRun: true ``` Use `dryRun: true` first, review the reported eligible backups and reclaimed bytes, then switch to `dryRun: false` after the policy has been approved. Use storage prefixes per tier when retention is enforced by the storage backend: ``` { "Rules": [ { "ID": "financial-archive-7-years", "Status": "Enabled", "Filter": { "Prefix": "financial/" }, "Expiration": { "Days": 2555 } }, { "ID": "production-90-days", "Status": "Enabled", "Filter": { "Prefix": "production/" }, "Expiration": { "Days": 90 } }, { "ID": "ephemeral-7-days", "Status": "Enabled", "Filter": { "Prefix": "ephemeral/" }, "Expiration": { "Days": 7 } } ] } ``` - **Use operator retention for backup-set pruning** — For PVC/local, S3/S3-compatible, and Azure Blob Storage, the operator can delete whole backup IDs without partially pruning manifests. - **Use storage lifecycle for backend-native controls** — Configure S3, S3-compatible, Azure Blob, or GCS lifecycle policies when you need object lock, legal hold, cross-account enforcement, or GCS support. - **Audit retention quarterly** — Review which topics are being backed up, how much storage each consumes, and whether the retention period is still appropriate. - **Document justification** — Record why each retention period was chosen. Compliance requirements change; undocumented policies cannot be reviewed. - **Support legal hold** — Ensure your retention automation can be overridden for specific backups when legal hold is required (e.g., litigation or regulatory investigation). > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Automated deletion is irreversible. Before enabling retention policies, verify that your compliance team has signed off on the retention periods for regulated data. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **Keep everything forever** — The most expensive and least compliant approach. Indefinite retention increases both cost and risk surface. > > - **Manual cleanup** — Relying on an engineer to periodically delete old backups. This never happens consistently. > > - **Same retention for all topics** — Applying a 7-year retention to ephemeral logs because "it's easier than classifying topics". * * * Implement tagging, dashboards, budgets, and review processes that make backup costs transparent and actionable. You cannot optimise what you cannot see. Without visibility, backup costs are absorbed into general cloud spend, optimisation opportunities go unnoticed, and there is no accountability for cost growth. **Tag all resources consistently:** ``` # Standard tagging schema tags: team: platform-engineering environment: production project: kafka-backup cost-centre: CC-4521 managed-by: terraform ``` **Build cost dashboards that show:** - Total backup cost per month (trend over 6+ months) - Breakdown by category: storage, compute, network - Cost per GB backed up (efficiency metric) - Cost by environment (production vs staging vs development) - Month-over-month growth rate **Set alerts and budgets:** ``` # AWS Budget example aws budgets create-budget \ --account-id 123456789012 \ --budget '{ "BudgetName": "kafka-backup-monthly", "BudgetLimit": {"Amount": "500", "Unit": "USD"}, "TimeUnit": "MONTHLY", "BudgetType": "COST", "CostFilters": { "TagKeyValue": ["user:project$kafka-backup"] } }' \ --notifications-with-subscribers '[{ "Notification": { "NotificationType": "ACTUAL", "ComparisonOperator": "GREATER_THAN", "Threshold": 80 }, "Subscribers": [{ "SubscriptionType": "EMAIL", "Address": "platform-team@example.com" }] }]' ``` - **Monthly cost review** — Include backup costs in your team's monthly operational review. Compare actual spend against budget and investigate variances. - **Cost per topic** — Where possible, attribute storage costs to individual topics. This surfaces topics that are disproportionately expensive and drives conversations about retention and compression. > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Use AWS Cost Explorer tag filtering, Azure Cost Management scopes, or GCP billing labels to create dedicated backup cost views without building custom dashboards. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **No tagging** — Backup resources are untagged, making it impossible to separate backup costs from general infrastructure spend. > > - **Lumped into general spend** — Backup costs are not broken out, so no one knows whether the $2,000/month increase is from backups, compute, or something else entirely. > > - **No budget or alerts** — The team discovers cost overruns during quarterly business review instead of when they happen. * * * Use these questions to evaluate the cost optimisation of your Kafka backup architecture: 1. Do you have lifecycle policies configured on all backup storage buckets? 2. Can you state the cost per GB of backed-up data for each environment? 3. Are compute resources right-sized to actual backup workload utilisation, with no more than 30% idle headroom? 4. Are you using Spot or Preemptible instances for non-critical backup workloads? 5. Is `kafka-backup` co-located in the same availability zone as the Kafka brokers it reads from? 6. Are VPC endpoints configured for all storage access paths? 7. Do you have documented and automated retention policies for every backed-up topic? 8. Are all backup resources tagged with a consistent schema (team, environment, project, cost-centre)? 9. Do you have budget alerts that fire before costs exceed your planned spend? 10. Is backup cost reviewed as a standing agenda item in your monthly operational review? * * * - [AWS S3 Pricing](https://aws.amazon.com/s3/pricing/) - [AWS Data Transfer Pricing](https://aws.amazon.com/ec2/pricing/on-demand/#Data_Transfer) - [Azure Blob Storage Pricing](https://azure.microsoft.com/en-gb/pricing/details/storage/blobs/) - [Google Cloud Storage Pricing](https://cloud.google.com/storage/pricing) - [AWS Cost Explorer](https://aws.amazon.com/aws-cost-management/aws-cost-explorer/) - [Azure Cost Management](https://azure.microsoft.com/en-gb/products/cost-management/) - [GCP Billing Reports](https://cloud.google.com/billing/docs/reports) - [kafka-backup Configuration Reference](https://kafkabackup.com/reference/config-yaml.md) --- title: Sustainability description: Minimising the environmental impact of Kafka backup operations through efficient resource utilisation, data lifecycle management, and considered infrastructure choices source_url: html: https://kafkabackup.com/well-architected/sustainability md: https://kafkabackup.com/well-architected/sustainability.md --- # Sustainability > "Minimising the environmental impact of Kafka backup operations through efficient resource utilisation, data lifecycle management, and considered infrastructure choices." Every compute cycle, stored byte, and network transfer has an environmental cost. While individual backup workloads are modest in isolation, the cumulative impact across environments, regions, and retention periods adds up. The Sustainability pillar helps you reduce this footprint by making deliberate choices about how, where, and how long you store and process backup data. 1. **Understand your impact** — Measure the carbon footprint of your backup infrastructure using cloud-provider tools. You cannot reduce what you do not measure. 2. **Maximise utilisation** — Right-size compute resources, use efficient compression algorithms, and eliminate idle capacity. Higher utilisation means less wasted energy per unit of useful work. 3. **Adopt more efficient technology** — `kafka-backup` is written in Rust, which delivers significantly lower CPU and memory consumption compared to JVM-based alternatives. Choosing inherently efficient tooling reduces energy consumption at the source. 4. **Reduce downstream impact** — Efficient storage tiering moves data to lower-energy cold storage over time, reducing the ongoing energy required to maintain backups. * * * Minimise the compute, memory, and energy consumed per unit of backed-up data by right-sizing resources, using efficient algorithms, and eliminating waste. Over-provisioned infrastructure consumes energy whether it is doing useful work or not. A right-sized, efficiently compressed backup pipeline can deliver the same data protection with a fraction of the environmental footprint. - **Rust is inherently efficient** — `kafka-backup` is written in Rust, which compiles to native machine code with no garbage collection overhead. This means significantly lower CPU and memory consumption compared to JVM-based backup tools, translating directly into smaller instances and lower energy usage. - **Right-size compute resources** — Follow the guidance in [PE-04](https://kafkabackup.com/well-architected/performance-efficiency.md) to size instances based on actual workload. Monitor utilisation and downsize when headroom exceeds 30%. ``` # Right-sized resource allocation resources: requests: cpu: "500m" memory: "512Mi" limits: cpu: "2000m" memory: "2Gi" ``` - **Enable autoscaling** — Scale compute to match backup windows rather than running at peak capacity 24/7. Use Kubernetes Horizontal Pod Autoscaler or cluster autoscaler to add and remove capacity dynamically. - **Use efficient compression** — `zstd` compression reduces stored data by 3–5x with minimal CPU overhead, reducing both storage energy and transfer energy. ``` backup: compression: enabled: true algorithm: zstd level: 3 ``` - **Choose regions with lower carbon intensity** — Where latency and data residency requirements allow, prefer regions powered by renewable energy. > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > The combination of Rust's efficiency, right-sized compute, and zstd compression can reduce the energy footprint of your backup pipeline by 5–10x compared to a naively provisioned JVM-based alternative. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **Large instances running 24/7** — Running oversized compute around the clock for a workload that runs for two hours per day wastes 90% of the energy consumed. > > - **No compression** — Storing and transferring uncompressed data when efficient compression is available at negligible CPU cost. > > - **No region consideration** — Deploying backup infrastructure based solely on latency without considering the carbon intensity of the region's energy grid. * * * Back up only the data you need, retain it only as long as required, and match recovery granularity to actual business requirements. Every byte stored consumes energy — for storage media, cooling, and redundancy. Reducing the volume of stored data through selective backup, appropriate retention, and right-sized granularity directly reduces the ongoing energy cost of your backup infrastructure. - **Implement retention policies** — Follow the tiered retention guidance in [CO-04](https://kafkabackup.com/well-architected/cost-optimisation.md) to ensure data is deleted when it is no longer needed. Every deleted backup is energy that no longer needs to be spent on storage. - **Use topic filtering** — Back up only the topics that require protection. Exclude ephemeral topics, internal Kafka topics, and any topic whose data can be trivially regenerated. ``` backup: topics: include: - "orders.*" - "payments.*" - "customer.*" exclude: - ".*\\.internal" - "logs\\.debug.*" - "__consumer_offsets" ``` - **Match PITR granularity to actual needs** — Point-in-time recovery with minute-level granularity generates far more backup data than hourly granularity. Choose the granularity your RTO and RPO actually require, not the finest granularity available. - **Archive to cold storage tiers** — Move older backups to cold storage (Glacier, Archive, Coldline). Cold storage tiers consume less energy per byte than hot storage because they use denser, less frequently accessed media. > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Before excluding topics from backup, confirm with application owners that the data is genuinely ephemeral or regenerable. Excluding a topic that turns out to be critical is not recoverable. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **Backing up all topics indiscriminately** — Including internal Kafka topics, debug logs, and ephemeral streams that have no recovery value. > > - **No retention policies** — Storing backups indefinitely, consuming energy for data that will never be accessed again. > > - **Over-specifying granularity** — Configuring minute-level PITR for a workload where hourly recovery is perfectly acceptable, generating 60x more backup data than necessary. * * * Choose cloud regions and storage tiers that minimise the carbon intensity of your backup infrastructure. Cloud regions differ significantly in carbon intensity depending on the local energy grid. A backup stored in a region powered primarily by renewables has a materially lower carbon footprint than the same backup in a coal-heavy region — with identical durability and availability. **Prefer regions with renewable energy:** | Provider | Lower-Carbon Regions | | --- | --- | | AWS | `eu-west-1` (Ireland), `eu-north-1` (Stockholm) | | Azure | North Europe (Ireland), Sweden Central | | GCP | `europe-north1` (Finland), `us-central1` (Iowa) | > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > AWS, Azure, and GCP all publish sustainability commitments and region-level carbon data. Use this information when choosing where to deploy backup infrastructure, especially for DR copies that are latency-insensitive. - **Use cold storage for long-term backups** — Cold and archive storage tiers use denser storage media that consumes less energy per byte. For backups older than 90 days that are rarely accessed, cold storage is both cheaper and more sustainable. - **Track your carbon footprint** — Use cloud-provider tools to measure the emissions attributable to your backup infrastructure: ``` # AWS: View Carbon Footprint in the Billing Console # Navigate to: AWS Billing > Carbon Footprint # Azure: View emissions via Emissions Dashboard # Navigate to: Azure Portal > Carbon Optimization # GCP: View Carbon Footprint in the Console # Navigate to: GCP Console > Carbon Footprint ``` - **Set sustainability targets** — Track carbon emissions per GB of backed-up data as a sustainability KPI. Review quarterly and set reduction targets aligned with your organisation's sustainability commitments. > [!CAUTION] > > [!NOTE] > > Anti-patterns > > [!NOTE] > > - **No carbon consideration in region selection** — Choosing regions based solely on cost or latency without evaluating carbon intensity. > > - **All hot storage, all the time** — Keeping every backup in hot/standard storage tiers when the vast majority will never be accessed after 30 days. > > - **No emissions tracking** — Operating backup infrastructure without any visibility into its environmental impact. * * * Use these questions to evaluate the sustainability of your Kafka backup architecture: 1. Do you know the carbon intensity of the regions where your backup infrastructure runs? 2. Are compute resources right-sized to actual workload utilisation, avoiding persistent over-provisioning? 3. Is compression enabled to reduce both storage volume and transfer energy? 4. Do you have retention policies that automatically delete backups when they are no longer needed? 5. Are you using cloud-provider carbon footprint tools to track and report on the emissions of your backup infrastructure? * * * - [AWS Customer Carbon Footprint Tool](https://aws.amazon.com/aws-cost-management/aws-customer-carbon-footprint-tool/) - [Azure Emissions Dashboard](https://azure.microsoft.com/en-gb/blog/empowering-cloud-sustainability-with-the-microsoft-emissions-dashboard/) - [Google Carbon Footprint](https://cloud.google.com/carbon-footprint) - [The Green Web Foundation](https://www.thegreenwebfoundation.org/) - [Cloud Carbon Footprint (open source)](https://www.cloudcarbonfootprint.org/) --- title: Cross-Cutting Concerns description: Architecture considerations that span multiple pillars of the Well-Architected Framework source_url: html: https://kafkabackup.com/well-architected/cross-cutting-concerns md: https://kafkabackup.com/well-architected/cross-cutting-concerns.md --- # Cross-Cutting Concerns > Architecture considerations that span multiple pillars of the Well-Architected Framework — topics that affect security, reliability, cost, operational excellence, and performance simultaneously. Some architectural decisions do not fit neatly into a single pillar. They cut across every dimension of the Well-Architected Framework and must be addressed holistically. This page covers the most important cross-cutting concerns for organisations running OSO Kafka Backup in production. * * * Many organisations operate Kafka across multiple clouds or hybrid on-prem/cloud environments. Your backup strategy must account for heterogeneous infrastructure, differing credential models, and the realities of cross-cloud networking. - **Storage portability** — kafka-backup supports S3, Azure Blob, GCS, and local filesystem. Backups created in one cloud can be restored in another. - **Credential management across clouds** — Each cloud has its own identity model (IAM roles, managed identities, workload identity). Backup configs must handle these differences. - **Network connectivity and latency** — Cross-cloud restores introduce network hops and potential bandwidth constraints. - **Data sovereignty and residency requirements** — Regulations may restrict where backup data can be stored or transferred. - **Cost of cross-cloud data transfer** — Egress charges can be significant when replicating backups between providers. - **Back up locally, replicate cross-cloud for DR** — Perform primary backups to the same cloud as the source cluster, then replicate to a secondary cloud for disaster recovery. - **Use S3-compatible storage (MinIO) as a universal intermediate format** — MinIO provides an S3-compatible API that runs on any cloud or on-prem, giving you a consistent storage interface. - **Consistent config across clouds using GitOps** — Store backup configurations in Git and deploy them identically across environments to reduce drift. > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Start with local backups to minimise latency and cost, then add cross-cloud replication as a second stage. This avoids paying egress fees on every backup cycle. Multi-cloud backup with primary S3 and secondary Azure Blob for DR: ``` # Primary backup — S3 in AWS (same region as source Kafka) backup: name: production-primary source: bootstrap-servers: "${KAFKA_BOOTSTRAP_SERVERS}" security-protocol: SASL_SSL sasl-mechanism: SCRAM-SHA-512 sasl-username: "${KAFKA_USERNAME}" sasl-password: "${KAFKA_PASSWORD}" storage: type: s3 bucket: "${AWS_BACKUP_BUCKET}" region: "${AWS_REGION}" prefix: kafka-backup/production topics: include: - ".*" --- # Secondary backup — Azure Blob for cross-cloud DR backup: name: production-dr source: bootstrap-servers: "${KAFKA_BOOTSTRAP_SERVERS}" security-protocol: SASL_SSL sasl-mechanism: SCRAM-SHA-512 sasl-username: "${KAFKA_USERNAME}" sasl-password: "${KAFKA_PASSWORD}" storage: type: azure-blob container: "${AZURE_BACKUP_CONTAINER}" account-name: "${AZURE_STORAGE_ACCOUNT}" account-key: "${AZURE_STORAGE_KEY}" prefix: kafka-backup/production topics: include: - ".*" ``` > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Cross-cloud credential environment variables must be managed carefully. Use a secrets manager (Vault, AWS Secrets Manager, Azure Key Vault) rather than storing credentials directly in config files or CI/CD pipelines. * * * kafka-backup provides a Kubernetes operator with Custom Resource Definitions (CRDs) for GitOps-native backup management. This allows you to declare backup and restore jobs as Kubernetes resources, managed alongside your application manifests. | CRD | Purpose | | --- | --- | | **KafkaBackup** | Defines a backup job — source cluster, storage target, schedule, and topic filters | | **KafkaRestore** | Defines a restore job — source backup, target cluster, and restore parameters | | **KafkaOffsetReset** | Manages consumer offset recovery after a restore operation | | **KafkaOffsetRollback** | Rolls back offset changes if a reset produces unexpected results | - **Store CRDs in Git alongside application manifests** — Backup definitions should live in the same repository as the services that produce and consume the data. - **Use ArgoCD or Flux for GitOps deployment** — Automate CRD deployment through your existing GitOps pipeline. - **Define resource requests and limits** — Prevent backup pods from starving other workloads or being OOM-killed during large backups. - **Use Kubernetes RBAC to control who can create restore CRDs** — Restores are destructive operations; limit access to authorised personnel. - **Monitor CRD status with kubectl and Prometheus** — The operator exposes metrics and CRD status conditions for observability. > [!NOTE] > > [!NOTE] > > info > > [!NOTE] > > The Kubernetes operator watches for CRD changes and reconciles the desired state automatically. This means you can trigger a backup or restore simply by applying a manifest — no imperative commands required. **KafkaBackup CRD:** ``` apiVersion: kafka.oso.dev/v1alpha1 kind: KafkaBackup metadata: name: production-daily namespace: kafka-backup spec: schedule: "0 2 * * *" source: bootstrapServers: kafka-cluster-kafka-bootstrap:9093 securityProtocol: SASL_SSL saslMechanism: SCRAM-SHA-512 credentialsSecret: name: kafka-backup-credentials storage: type: s3 bucket: my-backup-bucket region: eu-west-1 prefix: production/daily topics: include: - ".*" exclude: - "__.*" resources: requests: cpu: 500m memory: 1Gi limits: cpu: "2" memory: 4Gi ``` **KafkaRestore CRD:** ``` apiVersion: kafka.oso.dev/v1alpha1 kind: KafkaRestore metadata: name: restore-production-20260324 namespace: kafka-backup spec: backup: name: production-daily snapshot: "2026-03-24T02:00:00Z" target: bootstrapServers: kafka-cluster-kafka-bootstrap:9093 securityProtocol: SASL_SSL saslMechanism: SCRAM-SHA-512 credentialsSecret: name: kafka-restore-credentials topics: include: - "orders.*" - "payments.*" restoreOffsets: true resources: requests: cpu: "1" memory: 2Gi limits: cpu: "4" memory: 8Gi ``` > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Use `kubectl get kafkabackup` and `kubectl get kafkarestore` to check the status of backup and restore operations. The operator sets status conditions such as `Ready`, `Running`, `Completed`, and `Failed`. **See also:** [Operator Overview](https://kafkabackup.com/operator.md), [KafkaBackup CRD](https://kafkabackup.com/operator/crds/kafkabackup.md), [KafkaRestore CRD](https://kafkabackup.com/operator/crds/kafkarestore.md), [GitOps Guide](https://kafkabackup.com/operator/guides/gitops.md) * * * For clusters using a Schema Registry (Confluent, Apicurio), backup must include schemas alongside topic data. Without schemas, consumers cannot deserialise restored messages, and producers cannot validate new messages against the expected format. - **Schema IDs may not be preserved across clusters** — Schema IDs are auto-incremented integers assigned by the registry. A restore to a different cluster will likely produce different IDs. - **Schema evolution history should be backed up** — Consumers may depend on older schema versions for backward compatibility. - **Restore must handle schema ID remapping** — Messages reference schema IDs in their headers. After restore, these IDs must map to the correct schemas in the target registry. > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Restoring topic data without its associated schemas will result in deserialisation failures for all Avro, Protobuf, or JSON Schema consumers. Always include schema backup in your DR plan. The Enterprise edition provides integrated Schema Registry backup and restore with: - **Automatic ID remapping** — Schema IDs in restored messages are updated to match the target registry. - **Compatibility validation** — Schemas are validated against the target registry's compatibility settings before restore. - **Full evolution history** — All schema versions and their metadata are preserved. - **Back up Schema Registry independently** — Export schemas via the Schema Registry REST API as a supplementary backup. - **Use Enterprise for integrated schema backup/restore** — The Enterprise edition handles the complexity of ID remapping and compatibility checks automatically. - **Test schema compatibility after restore** — Verify that consumers can deserialise messages and producers can register new schemas. **See also:** [Schema Registry (Enterprise)](https://kafkabackup.com/enterprise/schema-registry.md) * * * Kafka Streams applications maintain local state stores backed by changelog topics. Backup and restore of a Streams application requires special consideration to ensure the application can recover its state correctly. - **Changelog topics must be included in backup** — These topics are the source of truth for Streams state stores. Without them, the application must reprocess all input data from scratch. - **State store rebuild time after restore** — Even with changelog topics restored, state stores must be rebuilt locally. Factor this time into your RTO calculations. - **Repartition topics may need to be excluded** — Repartition topics are intermediate topics generated by Streams. They can be regenerated from input data and do not need to be backed up. - **Consumer offset recovery is critical for Streams apps** — Streams applications use consumer offsets to track processing progress. Incorrect offsets can cause duplicate processing or data loss. - **Include all changelog topics in backup scope** — Use topic name patterns to capture changelog topics (typically suffixed with `-changelog`). - **Exclude repartition topics** — Repartition topics (typically suffixed with `-repartition`) regenerate automatically and waste storage if backed up. - **Test Streams app recovery as part of DR drills** — Streams recovery is more complex than simple consumer recovery. Validate it regularly. - **Use PITR to restore to a consistent state across all related topics** — Point-in-time recovery ensures that input topics, changelog topics, and output topics are restored to the same logical point. > [!NOTE] > > [!NOTE] > > info > > [!NOTE] > > Kafka Streams state store rebuild time depends on the volume of data in the changelog topics. For large state stores, this can take minutes to hours. Plan accordingly and consider standby replicas to reduce recovery time. Topic filtering for Kafka Streams applications — include changelogs, exclude repartition topics: ``` backup: name: streams-app-backup source: bootstrap-servers: kafka-cluster:9092 storage: type: s3 bucket: my-backup-bucket prefix: streams-app topics: include: # Input topics - "orders\\..*" - "payments\\..*" # Output topics - "enriched-orders" - "order-summaries" # Changelog topics (state stores) - "streams-app-.*-changelog" exclude: # Repartition topics (will regenerate) - "streams-app-.*-repartition" # Internal Streams topics - "__consumer_offsets" - "__transaction_state" ``` > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Use a naming convention for your Streams application ID (e.g., `streams-app-*`) so that changelog and repartition topics can be easily identified with wildcard patterns. **See also:** [Kafka Streams Example](https://kafkabackup.com/examples/kafka-streams.md) * * * Industries such as finance, healthcare, and retail have specific regulatory requirements for data backup, retention, and protection. kafka-backup can be configured to meet these requirements, with Enterprise features providing additional compliance capabilities. | Regulation | Requirement | kafka-backup Feature | | --- | --- | --- | | **GDPR** | Right to be forgotten, data minimisation | Data masking, field-level redaction (Enterprise) | | **SOX** | Financial data retention (7 years) | Long-term retention with lifecycle policies | | **HIPAA** | PHI protection, access logging | Storage-provider encryption at rest and access controls | | **PCI DSS** | Cardholder data protection | Existing Confluent CSFLE plus Enterprise metadata backup | | **DORA** | IT system resilience testing | DR testing framework, RTO/RPO tracking | - **Define compliance requirements per topic** — Not all topics carry regulated data. Tag topics with their compliance classification and apply appropriate backup policies. - **Preserve existing encryption metadata** — If applications use Confluent CSFLE, use the Enterprise metadata backup to capture KEKs, encrypted DEKs, and schema encryption rules alongside Kafka data. - **Implement audit logging for all operations** — Every backup, restore, and configuration change should be logged with the operator identity, timestamp, and outcome. - **Conduct regular compliance audits** — Periodically review backup configurations, retention policies, and access controls against regulatory requirements. - **Maintain evidence of DR testing for auditors** — Regulators such as those enforcing DORA require documented evidence that disaster recovery procedures have been tested. > [!WARNING] > > [!NOTE] > > warning > > [!NOTE] > > Regulatory non-compliance can result in significant fines and reputational damage. Treat compliance requirements as hard constraints, not aspirational goals. If in doubt, consult your compliance or legal team before finalising backup configurations. > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Use the Enterprise audit logging feature to generate compliance reports automatically. These reports can be exported in formats suitable for external auditors and regulators. **See also:** [Audit Logging (Enterprise)](https://kafkabackup.com/enterprise/audit-logging.md), [Encryption (Enterprise)](https://kafkabackup.com/enterprise/encryption.md), [RBAC (Enterprise)](https://kafkabackup.com/enterprise/rbac.md), [Compliance Audit Use Case](https://kafkabackup.com/use-cases/compliance-audit.md) --- title: Reference Architectures description: Proven deployment patterns for Kafka backup across single-region, cross-region, multi-cloud, air-gapped, and GitOps scenarios source_url: html: https://kafkabackup.com/well-architected/reference-architectures md: https://kafkabackup.com/well-architected/reference-architectures.md --- # Reference Architectures > Proven deployment patterns that combine the principles from every Well-Architected pillar into end-to-end, production-ready configurations you can adopt or adapt. Each reference architecture below includes a complete topology, configuration, cost estimate, and known limitations so you can evaluate trade-offs before committing to a design. Pick the architecture closest to your constraints, then adjust RPO, RTO, and storage tiers to match your specific requirements. * * * | Architecture | RPO | RTO | Complexity | Est. Monthly Cost | Best For | | --- | --- | --- | --- | --- | --- | | **1\. Single-Region S3** | [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Start with **Architecture 1** to validate your backup strategy, then evolve toward cross-region or multi-cloud patterns as your availability requirements grow. Each architecture builds on the configuration patterns established in the simpler designs. * * * The simplest production-ready pattern. A single kafka-backup deployment runs continuously inside the same region as your Kafka cluster, streaming data to an S3 bucket with versioning enabled. Prometheus scrapes the built-in metrics endpoint for alerting and dashboards. - Single-region Kafka deployment - **RPO [!NOTE] > > [!NOTE] > > info > > [!NOTE] > > **COMPLIANCE mode** prevents anyone — including the root user — from deleting or overwriting objects before the retention period expires. Use **GOVERNANCE mode** if you need the ability to override with special permissions during testing. ``` { "Version": "2012-10-17", "Statement": [ { "Sid": "DenyAllDeleteOperations", "Effect": "Deny", "Principal": "*", "Action": [ "s3:DeleteObject", "s3:DeleteObjectVersion", "s3:PutBucketPolicy", "s3:DeleteBucketPolicy" ], "Resource": [ "arn:aws:s3:::my-org-kafka-backup-airgap", "arn:aws:s3:::my-org-kafka-backup-airgap/*" ], "Condition": { "StringNotEquals": { "aws:PrincipalArn": "arn:aws:iam::111111111111:root" } } }, { "Sid": "AllowWriteFromProductionAccount", "Effect": "Allow", "Principal": { "AWS": "arn:aws:iam::222222222222:role/kafka-backup-transfer-role" }, "Action": [ "s3:PutObject", "s3:ListBucket" ], "Resource": [ "arn:aws:s3:::my-org-kafka-backup-airgap", "arn:aws:s3:::my-org-kafka-backup-airgap/*" ] } ] } ``` | Item | Monthly Cost | | --- | --- | | Primary backup (Architecture 1) | ~$105 | | Air-gapped S3 storage (Glacier + Object Lock) | ~$90 | | **Total** | **~$195** | > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Use S3 Intelligent-Tiering or lifecycle policies to transition older backups to Glacier Deep Archive after 90 days. This can reduce air-gapped storage costs by up to 70% for long-retention requirements. - Higher RTO due to the air gap — restoring requires transferring data back from the isolated account - Transfer scheduling adds complexity (S3 Batch operations, DataSync jobs) - MFA-protected root account access is required for emergency operations in the air-gapped account - Object Lock retention cannot be shortened once set in COMPLIANCE mode - Testing restores from the air-gapped account requires careful planning to avoid violating the air gap * * * A fully declarative, Kubernetes-native approach where backup and restore operations are managed through Custom Resource Definitions (CRDs) and reconciled by a GitOps controller such as ArgoCD or Flux. All configuration lives in a Git repository, providing version history, peer review, and automated rollout for every change. - Your team already operates a Kubernetes platform with GitOps tooling - **RPO [!NOTE] > > [!NOTE] > > info > > [!NOTE] > > With GitOps, every configuration change goes through a pull request. This gives you a full audit trail, peer review, and the ability to roll back any change by reverting a commit. | Item | Monthly Cost | | --- | --- | | Base backup infrastructure (Architecture 1) | ~$105 | | GitOps tooling (ArgoCD/Flux — typically already deployed) | ~$0 | | **Total** | **~$105** | - Requires Kubernetes and GitOps expertise on the team - Operator learning curve — custom resources add an abstraction layer - CRD schema changes require careful upgrade planning - ArgoCD/Flux must be operational for configuration changes to propagate (backup continues running if GitOps is temporarily down) * * * > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > Use the comparison table at the top of this page as a starting point. Then consider these questions: > > > > 1. **What is your RPO/RTO budget?** If > 2. **Do you need multi-cloud protection?** Architecture 3 is the only option that survives a full cloud provider outage. > > 3. **Are you in a regulated industry?** Architecture 4 (Air-Gapped) provides the immutability guarantees auditors look for. > > 4. **Is your team already running GitOps?** Architecture 5 adds minimal overhead and maximum auditability. > > 5. **Just getting started?** Architecture 1 is the fastest path to a working, production-grade backup. All architectures can be combined. For example, you can run Architecture 5 (GitOps) as your deployment model while using Architecture 2 (Cross-Region) as your storage topology and Architecture 4 (Air-Gapped) as an additional compliance layer. --- title: Self-Assessment Checklist description: Score your Kafka backup architecture maturity across all six pillars of the Well-Architected Framework source_url: html: https://kafkabackup.com/well-architected/self-assessment md: https://kafkabackup.com/well-architected/self-assessment.md --- # Self-Assessment Checklist Use this checklist to evaluate the maturity of your Kafka backup architecture across all six pillars. Score each item honestly — the goal is to identify improvement areas, not to achieve a perfect score on day one. Rate each item on a 0–3 scale: | Score | Level | Description | | --- | --- | --- | | **0** | Not implemented | No action taken | | **1** | Basic | Partially implemented, manual processes | | **2** | Advanced | Fully implemented, mostly automated | | **3** | Expert | Fully automated, continuously improved, measured | | Total Score | Maturity | Action | | --- | --- | --- | | **0–25** | Critical gaps | Address immediately — your backup infrastructure has significant risk | | **26–50** | Developing | Create a prioritised improvement plan targeting the lowest-scoring pillars | | **51–70** | Mature | Focus on optimisation and automation of remaining manual processes | | **71–87** | Well-Architected | Maintain through continuous improvement and regular reassessment | > [!TIP] > > [!NOTE] > > Reassess Regularly > > [!NOTE] > > Re-run this assessment quarterly, or after significant changes to your Kafka environment (new topics, increased throughput, new compliance requirements). Track your score over time to measure improvement. * * * | # | Check | Score (0–3) | | --- | --- | --- | | 1 | Backup operations have a designated owner with clear escalation paths | | | 2 | Backup schedules are fully automated (no manual runs required) | | | 3 | Monitoring and alerting covers all key backup metrics (lag, throughput, errors, checkpoint age) | | | 4 | DR runbooks exist with exact `kafka-backup` CLI commands and have been tested | | | 5 | All backup configuration is version-controlled and deployed via GitOps or CI/CD | | **Pillar subtotal: \_\_\_ / 15** * * * | # | Check | Score (0–3) | | --- | --- | --- | | 6 | Least-privilege IAM policies are enforced for backup and restore processes separately | | | 7 | All backup data is encrypted at rest (SSE or client-side encryption) | | | 8 | All connections are encrypted in transit (TLS 1.2+ for Kafka, HTTPS for storage) | | | 9 | No hardcoded credentials — all secrets managed via a secrets manager or environment variables | | | 10 | Audit logging is enabled for all backup and restore operations | | **Pillar subtotal: \_\_\_ / 15** * * * | # | Check | Score (0–3) | | --- | --- | --- | | 11 | Backup integrity is validated automatically after every run (`kafka-backup validate --deep`) | | | 12 | RPO and RTO targets are defined per topic tier and documented | | | 13 | Consumer offset recovery has been tested and is part of the restore procedure | | | 14 | DR drills are conducted at least quarterly with documented results | | | 15 | Backup storage is geographically separated from the primary Kafka cluster | | **Pillar subtotal: \_\_\_ / 15** * * * | # | Check | Score (0–3) | | --- | --- | --- | | 16 | Backup throughput has been benchmarked and meets RPO requirements | | | 17 | Compression algorithm and level have been optimised for your data formats | | | 18 | kafka-backup is co-located with Kafka brokers (same AZ/region) | | | 19 | Compute resources are right-sized based on measured utilisation | | | 20 | Restore performance has been benchmarked and meets RTO requirements | | **Pillar subtotal: \_\_\_ / 15** * * * | # | Check | Score (0–3) | | --- | --- | --- | | 21 | Storage lifecycle policies are active (tiering from Standard → IA → Glacier) | | | 22 | Retention policies are defined per topic tier and enforced automatically | | | 23 | Backup costs are tracked, tagged, and attributed to teams or projects | | | 24 | VPC endpoints are used for storage access (no public internet transfer costs) | | | 25 | Compute is right-sized and scales down when not actively backing up | | **Pillar subtotal: \_\_\_ / 15** * * * | # | Check | Score (0–3) | | --- | --- | --- | | 26 | Compute resources scale down or terminate when not in use | | | 27 | Topic filtering excludes unnecessary topics from backup | | | 28 | Cold storage tiers are used for long-term retention | | | 29 | Compression is enabled to reduce storage and network resource consumption | | **Pillar subtotal: \_\_\_ / 12** * * * | Pillar | Score | | --- | --- | | Operational Excellence | \_\_\_ / 15 | | Security | \_\_\_ / 15 | | Reliability | \_\_\_ / 15 | | Performance Efficiency | \_\_\_ / 15 | | Cost Optimisation | \_\_\_ / 15 | | Sustainability | \_\_\_ / 12 | | **Total** | **\_\_\_ / 87** | * * * Based on your score, prioritise improvements in the lowest-scoring pillars: 1. **Identify the pillar with the lowest score** — this is your highest-risk area 2. **Review the corresponding pillar page** for detailed best practices and implementation guidance 3. **Start with the highest-impact, lowest-effort items** — typically monitoring (OE-03), encryption at rest (SEC-02), and backup validation (REL-01) 4. **Set a target score** for your next quarterly assessment 5. **Track progress** over time and celebrate improvements > [!NOTE] > > [!NOTE] > > Need Help? > > [!NOTE] > > If your assessment reveals critical gaps, the [Reference Architectures](https://kafkabackup.com/well-architected/reference-architectures.md) provide proven deployment patterns you can adopt. For Enterprise features like encryption, RBAC, and audit logging, [contact OSO](https://oso.sh) for a consultation. --- title: Glossary description: Definitions of key terms used throughout the Well-Architected Framework source_url: html: https://kafkabackup.com/well-architected/glossary md: https://kafkabackup.com/well-architected/glossary.md --- # Glossary Definitions of key terms used throughout the OSO Kafka Backup Well-Architected Framework documentation. | Term | Definition | | --- | --- | | **ACL** | Access Control List. A set of rules in Apache Kafka that define which users or service accounts are permitted to perform specific operations (read, write, describe) on topics, consumer groups, and other resources. | | **Air-Gapped Backup** | A backup stored in a location that is physically or logically isolated from the primary environment, preventing compromise of both primary data and backups in a single incident. | | **Audit Log** | A chronological record of operations performed by OSO Kafka Backup, including who initiated each action, what was affected, and the outcome. Available in the Enterprise edition. | | **Backup** | A durable copy of Kafka topic data and metadata stored in external object storage, created by OSO Kafka Backup for disaster recovery and compliance purposes. | | **Backup ID** | A unique identifier assigned to each backup run, used to reference and manage specific backup snapshots. | | **Backup Window** | The time period during which a backup operation runs. Shorter backup windows reduce the risk of data loss but may require more resources. | | **Bootstrap Servers** | A comma-separated list of Kafka broker addresses (host:port) used by clients to establish an initial connection to the Kafka cluster and discover the full cluster topology. | | **Broker** | A Kafka server that stores topic partitions and serves client requests. A Kafka cluster consists of one or more brokers. | | **Chaos Engineering** | The discipline of experimenting on a system to build confidence in its ability to withstand turbulent conditions in production, such as simulating broker failures or network partitions. | | **Checkpoint** | A record of the last successfully committed offset for each topic-partition, stored in a local SQLite database. Checkpoints enable incremental backups and crash-resilient resume. | | **Consumer Group** | A named group of Kafka consumers that coordinate to consume messages from one or more topics, with each partition assigned to exactly one consumer in the group. | | **CRD** | Custom Resource Definition. A Kubernetes extension mechanism used by the OSO Kafka Backup Operator to define custom resources such as `KafkaBackup`, `KafkaRestore`, `KafkaOffsetReset`, `KafkaOffsetRollback`, and `KafkaBackupValidation`. | | **Customer-Managed Key (CMK)** | An encryption key owned and managed by the customer (rather than the cloud provider) used for encrypting backup data at rest, providing full control over key lifecycle and access. | | **Data Masking** | The process of obfuscating or redacting sensitive fields within Kafka messages during backup, ensuring that personally identifiable information (PII) is not stored in plain text. Available in the Enterprise edition. | | **DR Drill** | Disaster Recovery Drill. A planned exercise that tests the end-to-end restore process, validating that backups are viable and that the team can meet RTO and RPO targets. | | **Encryption at Rest** | Protection of stored data using encryption algorithms (e.g., AES-256) so that data on disk or in object storage is unreadable without the decryption key. | | **Encryption in Transit** | Protection of data as it moves between systems using TLS, ensuring that data exchanged between Kafka brokers, backup tools, and storage backends cannot be intercepted. | | **Full Backup** | A backup that captures all messages in the configured topics from the earliest available offset through to the current offset. Contrast with incremental backup. | | **GCS** | Google Cloud Storage. An object storage service from Google Cloud Platform, supported as a backup storage backend by OSO Kafka Backup. | | **GitOps** | An operational model where the desired state of infrastructure and applications is declared in Git repositories, with automated tooling (e.g., ArgoCD, Flux) reconciling the live state to match. | | **Grafana** | An open-source observability platform used to visualise Prometheus metrics from OSO Kafka Backup through pre-built dashboards. | | **IAM** | Identity and Access Management. Cloud provider services (AWS IAM, Azure RBAC, GCP IAM) that control which identities can access storage buckets, encryption keys, and other resources. | | **Incremental Backup** | A backup that captures only messages produced since the last checkpoint, reducing backup duration and storage consumption compared to a full backup. | | **ISR (In-Sync Replicas)** | The set of partition replicas that are fully caught up with the leader replica. A message is considered committed only when all ISR members have acknowledged it. | | **KRaft** | Kafka Raft. The consensus protocol that replaces ZooKeeper for Kafka cluster metadata management, available from Kafka 3.3 and the default from Kafka 4.0. | | **Kubernetes Operator** | A software extension to Kubernetes that uses CRDs and custom controllers to manage the lifecycle of OSO Kafka Backup resources, including scheduling, monitoring, and reconciliation. | | **Lifecycle Policy** | A storage backend rule that automatically transitions or deletes objects based on age. Used to manage backup retention by moving older backups to cheaper storage tiers or expiring them. | | **Manifest** | A JSON file (`manifest.json`) stored at the root of a backup that contains metadata about the backup, including topics, partitions, offset ranges, and timestamps. | | **MinIO** | An open-source, S3-compatible object storage system that can serve as a self-hosted backup storage backend for OSO Kafka Backup. | | **mTLS** | Mutual TLS. A TLS configuration where both the client and server present certificates and verify each other's identity, providing stronger authentication than one-way TLS. | | **Object Storage** | A storage architecture that manages data as objects (with metadata and a unique identifier) rather than as files in a hierarchy. Examples include S3, GCS, and Azure Blob Storage. | | **Offset** | A sequential integer assigned to each message within a Kafka partition, uniquely identifying the message's position. Offsets are used to track consumer progress and enable point-in-time recovery. | | **Partition** | A subdivision of a Kafka topic that provides parallelism. Each partition is an ordered, immutable sequence of messages, and each message within a partition has a unique offset. | | **Point-in-Time Recovery (PITR)** | The ability to restore Kafka topic data to any arbitrary timestamp by filtering backed-up messages based on their timestamps. | | **Prometheus** | An open-source monitoring and alerting toolkit used to collect and query metrics exposed by OSO Kafka Backup on its metrics endpoint (default port 8080). | | **RBAC** | Role-Based Access Control. A security model that restricts operations based on the roles assigned to users or service accounts. Available in the Enterprise edition. | | **Recovery Point Objective (RPO)** | The maximum acceptable amount of data loss measured in time. An RPO of 1 hour means that up to 1 hour of data may be lost in a disaster. | | **Recovery Time Objective (RTO)** | The maximum acceptable time to restore service after a disaster. An RTO of 4 hours means the system must be operational within 4 hours of an incident. | | **Replication Factor** | The number of copies of each partition maintained across Kafka brokers. A replication factor of 3 means each partition has three replicas, providing fault tolerance. | | **Restore** | The process of reading backed-up data from object storage and producing it to a target Kafka cluster, optionally filtered by time window, topic, or partition. | | **Runbook** | A documented procedure for performing a specific operational task, such as restoring a Kafka topic from backup or responding to a backup failure alert. | | **S3** | Amazon Simple Storage Service. An object storage service from AWS, and the most commonly used backup storage backend for OSO Kafka Backup. | | **SASL** | Simple Authentication and Security Layer. A framework for authentication used by Kafka clients, supporting mechanisms such as PLAIN, SCRAM-SHA-256, and SCRAM-SHA-512. | | **Segment** | A compressed file within a backup that contains a range of Kafka messages for a specific topic-partition. Segments are named by their starting offset (e.g., `segment-000000001000.zst`). | | **Server-Side Encryption (SSE)** | Encryption performed by the storage provider (e.g., S3, GCS, Azure Blob) at rest, transparently encrypting and decrypting objects without changes to the client. | | **SLA** | Service Level Agreement. A formal commitment defining the expected availability, performance, and support response times for a service. | | **SLI** | Service Level Indicator. A quantitative metric used to measure system behaviour, such as backup success rate, restore latency, or storage write throughput. | | **SLO** | Service Level Objective. A target value or range for an SLI, such as "99.9% backup success rate" or "restore completes within 4 hours." | | **Storage Tier** | A class of storage with specific cost and performance characteristics. For example, S3 Standard for active backups and S3 Glacier for long-term archival. | | **TLS** | Transport Layer Security. A cryptographic protocol that provides encrypted communication between Kafka clients and brokers, and between the backup tool and storage backends. | | **Topic** | A named category or feed in Apache Kafka to which messages are published. Topics are divided into partitions for scalability and parallelism. | | **ZooKeeper** | A centralised coordination service historically used by Apache Kafka for cluster metadata management, being replaced by KRaft in modern Kafka deployments. | --- title: Further Reading description: Curated resources for deepening your Kafka backup and disaster recovery knowledge source_url: html: https://kafkabackup.com/well-architected/further-reading md: https://kafkabackup.com/well-architected/further-reading.md --- # Further Reading Curated resources for deepening your understanding of Kafka backup, disaster recovery, and operational best practices. * * * - **Official Documentation:** [kafkabackup.com](https://kafkabackup.com) -- Comprehensive guides covering installation, configuration, backup, restore, and monitoring. - **GitHub Repository:** [github.com/osodevops/kafka-backup](https://github.com/osodevops/kafka-backup) -- Source code, issue tracker, and contribution guidelines. - **Demo Repository:** [github.com/osodevops/kafka-backup-demos](https://github.com/osodevops/kafka-backup-demos) -- End-to-end examples, benchmark suites, and reference architectures for testing OSO Kafka Backup in various environments. - **Changelog:** [CHANGELOG.md](https://github.com/osodevops/kafka-backup/blob/main/CHANGELOG.md) -- Release notes and version history for all OSO Kafka Backup releases. * * * These cloud provider frameworks informed the structure and principles of this Well-Architected guide: - **AWS Well-Architected Framework:** [docs.aws.amazon.com/wellarchitected/latest/framework/welcome.html](https://docs.aws.amazon.com/wellarchitected/latest/framework/welcome.html) -- Amazon's framework covering operational excellence, security, reliability, performance efficiency, cost optimisation, and sustainability. - **Azure Well-Architected Framework:** [learn.microsoft.com/en-us/azure/well-architected/](https://learn.microsoft.com/en-us/azure/well-architected/) -- Microsoft's guidance for designing and operating reliable, secure, and efficient workloads on Azure. - **Google Cloud Architecture Framework:** [cloud.google.com/architecture/framework](https://cloud.google.com/architecture/framework) -- Google's best practices for building well-architected systems on Google Cloud Platform. * * * - **Apache Kafka Documentation:** [kafka.apache.org/documentation/](https://kafka.apache.org/documentation/) -- The official Apache Kafka documentation, including broker configuration, client APIs, security, and operations guides. - **KRaft Mode:** [kafka.apache.org/documentation/#kraft](https://kafka.apache.org/documentation/#kraft) -- Documentation for Kafka's ZooKeeper-free consensus mode, which simplifies cluster management and is the default from Kafka 4.0. * * * The following standards and frameworks are relevant to organisations implementing Kafka backup and disaster recovery in regulated environments: - **NIST Cybersecurity Framework** -- A voluntary framework from the U.S. National Institute of Standards and Technology providing guidelines for managing and reducing cybersecurity risk. Relevant to backup encryption, access controls, and incident response planning. - **SOC 2 Type II** -- An auditing standard from the American Institute of CPAs (AICPA) that evaluates the effectiveness of an organisation's controls over security, availability, processing integrity, confidentiality, and privacy. Backup and restore procedures are a key component of SOC 2 compliance. - **ISO 27001** -- An international standard for information security management systems (ISMS). It provides a systematic approach to managing sensitive data, including requirements for backup procedures, access control, and disaster recovery. - **GDPR (General Data Protection Regulation)** -- The European Union's data protection regulation that governs how personal data is collected, processed, and stored. Relevant to backup data retention policies, right to erasure, and data masking requirements. The OSO Kafka Backup Enterprise edition provides specific GDPR compliance tools. - **PCI DSS (Payment Card Industry Data Security Standard)** -- A set of security standards for organisations handling credit card data. Relevant to encryption of backup data, access control, audit logging, and secure storage of cardholder information. --- title: Common Errors description: Troubleshooting common OSO Kafka Backup errors source_url: html: https://kafkabackup.com/troubleshooting/common-errors md: https://kafkabackup.com/troubleshooting/common-errors.md --- # Common Errors This guide covers the most frequently encountered errors and their solutions. ``` Error: Failed to connect to Kafka broker Cause: Connection refused (os error 111) Broker: kafka:9092 ``` **Causes:** - Broker is not running - Wrong hostname/port - Network/firewall issue **Solutions:** ``` # Test connectivity nc -zv kafka 9092 # Check if Kafka is running kafka-broker-api-versions --bootstrap-server kafka:9092 # Verify DNS resolution nslookup kafka ``` ``` # Fix: Correct broker addresses source: bootstrap_servers: - kafka-0.kafka.svc.cluster.local:9092 - kafka-1.kafka.svc.cluster.local:9092 ``` ``` Error: SSL handshake failed Cause: certificate verify failed ``` **Causes:** - Missing CA certificate - Expired certificate - Hostname mismatch > [!NOTE] > > [!NOTE] > > Fixed in v0.4.0 > > [!NOTE] > > Prior to v0.4.0, custom CA certificates configured via `ssl_ca_location` were not being used during TLS connections. This caused "UnknownIssuer" errors when connecting to Kafka brokers with self-signed or internal CA certificates. Upgrade to v0.4.0 or later to resolve this issue. See [GitHub Issue #3](https://github.com/osodevops/kafka-backup/issues/3) for details. **Solutions:** ``` # Verify certificate openssl s_client -connect kafka:9093 -CAfile ca.crt # Check certificate expiry openssl x509 -in ca.crt -noout -dates # Check hostname in certificate openssl x509 -in server.crt -noout -text | grep -A1 "Subject Alternative Name" ``` ``` # Fix: Provide correct CA certificate source: security: security_protocol: SSL ssl_ca_location: /certs/ca.crt ``` ``` Error: SASL authentication failed Cause: Authentication failed during SASL handshake ``` **Causes:** - Wrong username/password - User doesn't exist - Wrong SASL mechanism **Solutions:** ``` # Test authentication with kafka-console-consumer kafka-console-consumer \ --bootstrap-server kafka:9092 \ --consumer.config client.properties \ --topic test --max-messages 1 ``` ``` # Fix: Verify credentials source: security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA256 # Ensure correct mechanism sasl_username: backup-user sasl_password: ${KAFKA_PASSWORD} # Check environment variable ``` ``` Error: Broken pipe (os error 32) Cause: Connection terminated by broker ``` Or: ``` Error: Failed to send request: Connection reset by peer ``` **Causes:** - Cloud Kafka service terminated idle connection (Confluent Cloud, AWS MSK, etc.) - Network idle timeout - TCP keepalive not enabled **Symptoms:** - Errors appear after ~5 minutes of low topic activity - Topics with infrequent messages fail while active topics work - Restarting the backup temporarily fixes it **Solutions:** ``` # Fix: Enable TCP keepalive (enabled by default in v0.6.0+) source: bootstrap_servers: - kafka.confluent.cloud:9092 connection: tcp_keepalive: true # Keep connections alive keepalive_time_secs: 60 # Probe after 60s idle keepalive_interval_secs: 20 # Probe every 20s tcp_nodelay: true # Reduce latency ``` > [!TIP] > > [!NOTE] > > Confluent Cloud Users > > [!NOTE] > > Confluent Cloud terminates idle TCP connections after approximately 5 minutes. The default connection settings (enabled since v0.6.0) prevent this. If you're upgrading from an earlier version, ensure `tcp_keepalive: true` is set. ``` Error: Storage error: Backend error: error returned from database: (code: 14) unable to open database file ``` **Causes:** - Running with `readOnlyRootFilesystem: true` in Kubernetes without a writable `/tmp` volume - The working directory is not writable (common in container environments) **Solutions:** Mount `/tmp` as an `emptyDir` in your Kubernetes pod spec: ``` volumeMounts: - name: tmp mountPath: /tmp volumes: - name: tmp emptyDir: {} ``` Or configure a custom offset database path via `offset_storage.db_path` pointing to a writable volume: ``` offset_storage: db_path: /data/offsets.db ``` > [!NOTE] > > [!NOTE] > > note > > [!NOTE] > > Since **v0.11.4**, the automatically created continuous-mode offset store uses `$TMPDIR`. An explicit `offset_storage` block still defaults `db_path` to `./offsets.db`, so set a writable path when the container filesystem is read only. See [GitHub Issue #62](https://github.com/osodevops/kafka-backup/issues/62). ``` Error: Access denied to S3 bucket Bucket: my-kafka-backups Operation: PutObject ``` **Causes:** - Missing IAM permissions - Wrong credentials - Bucket policy restriction **Solutions:** ``` # Test AWS credentials aws sts get-caller-identity # Test bucket access aws s3 ls s3://my-kafka-backups/ # Check bucket policy aws s3api get-bucket-policy --bucket my-kafka-backups ``` ``` // Fix: Add IAM policy { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "s3:PutObject", "s3:GetObject", "s3:ListBucket", "s3:DeleteObject" ], "Resource": [ "arn:aws:s3:::my-kafka-backups", "arn:aws:s3:::my-kafka-backups/*" ] } ] } ``` ``` Error: Bucket not found Bucket: my-kafka-backups Region: us-west-2 ``` **Solutions:** ``` # Verify bucket exists aws s3api head-bucket --bucket my-kafka-backups # Check region aws s3api get-bucket-location --bucket my-kafka-backups ``` ``` # Fix: Correct bucket and region storage: backend: s3 bucket: my-kafka-backups region: us-east-1 # Correct region ``` ``` Error: Azure storage operation failed Container: kafka-backups Status: 403 Forbidden ``` **Solutions:** ``` # Test Azure credentials az storage container list --account-name mystorageaccount # Check connection string az storage container show --name kafka-backups --account-name mystorageaccount ``` ``` # Fix: Verify connection string storage: backend: azure container: kafka-backups connection_string: ${AZURE_STORAGE_CONNECTION_STRING} ``` ``` Error: Topic not found Topic: orders Cluster: kafka:9092 ``` **Causes:** - Topic doesn't exist - ACL restriction - Typo in topic name **Solutions:** ``` # List topics kafka-topics --bootstrap-server kafka:9092 --list # Check if topic exists kafka-topics --bootstrap-server kafka:9092 --describe --topic orders ``` ``` # Fix: Use correct topic pattern source: topics: include: - orders # Exact name - "orders-*" # Or pattern ``` ``` Error: Not authorized to access topic Topic: production-orders User: backup-user ``` **Solutions:** ``` # Check ACLs kafka-acls --bootstrap-server kafka:9092 \ --list --topic production-orders # Add read permission kafka-acls --bootstrap-server kafka:9092 \ --add --allow-principal User:backup-user \ --operation Read --operation Describe \ --topic production-orders ``` ``` WARN Best-effort offset sync failed after fetch error: ... Path: s3://bucket/backups//offsets.db ``` **Causes:** - Storage permission issue - Disk full - Network timeout **Solutions:** ``` # Test write access echo "test" | aws s3 cp - s3://bucket/backups/test.txt aws s3 rm s3://bucket/backups/test.txt ``` ``` # Fix: Adjust checkpoint settings backup: checkpoint_interval_secs: 60 # Less frequent checkpoint_retries: 3 ``` ``` Error: Produce error: Kafka error: Partition 0 not available for topic test-backup_restored ``` **Causes:** - Topic doesn't exist when using `topic_mapping` to restore to a new topic name - Topic was just created but Kafka hasn't finished propagating metadata - Insufficient replication factor for the cluster **Solutions:** ``` # Fix: Enable auto topic creation (v0.3.0+) restore: create_topics: true default_replication_factor: 3 # Match your cluster's requirements topic_mapping: source-topic: target-topic ``` ``` # Alternative: Pre-create the topic before restore kafka-topics --bootstrap-server kafka:9092 \ --create --topic test-backup_restored \ --partitions 6 --replication-factor 3 ``` > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > When using `topic_mapping`, always enable `create_topics: true` to ensure target topics are created automatically before the restore begins. This was fixed in v0.3.0. ``` Error: Backup not found Backup ID: production-backup-20241201 Path: s3://bucket/backups ``` **Solutions:** ``` # List available backups kafka-backup list --path s3://bucket/backups # Check exact backup ID kafka-backup list --path s3://bucket/backups --backup-id my-backup-id ``` ``` Error: Invalid time range for PITR Start: 1701388800000 End: 1701302400000 Reason: Start time is after end time ``` **Solution:** ``` # Fix: Correct time order restore: time_window_start: 1701302400000 # Earlier time time_window_end: 1701388800000 # Later time ``` ``` Error: Topic already exists with different configuration Topic: orders Existing partitions: 12 Backup partitions: 6 ``` **Solutions:** ``` # Option 1: Use topic mapping restore: topic_mapping: orders: restored-orders # Option 2: Delete existing topic first # (manual step before restore) # Option 3: Allow partition mismatch restore: allow_partition_mismatch: true ``` ``` Error: Schema ID not found Schema ID: 42 Topic: orders ``` **Causes:** - Schema Registry not backed up - Schema deleted - Wrong Schema Registry URL **Solutions:** ``` # Check if schema exists curl https://schema-registry:8081/schemas/ids/42 # List all schemas curl https://schema-registry:8081/subjects ``` ``` # Fix: Enable Schema Registry sync (Enterprise) enterprise: schema_registry: enabled: true url: https://schema-registry:8081 ``` ``` Error: Consumer group not found Group: order-processor ``` **Solutions:** ``` # List consumer groups kafka-consumer-groups --bootstrap-server kafka:9092 --list # Describe group kafka-consumer-groups --bootstrap-server kafka:9092 \ --describe --group order-processor ``` ``` Error: Offset out of range Topic: orders Partition: 0 Requested: 50000 Available: 0-45000 ``` **Causes:** - Data retention cleaned old messages - PITR window mismatch - Wrong offset mapping **Solutions:** ``` # Check available offsets kafka-run-class kafka.tools.GetOffsetShell \ --broker-list kafka:9092 \ --topic orders ``` ``` # Fix: Use timestamp-based reset offset_reset: strategy: timestamp timestamp: 1701388800000 # Within available range ``` ``` Error: Original offset header not found Topic: orders Expected header: x-original-offset ``` **Causes:** - Backup was taken with `include_offset_headers: false` (the default is `true`) - Headers stripped during processing (e.g. `restore.strip_offset_headers: true`) **Solutions:** ``` # Fix: Re-backup with headers enabled (the default) backup: include_offset_headers: true # Alternative: Use timestamp-based offset reset offset_reset: strategy: timestamp ``` ``` Error: Invalid configuration Field: compression_level Value: 25 Valid range: 1-22 ``` **Solution:** Check the [Configuration Reference](https://kafkabackup.com/reference/config-yaml.md) for valid values. ``` Error: Missing required field Field: source.bootstrap_servers ``` **Solution:** ``` # Fix: Add required field source: bootstrap_servers: - kafka:9092 ``` ``` Error: Environment variable not set Variable: KAFKA_PASSWORD Location: source.security.sasl_password ``` **Solution:** ``` # Set the environment variable export KAFKA_PASSWORD="your-password" # Or use a different approach source: security: sasl_password: "direct-value" # Not recommended ``` ``` Error: Backup validation failed Backup ID: my-backup Reason: Checksum mismatch in segment-0005.dat ``` **Causes:** - Corrupted storage - Incomplete upload - Storage system issue **Solutions:** ``` # Run deep validation kafka-backup validate \ --path s3://bucket/backups \ --backup-id my-backup \ --deep # Re-run backup kafka-backup backup --config backup.yaml --force ``` ``` Error: Restore validation failed Reason: Target cluster unreachable ``` **Solution:** ``` # Validate configuration first kafka-backup validate-restore --config restore.yaml # Check target connectivity kafka-broker-api-versions --bootstrap-server target-kafka:9092 ``` | Error | Likely Cause | First Step | | --- | --- | --- | | Connection refused | Broker down | Check broker status | | SSL handshake failed | Certificate issue | Verify certificates | | SASL auth failed | Wrong credentials | Test with console tools | | Broken pipe | Idle connection timeout | Enable TCP keepalive | | Access denied (S3) | IAM permissions | Check IAM policy | | Topic not found | ACL or typo | List topics | | Partition not available | Topic doesn't exist | Enable `create_topics` | | Backup not found | Wrong path/ID | List backups | | Schema ID not found | SR not backed up | Enable SR sync | - [Performance Issues](https://kafkabackup.com/troubleshooting/performance-issues.md) - Slow backup/restore - [Offset Discontinuity](https://kafkabackup.com/troubleshooting/offset-discontinuity.md) - Offset problems - [Debug Mode](https://kafkabackup.com/troubleshooting/debug-mode.md) - Enable verbose logging - [Support](https://kafkabackup.com/troubleshooting/support.md) - Get help --- title: Performance Issues description: Troubleshooting slow backup and restore operations source_url: html: https://kafkabackup.com/troubleshooting/performance-issues md: https://kafkabackup.com/troubleshooting/performance-issues.md --- # Performance Issues This guide helps diagnose and resolve slow backup and restore operations. ``` # Monitor backup progress in real-time kafka-backup status --config backup.yaml --watch # Example output: # ================================================================ # OSO Kafka Backup - Live Status # ================================================================ # Backup ID: production-backup Uptime: 00:15:32 # Status: RUNNING # ================================================================ # Progress # |- Records: 1,234,567 # |- Bytes: 256.0 MB (compressed) # |- Throughput: 15234 rec/s | 3.2 MB/s # |- Lag: 45,000 records (orders-0) # ================================================================ # Components # |- kafka: [OK] ok # |- storage: [OK] ok # ================================================================ # Compression: 3.2x ratio | Errors: 0 # ================================================================ # One-shot status (no continuous refresh) kafka-backup status --config backup.yaml # Custom refresh interval (5 seconds) kafka-backup status --config backup.yaml --watch --interval 5 ``` | Operation | Per Partition | With 10 Partitions | | --- | --- | --- | | Backup | 50-100 MB/s | 500 MB/s - 1 GB/s | | Restore | 75-150 MB/s | 750 MB/s - 1.5 GB/s | If you're seeing significantly lower numbers, continue troubleshooting. **Symptoms:** - Low records/sec - High CPU idle - Storage I/O is fine **Diagnosis:** ``` # Check consumer lag kafka-consumer-groups --bootstrap-server kafka:9092 \ --describe --group kafka-backup-$BACKUP_ID ``` **Solutions:** ``` # Increase fetch sizes source: kafka_config: fetch.max.bytes: 104857600 # 100 MB max.partition.fetch.bytes: 10485760 # 10 MB per partition fetch.min.bytes: 1048576 # 1 MB minimum fetch.max.wait.ms: 500 # Wait for batches ``` **Symptoms:** - Throughput limited regardless of settings - Network utilization at 100% **Diagnosis:** ``` # Check network utilization iftop -i eth0 # Test network bandwidth iperf3 -c kafka-broker-0 -p 5201 ``` **Solutions:** 1. **Compress before transfer:** ``` backup: compression: zstd compression_level: 1 # Fast compression ``` 2. **Use closer storage region:** ``` storage: backend: s3 region: us-west-2 # Same region as Kafka ``` 3. **Enable VPC endpoints:** ``` storage: backend: s3 endpoint: https://s3.us-west-2.amazonaws.com use_vpc_endpoint: true ``` **Symptoms:** - High CPU utilization - Compression taking most of the time **Diagnosis:** ``` # Check CPU during backup top -p $(pgrep kafka-backup) # Check metrics curl localhost:9090/metrics | grep compression ``` **Solutions:** ``` # Use faster compression backup: compression: lz4 # Faster than zstd # Or reduce level compression: zstd compression_level: 1 # Fastest ``` **Symptoms:** - Low storage throughput - High storage latency **Diagnosis:** ``` # Test S3 write speed dd if=/dev/zero bs=1M count=100 | aws s3 cp - s3://bucket/test-file # Check CloudWatch metrics for S3 ``` **Solutions:** ``` # Optimize multipart uploads storage: backend: s3 multipart_threshold: 52428800 # 50 MB multipart_part_size: 10485760 # 10 MB max_concurrent_uploads: 10 ``` **Symptoms:** - One partition much slower than others - Uneven partition sizes **Diagnosis:** ``` # Check partition sizes kafka-log-dirs --bootstrap-server kafka:9092 --describe ``` **Solutions:** 1. For future: Use better partition keys 2. For now: Accept longer backup time for skewed partitions **Symptoms:** - Low write throughput - Producer backpressure **Diagnosis:** ``` # Check producer metrics kafka-backup status --config restore.yaml --watch ``` **Solutions:** ``` # Optimize producer settings target: kafka_config: batch.size: 1048576 # 1 MB batches linger.ms: 100 # Wait for batches buffer.memory: 67108864 # 64 MB buffer acks: 1 # Trade durability for speed compression.type: lz4 # Producer-side compression ``` **Symptoms:** - Reading from storage is fine - CPU bound during restore **Solutions:** ``` # Use faster decompression (for future backups) backup: compression: lz4 # Decompresses faster than zstd ``` **Symptoms:** - PITR restore much slower than full restore - Many segments being read but few records restored **Diagnosis:** ``` # Check how many records match time window kafka-backup describe \ --path s3://bucket/backups \ --backup-id my-backup \ --format json | jq '.segments[] | select(.start_timestamp < 1701388800000)' ``` **Solutions:** ``` # Ensure time window aligns with segment boundaries # Backups with shorter checkpoint intervals = better PITR performance backup: checkpoint_interval_secs: 30 # More granular segments ``` ``` # Reduce memory for constrained environments backup: batch_size: 10000 # Fewer records per batch max_batch_bytes: 52428800 # 50 MB max # Increase memory for better throughput backup: batch_size: 100000 # More records per batch max_batch_bytes: 209715200 # 200 MB max ``` ``` # Reduce CPU usage backup: compression: lz4 # Lower CPU compression_level: 1 # Maximize throughput (more CPU) backup: compression: zstd compression_level: 3 # Parallel compression across partitions ``` ``` # More parallelism (higher resource usage) backup: parallel_partitions: 20 # Process more partitions concurrently # Less parallelism (lower resource usage) backup: parallel_partitions: 4 ``` ``` storage: backend: s3 bucket: my-bucket region: us-west-2 # Transfer acceleration use_accelerate_endpoint: true # Multipart optimization multipart_threshold: 52428800 multipart_part_size: 10485760 max_concurrent_uploads: 20 # Connection pooling max_connections: 50 ``` ``` storage: backend: azure container: kafka-backups # Block blob settings block_size: 10485760 # 10 MB blocks max_concurrency: 10 ``` ``` storage: backend: gcs bucket: kafka-backups # Parallel composite uploads parallel_composite_upload: true chunk_size: 10485760 # 10 MB chunks ``` ``` storage: backend: local path: /backups # Use SSD if available # Ensure sufficient IOPS ``` ``` # Backup throughput rate(kafka_backup_bytes_total[5m]) # Records per second rate(kafka_backup_records_total[5m]) # Compression ratio kafka_backup_compression_ratio # Batch processing time (p99) histogram_quantile(0.99, kafka_backup_batch_duration_seconds_bucket) ``` Key panels to monitor: - Throughput (MB/s) over time - Records/sec per partition - Compression ratio - Checkpoint latency - Storage write latency ``` # Watch backup progress kafka-backup backup --config backup.yaml --progress # Output: # Progress: 45% (450,000/1,000,000 records) # Speed: 85 MB/s # ETA: 10 minutes ``` - Kafka cluster has capacity for backup consumer - Storage bucket is in same region as Kafka - Network allows sufficient bandwidth - Compression level matches needs (speed vs size) - Monitor throughput metrics - Check for consumer lag - Verify checkpoint writes succeeding - Watch for rate limiting errors - Verify backup size is reasonable - Check compression ratio - Note total duration for planning - Run validation ``` # Test maximum throughput to storage dd if=/dev/zero bs=1M count=1000 | \ aws s3 cp - s3://bucket/benchmark/test-1gb # Test maximum throughput from Kafka kafka-consumer-perf-test \ --bootstrap-server kafka:9092 \ --topic orders \ --messages 1000000 \ --threads 4 ``` ``` # Test with different compression for level in 1 3 6 9; do kafka-backup backup \ --config backup.yaml \ --compression-level $level \ --dry-run \ --benchmark done ``` - [Debug Mode](https://kafkabackup.com/troubleshooting/debug-mode.md) - Enable verbose logging - [Performance Tuning Guide](https://kafkabackup.com/guides/performance-tuning.md) - Optimization strategies - [Architecture](https://kafkabackup.com/architecture/zero-copy-optimization.md) - Understanding internals --- title: Offset Discontinuity description: Troubleshooting offset-related issues after restore source_url: html: https://kafkabackup.com/troubleshooting/offset-discontinuity md: https://kafkabackup.com/troubleshooting/offset-discontinuity.md --- # Offset Discontinuity After restoring Kafka data, consumers may encounter offset discontinuities. This guide explains the causes and solutions. Kafka offsets are sequential numbers assigned to each message. After restore: ``` Source Cluster Offsets: Target Cluster Offsets: 0, 1, 2, 3, 4, 5, ... 0, 1, 2, 3, 4, 5, ... │ │ ▼ ▼ Consumer at offset 1000 Consumer at offset ??? ``` The problem: Consumer offset 1000 from source may not correspond to the same message in target. 1. **Fresh topic starts at 0** - New topic always begins at offset 0 2. **PITR filtering** - Not all messages restored 3. **Compacted topics** - Different compaction state 4. **Different partition count** - Records redistributed ``` # Source cluster (before migration) kafka-consumer-groups \ --bootstrap-server source-kafka:9092 \ --group order-processor \ --describe # Output: # TOPIC PARTITION CURRENT-OFFSET LOG-END-OFFSET # orders 0 1000 1500 # orders 1 800 1200 ``` ``` # Target cluster (after restore) kafka-run-class kafka.tools.GetOffsetShell \ --broker-list target-kafka:9092 \ --topic orders # Output: # orders:0:0:500 (partition 0 has 500 messages, starting at 0) # orders:1:0:400 (partition 1 has 400 messages, starting at 0) ``` ``` # Use OSO Kafka Backup to find mapping kafka-backup show-offset-mapping \ --bootstrap-servers target-kafka:9092 \ --topic orders \ --source-cluster "source-cluster-id" # Output: # Partition 0: Source offset 500-1000 → Target offset 0-500 # Partition 1: Source offset 400-800 → Target offset 0-400 ``` OSO Kafka Backup stores original offsets in headers. Use this to find correct position: ``` # Generate offset reset plan kafka-backup offset-reset plan \ --bootstrap-servers target-kafka:9092 \ --groups order-processor \ --strategy header-based \ --source-cluster "source-cluster-id" \ --output reset-plan.json # Review plan cat reset-plan.json # Execute kafka-backup offset-reset execute \ --plan reset-plan.json \ --bootstrap-servers target-kafka:9092 ``` Configuration: ``` offset_reset: strategy: header-based source_cluster: "source-cluster-id" groups: - order-processor - payment-processor ``` If headers aren't available, use timestamps: ``` # Reset to timestamp kafka-consumer-groups \ --bootstrap-server target-kafka:9092 \ --group order-processor \ --reset-offsets \ --to-datetime 2024-12-01T10:00:00.000 \ --all-topics \ --execute ``` Or using OSO Kafka Backup: ``` offset_reset: strategy: timestamp timestamp: 1701421200000 # Unix timestamp in milliseconds groups: - order-processor ``` If reprocessing is acceptable: ``` kafka-consumer-groups \ --bootstrap-server target-kafka:9092 \ --group order-processor \ --reset-offsets \ --to-earliest \ --all-topics \ --execute ``` For precise control: ``` # Generate mapping during restore kafka-backup restore \ --config restore.yaml \ --generate-offset-mapping \ --mapping-output offset-mapping.json # Apply mapping kafka-backup offset-reset execute \ --strategy from-mapping \ --mapping-file offset-mapping.json \ --groups order-processor ``` The three-phase restore handles offset translation automatically: ``` kafka-backup three-phase-restore --config restore.yaml ``` What it does: ``` Phase 1: Restore Data ├── Restore messages to target └── Include offset headers Phase 2: Build Mapping ├── Scan restored messages └── Build source→target offset map Phase 3: Reset Offsets ├── Read consumer group positions ├── Translate to target offsets └── Reset consumer groups ``` Configuration: ``` mode: restore backup_id: "my-backup" target: bootstrap_servers: - target-kafka:9092 restore: include_original_offset_header: true consumer_group_strategy: header-based reset_consumer_offsets: true consumer_groups: - order-processor - payment-processor - notification-service storage: backend: s3 bucket: kafka-backups prefix: production ``` When using point-in-time recovery, some messages are excluded: ``` Source: Messages at T=1,2,3,4,5,6,7,8,9,10 PITR: Only restore T=3-7 Target: Messages at T=3,4,5,6,7 (offsets 0-4) Consumer was at: T=5 (source offset 5) Should be at: T=5 (target offset 2) ``` Solution: ``` restore: time_window_start: 1701388800000 # T=3 time_window_end: 1701410400000 # T=7 include_original_offset_header: true offset_reset: strategy: header-based # Will find T=5 in target and map correctly ``` When partition count differs: ``` Source: 6 partitions Target: 12 partitions Message with key "order-123": Source: Partition 2, Offset 500 Target: Partition 8, Offset 50 (different partition!) ``` Solution: ``` offset_reset: strategy: header-based scan_all_partitions: true # Required for partition changes ``` When topics are renamed: ``` restore: topic_mapping: orders: restored-orders payments: restored-payments offset_reset: strategy: header-based topic_mapping: orders: restored-orders payments: restored-payments ``` Compacted topics may have different records: ``` Source: Keys A(v1), B(v1), A(v2), C(v1), B(v2) Compacted: A(v2), C(v1), B(v2) Backup: A(v2), C(v1), B(v2) (3 records) Target after restore: A(v2), C(v1), B(v2) at offsets 0,1,2 ``` For compacted topics, timestamp or header-based reset works best. If offset reset causes problems, rollback to previous position: ``` # Snapshot current offsets kafka-backup offset-rollback snapshot \ --bootstrap-servers target-kafka:9092 \ --groups order-processor \ --output pre-reset-snapshot.json # Perform reset kafka-backup offset-reset execute ... # If problems, rollback kafka-backup offset-rollback rollback \ --bootstrap-servers target-kafka:9092 \ --snapshot pre-reset-snapshot.json ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaOffsetRollback metadata: name: rollback-order-processor spec: kafkaCluster: bootstrapServers: - kafka:9092 consumerGroups: - order-processor operation: snapshot snapshotId: pre-reset ``` ``` # Get message at specific offset kafka-console-consumer \ --bootstrap-server target-kafka:9092 \ --topic orders \ --partition 0 \ --offset 100 \ --max-messages 1 \ --property print.headers=true ``` Look for the `x-original-offset` header (8-byte little-endian `i64`) to verify the original offset. ``` # Source message at offset 1000 kafka-console-consumer \ --bootstrap-server source-kafka:9092 \ --topic orders \ --partition 0 \ --offset 1000 \ --max-messages 1 \ --property print.key=true # Target message (should be same content) kafka-console-consumer \ --bootstrap-server target-kafka:9092 \ --topic orders \ --partition 0 \ --offset 500 \ --max-messages 1 \ --property print.key=true ``` ``` # Check consumer group can consume kafka-consumer-groups \ --bootstrap-server target-kafka:9092 \ --group order-processor \ --describe # LAG should be reasonable after reset ``` 1. **Always include offset headers** during backup 2. **Use three-phase restore** for complete migrations 3. **Take offset snapshots** before any reset 4. **Test with single consumer group** before bulk reset 5. **Verify data correctness** after reset 6. **Monitor consumer lag** after restart - Was backup created with `include_offset_headers: true` (the default — only an explicit `false` disables it)? - Is `source_cluster_id` correctly specified? - Are consumer groups stopped before reset? - Is the correct strategy being used? - For partition changes, is `scan_all_partitions` enabled? - Are topic mappings consistent between restore and offset reset? - [Offset Management Guide](https://kafkabackup.com/guides/offset-management.md) - Detailed offset operations - [Offset Translation Architecture](https://kafkabackup.com/architecture/offset-translation.md) - How it works - [CLI Reference](https://kafkabackup.com/reference/cli-reference.md#offset-reset) - Command options --- title: Debug Mode description: Enable verbose logging and debugging for OSO Kafka Backup source_url: html: https://kafkabackup.com/troubleshooting/debug-mode md: https://kafkabackup.com/troubleshooting/debug-mode.md --- # Debug Mode When troubleshooting issues, debug mode provides detailed logging to help identify problems. ``` # Basic debug logging export RUST_LOG=debug kafka-backup backup --config backup.yaml # Verbose debug logging export RUST_LOG=trace kafka-backup backup --config backup.yaml # Module-specific logging export RUST_LOG=kafka_backup=debug,rdkafka=info kafka-backup backup --config backup.yaml ``` ``` # Debug mode kafka-backup backup --config backup.yaml --debug # Trace mode (most verbose) kafka-backup backup --config backup.yaml --trace # Quiet mode (errors only) kafka-backup backup --config backup.yaml --quiet ``` ``` logging: level: debug # error, warn, info, debug, trace format: json # text or json output: stderr # or file path # Module-specific levels modules: kafka_backup: debug rdkafka: warn aws_sdk: info ``` | Level | Description | When to Use | | --- | --- | --- | | `error` | Critical failures only | Production (minimal logging) | | `warn` | Warnings and errors | Production (recommended) | | `info` | Progress and status | Normal operation | | `debug` | Detailed operation info | Troubleshooting | | `trace` | All internal details | Deep debugging | ``` 2024-12-01T10:00:00.123Z DEBUG kafka_backup::backup Starting backup 2024-12-01T10:00:00.124Z DEBUG kafka_backup::kafka Connecting to bootstrap servers: ["kafka:9092"] 2024-12-01T10:00:00.200Z DEBUG rdkafka::client Broker kafka:9092 connected (version 3.5.0) 2024-12-01T10:00:00.201Z DEBUG kafka_backup::kafka Fetching topic metadata 2024-12-01T10:00:00.250Z DEBUG kafka_backup::kafka Topics found: ["orders", "payments"] 2024-12-01T10:00:00.251Z DEBUG kafka_backup::backup Starting partition consumers 2024-12-01T10:00:00.252Z DEBUG kafka_backup::consumer Partition 0: seeking to offset 0 2024-12-01T10:00:00.253Z DEBUG kafka_backup::consumer Partition 1: seeking to offset 0 2024-12-01T10:00:01.000Z DEBUG kafka_backup::consumer Partition 0: received batch of 1000 records 2024-12-01T10:00:01.001Z DEBUG kafka_backup::compress Compressing batch (10 MB → 2 MB) 2024-12-01T10:00:01.050Z DEBUG kafka_backup::storage Writing segment: s3://bucket/backups/segment-0000.dat 2024-12-01T10:00:01.200Z DEBUG kafka_backup::storage Segment written successfully (2 MB) ``` 1. **Connection details** ``` DEBUG kafka_backup::kafka Connecting to bootstrap servers DEBUG rdkafka::client Broker connected ``` 2. **Topic/partition discovery** ``` DEBUG kafka_backup::kafka Topics found: [...] DEBUG kafka_backup::backup Partitions to backup: [...] ``` 3. **Progress information** ``` DEBUG kafka_backup::consumer Partition X: received batch DEBUG kafka_backup::storage Writing segment ``` 4. **Errors and retries** ``` WARN kafka_backup::storage Write failed, retrying (attempt 2/3) ERROR kafka_backup::storage Write failed: Access denied ``` ``` # Enable rdkafka debug export RUST_LOG=rdkafka=debug kafka-backup backup --config backup.yaml # Or via config source: kafka_config: debug: "all" # broker,topic,msg,protocol,cgrp,security,fetch,feature,interceptor,all ``` ``` source: kafka_config: debug: "broker,security,fetch" ``` | Category | What it Shows | | --- | --- | | `broker` | Broker connections | | `security` | SASL/SSL handshakes | | `fetch` | Fetch requests/responses | | `msg` | Message processing | | `protocol` | Wire protocol details | | `cgrp` | Consumer group operations | ``` # Enable OpenSSL debugging export RUST_LOG=rdkafka::client=debug # Check SSL handshake kafka-backup backup --config backup.yaml --debug 2>&1 | grep -i ssl ``` ``` # Enable AWS SDK debug logging export RUST_LOG=aws_sdk_s3=debug,aws_config=debug kafka-backup backup --config backup.yaml ``` ``` export RUST_LOG=azure_storage=debug kafka-backup backup --config backup.yaml ``` ``` export RUST_LOG=google_cloud_storage=debug kafka-backup backup --config backup.yaml ``` ``` # Test basic connectivity nc -zv kafka 9092 nc -zv kafka 9093 # SSL port # DNS resolution nslookup kafka # TCP dump (requires root) tcpdump -i any port 9092 -w kafka-traffic.pcap ``` ``` # Check server certificate openssl s_client -connect kafka:9093 -CAfile ca.crt # Verify certificate chain openssl verify -CAfile ca.crt server.crt # Check certificate expiry openssl x509 -in server.crt -noout -dates ``` Test configuration without making changes: ``` # Validate a restore without writing records kafka-backup validate-restore --config restore.yaml --format json ``` ``` # Show the accepted command options kafka-backup backup --help kafka-backup validate-restore --help ``` Version `0.15.11` has no `validate-config` command and no backup `--dry-run` flag. A backup command validates its YAML before connecting, while `validate-restore` is the non-destructive restore check. ``` # Quick validation kafka-backup validate \ --path s3://bucket/backups \ --backup-id my-backup # Deep validation (slower, checks all data) kafka-backup validate \ --path s3://bucket/backups \ --backup-id my-backup \ --deep # Verbose validation kafka-backup validate \ --path s3://bucket/backups \ --backup-id my-backup \ --deep \ --verbose ``` ``` kafka-backup validate-restore --config restore.yaml # Output: # Backup: Found # Target cluster: Reachable # Topics: All exist / Will be created # PITR window: Valid (250,000 records match) # Offset headers: Present ``` For scripting and analysis: ``` # JSON output kafka-backup backup --config backup.yaml --output json > backup-log.json # JQ filtering cat backup-log.json | jq 'select(.level == "ERROR")' cat backup-log.json | jq 'select(.target == "kafka_backup::storage")' ``` ``` logging: level: debug output: /var/log/kafka-backup/backup.log rotation: max_size_mb: 100 max_files: 5 ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: debug-backup spec: logging: level: debug # Or via environment env: - name: RUST_LOG value: "kafka_backup=debug,rdkafka=info" ``` View logs: ``` kubectl logs -n kafka-backup deployment/kafka-backup-operator -f # Filter for errors kubectl logs -n kafka-backup deployment/kafka-backup-operator | grep ERROR ``` ``` export RUST_LOG=rdkafka::client=debug,kafka_backup::kafka=debug kafka-backup backup --config backup.yaml 2>&1 | grep -E "(timeout|connect|broker)" ``` ``` export RUST_LOG=rdkafka::client=debug kafka-backup backup --config backup.yaml 2>&1 | grep -iE "(auth|sasl|ssl|security)" ``` ``` export RUST_LOG=kafka_backup=debug kafka-backup backup --config backup.yaml 2>&1 | grep -E "(batch|compress|write|duration)" ``` ``` export RUST_LOG=kafka_backup::storage=debug,aws_sdk_s3=debug kafka-backup backup --config backup.yaml 2>&1 | grep -iE "(s3|storage|write|upload)" ``` When reporting issues, include: 1. **Debug logs** with sensitive data redacted 2. **Configuration file** (without credentials) 3. **Version information**: `kafka-backup --version` 4. **Environment details**: OS, Kafka version, storage backend 5. **Steps to reproduce** - [Common Errors](https://kafkabackup.com/troubleshooting/common-errors.md) - Error reference - [Performance Issues](https://kafkabackup.com/troubleshooting/performance-issues.md) - Performance debugging - [Support](https://kafkabackup.com/troubleshooting/support.md) - Get help --- title: Support description: Get help with OSO Kafka Backup source_url: html: https://kafkabackup.com/troubleshooting/support md: https://kafkabackup.com/troubleshooting/support.md --- # Support Get help with OSO Kafka Backup through community resources or enterprise support. Report bugs and request features: - **Repository**: [https://github.com/osodevops/kafka-backup](https://github.com/osodevops/kafka-backup) - **Issues**: [https://github.com/osodevops/kafka-backup/issues](https://github.com/osodevops/kafka-backup/issues) **Before opening an issue:** 1. Search existing issues for similar problems 2. Check the documentation for solutions 3. Gather relevant information (see below) **What to include:** ``` ## Environment - OSO Kafka Backup version: X.Y.Z - Kafka version: X.Y.Z - Storage backend: S3 / Azure / GCS / Local - OS: Linux / macOS / Windows - Deployment: CLI / Docker / Kubernetes ## Problem Description What you expected vs. what happened ## Steps to Reproduce 1. Step one 2. Step two 3. ... ## Configuration (redact sensitive data) ```yaml mode: backup source: bootstrap_servers: - kafka:9092 # ... ``` ``` [Relevant debug logs] ``` Any other relevant information ``` ### GitHub Discussions For questions, ideas, and community help: - **Discussions**: https://github.com/osodevops/kafka-backup/discussions Categories: - **Q&A**: Ask questions - **Ideas**: Feature suggestions - **Show and Tell**: Share your use cases - **General**: General discussion ### Community Slack Join the community Slack for real-time help: - **Invite link**: https://oso.sh/slack - **Channel**: #kafka-backup ## Enterprise Support Support for Kafka Backup, the Enterprise binary and the Strimzi Backup Operator is included with the Kafka Backup Enterprise licence. There is one support model, not tiers, and it is delivered by OSO's own engineers, the team that builds and maintains the software. ### Hours and channels | Item | Detail | |------|--------| | Support hours | 08:00 to 18:00 Central European Time, Monday to Friday, excluding UK public holidays. Response targets are measured within these hours. | | Outside support hours | There is no staffed 24x7 desk. A ticket raised outside support hours is picked up at the start of the next support period. Out-of-hours P1 call-out and on-call cover for agreed days are available as priced options in the licence agreement. | | Channels | support@oso.sh, the OSO support portal (access details are issued at onboarding), and a shared Slack Connect or Microsoft Teams channel set up at onboarding. | | Access to your systems | None. OSO has no standing access to your clusters, storage or data; diagnostics are shared by you at your discretion. | ### Priorities and response targets | Priority | Definition | Initial response | Resolution target | |----------|------------|------------------|-------------------| | **P1 Critical** | A restore needed for recovery is failing or blocked, production backups have stopped with no workaround, or backup data is suspected incomplete or corrupt | 60 minutes | Workaround or restore path within 4 hours | | **P2 High** | Backup or restore degraded or partially failing; a workaround exists but is not sustainable | 4 hours | 1 business day | | **P3 Medium** | A non-critical function is impaired, or a production issue has an acceptable workaround | 1 business day | 3 business days | | **P4 Low** | Questions, guidance, documentation and feature requests | 2 business days | 5 business days or the next scheduled release | Resolution targets are working objectives rather than guarantees. Every P1 incident receives a documented root cause analysis within 5 business days of resolution. Escalation runs from the responding engineer to a lead engineer and then to the CTO. The full escalation path, the supported-version window, the security-fix policy and the maintenance commitments are published in the operator repository's [SUPPORT.md](https://github.com/osodevops/strimzi-backup-operator/blob/main/SUPPORT.md) and apply equally to the engine. ### Raising a ticket Include the kafka-backup or operator version, the storage backend and Kafka version, the configuration with secrets redacted, the relevant log excerpt, and for the operator the output of `kubectl get kafkabackups,kafkarestores -o yaml` together with the operator log. Set the priority using the definitions above; the responding engineer confirms it, and where the two views differ the higher priority applies until the initial assessment is complete. ## Self-Service Resources ### Documentation - **This site**: Comprehensive documentation - **CLI Help**: `kafka-backup --help` - **Command Help**: `kafka-backup backup --help` ### Troubleshooting Guides - [Common Errors](https://kafkabackup.com/troubleshooting/common-errors.md) - [Performance Issues](https://kafkabackup.com/troubleshooting/performance-issues.md) - [Offset Discontinuity](https://kafkabackup.com/troubleshooting/offset-discontinuity.md) - [Debug Mode](https://kafkabackup.com/troubleshooting/debug-mode.md) ### Version Information ```bash # Check version kafka-backup --version # Output: # kafka-backup 0.15.11 ``` When metrics are enabled for a running backup or restore, verify the HTTP server and inspect the live series directly: ``` curl --fail http://127.0.0.1:8080/health curl --fail http://127.0.0.1:8080/metrics ``` There is no `doctor` command. Run `validate-restore` before a restore, or `validate --path ... --backup-id ...` to verify an existing backup. ``` # Version kafka-backup --version # Environment uname -a cat /etc/os-release # Linux sw_vers # macOS # Resources free -h # Memory df -h # Disk nproc # CPUs ``` ``` # Export config (redact passwords!) kafka-backup config show --config backup.yaml --redact ``` ``` # Enable debug logging export RUST_LOG=debug kafka-backup backup --config backup.yaml 2> debug.log # Compress for upload gzip debug.log ``` ``` # Test Kafka connectivity kafka-broker-api-versions --bootstrap-server kafka:9092 # Test storage connectivity aws s3 ls s3://my-bucket/ # S3 az storage blob list --container-name my-container # Azure gsutil ls gs://my-bucket/ # GCS ``` ``` # Prometheus metrics curl http://localhost:9090/metrics > metrics.txt ``` **Q: Is there a size limit for backups?** A: No inherent limit. Limited only by storage capacity. **Q: Can I backup to multiple storage backends?** A: Not simultaneously. Run separate backups for different backends. **Q: Is the data encrypted at rest?** A: Depends on storage backend configuration. S3 SSE, Azure encryption, etc. **Q: Why is my backup slow?** A: See [Performance Issues](https://kafkabackup.com/troubleshooting/performance-issues.md). Common causes: network, compression, storage. **Q: How can I speed up restore?** A: Use parallel restores, reduce compression level during backup, use faster storage tier. **Q: How do I handle consumer offset reset?** A: Use three-phase restore or header-based offset reset. See [Offset Management](https://kafkabackup.com/guides/offset-management.md). **Q: Can I restore to a different number of partitions?** A: Yes, with `scan_all_partitions: true` for offset reset. **Q: How do I get an enterprise trial?** A: Visit [https://oso.sh/kafka-backup/trial](https://oso.sh/kafka-backup/trial) or run `kafka-backup license request`. **Q: What features require enterprise?** A: Encryption, masking, RBAC, audit logging, Schema Registry sync, plugins. Stay updated with new releases: - **Releases**: [https://github.com/osodevops/kafka-backup/releases](https://github.com/osodevops/kafka-backup/releases) - **Changelog**: [https://github.com/osodevops/kafka-backup/blob/main/CHANGELOG.md](https://github.com/osodevops/kafka-backup/blob/main/CHANGELOG.md) **Do not** report security vulnerabilities through public GitHub issues. **Email**: [security@oso.sh](mailto:security@oso.sh) Include: - Description of the vulnerability - Steps to reproduce - Potential impact - Suggested fix (if any) Reports are acknowledged within two business days and given a severity assessment within five. Coordinated disclosure, the supported-version window and how fixes are published are set out in the [security policy](https://github.com/osodevops/kafka-backup/blob/main/SECURITY.md). Published at: [https://github.com/osodevops/kafka-backup/security/advisories](https://github.com/osodevops/kafka-backup/security/advisories) We welcome contributions! - **Contributing Guide**: [https://github.com/osodevops/kafka-backup/blob/main/CONTRIBUTING.md](https://github.com/osodevops/kafka-backup/blob/main/CONTRIBUTING.md) - **Code of Conduct**: [https://github.com/osodevops/kafka-backup/blob/main/CODE\_OF\_CONDUCT.md](https://github.com/osodevops/kafka-backup/blob/main/CODE_OF_CONDUCT.md) Ways to contribute: - Report bugs - Suggest features - Improve documentation - Submit pull requests | Purpose | Contact | | --- | --- | | General questions | [support@oso.sh](mailto:support@oso.sh) | | Sales inquiries | [sales@oso.sh](mailto:sales@oso.sh) | | Security issues | [security@oso.sh](mailto:security@oso.sh) | | Partnerships | [partners@oso.sh](mailto:partners@oso.sh) | - **Website**: [https://oso.sh/kafka-backup](https://oso.sh/kafka-backup) - **Documentation**: [https://docs.oso.sh/kafka-backup](https://docs.oso.sh/kafka-backup) - **GitHub**: [https://github.com/osodevops/kafka-backup](https://github.com/osodevops/kafka-backup) - **Slack**: [https://oso.sh/slack](https://oso.sh/slack) - **Twitter**: [https://twitter.com/osodevops](https://twitter.com/osodevops) - [Common Errors](https://kafkabackup.com/troubleshooting/common-errors.md) - Error reference - [Debug Mode](https://kafkabackup.com/troubleshooting/debug-mode.md) - Enable debugging - [Getting Started](https://kafkabackup.com/getting-started/quickstart.md) - Quick start guide --- title: MSK KRaft Migration Troubleshooting description: Troubleshooting guide for kafka-backup Enterprise MSK ZooKeeper to KRaft migration errors and common issues. source_url: html: https://kafkabackup.com/troubleshooting/msk-kraft-migration md: https://kafkabackup.com/troubleshooting/msk-kraft-migration.md --- # MSK KRaft Migration Troubleshooting Precheck reports blockers that must be resolved before migration can proceed. See the [Precheck Codes Reference](https://kafkabackup.com/enterprise/msk-kraft-precheck-codes.md) for every code with detailed remediation. **Most common blockers:** | Code | Issue | Quick fix | | --- | --- | --- | | B09/B10 | Kafka brokers unreachable | Check security groups, VPC peering, bootstrap servers | | B07/B08 | S3 bucket unreachable | Create bucket or fix IAM policy | | B11/B12 | Target message size too small | Increase `message.max.bytes` on target | ``` Error: AccessDenied when calling PutObject on s3://bucket/prefix/... ``` **Cause:** The migration runner's IAM role lacks S3 permissions. **Fix:** Use the IAM policy generated by `plan --format iam-policy`. Ensure the role has `s3:PutObject`, `s3:GetObject`, `s3:ListBucket`, `s3:DeleteObject` on both the segments and evidence buckets. ``` Error: Timeout waiting for metadata from source cluster ``` **Cause:** Source cluster is under heavy load or network path is slow. **Fix:** - Reduce `seed.max_concurrent_partitions` (default 4) to lower source cluster load - Verify network bandwidth between runner and source - Check source broker CPU and memory in CloudWatch ``` Error: DescribeConfigs failed on target broker: NOT_CONTROLLER ``` **Cause:** KRaft clusters route admin requests to the controller. Metadata may briefly point to a non-controller broker after elections. **Fix:** Re-run execute. The retry will typically succeed. If persistent, check target cluster health in the MSK console. The tail phase reports per-partition lag. If lag is not decreasing: **Cause 1:** Source producer throughput exceeds replication speed. - Reduce source producer throughput during migration, or - Increase runner resources (CPU, network bandwidth) **Cause 2:** Compacted topics with high churn. - Compaction deletes records between seed and tail, causing apparent drift - This is expected — precheck warning W07 flags compacted topics **Cause 3:** Large messages consuming bandwidth. - Check `message.max.bytes` on source topics - Consider scheduling migration during off-peak hours ``` Error: drain_timeout exceeded — lag not within threshold after 30m ``` **Fix:** Increase `cutover.drain_timeout` in config, or reduce `cutover.drain_max_partition_lag` threshold. ``` cutover failed: Cutover failed: producer-freeze freeze: stdin is not a TTY — configure cutover.producer_freeze_webhook or run in an interactive shell ``` **Cause:** No `cutover.producer_freeze_webhook` is configured and the cutover command is running without an interactive terminal, so the tool cannot prompt for manual freeze confirmation. **Fix:** Configure `cutover.producer_freeze_webhook`, or rerun `cutover` from an interactive shell where the operator can confirm the producer freeze. ``` Error: producer freeze webhook did not respond within max_producer_freeze ``` **Cause:** The webhook URL is unreachable or took too long to respond. **Fix:** - Verify the webhook URL is accessible from the migration runner - Increase `cutover.max_producer_freeze` if producers need more time to drain - Check webhook service health and logs ``` Error: producer freeze webhook returned HTTP 500 ``` **Cause:** The freeze endpoint returned a non-2xx status. **Fix:** Check the webhook service logs. The tool expects HTTP 2xx for success. On failure, producers are automatically unfrozen (best-effort) and the migration enters `failed` state. Resume after fixing the webhook. ``` Error: failed to commit translated offsets: CoordinatorNotAvailable ``` **Cause:** The target cluster's consumer group coordinator is unavailable. **Fix:** Check target cluster health. Resume the migration — the cutover phase is idempotent and will retry the offset commit. ``` Validation FAILED: counts_and_offsets: 3 partitions exceed count_tolerance=1 ``` **Cause 1:** Records were produced to the target after cutover but before finalize. The validation compares source snapshot to current target state. **Fix:** Run finalize immediately after cutover-ack, before any applications write to the target. If this already happened, the extra records are from your applications — increase `validation.count_tolerance` or accept this as expected. **Cause 2:** Compacted topics have different retention on source and target. **Fix:** Compacted topic drift is expected and reported as W07. The validation downgrades count mismatches on compacted topics to warnings. **Cause 3:** A delete-retention topic aged out restored records on the target before the final switch. This can happen when source records are restored with their original `CreateTime` timestamps and the target topic has finite `retention.ms`. **Fix:** Temporarily extend topic retention for the migration window, rerun the repair/replay step for the affected partition, and finalize again. Recent versions also run a cutover log-start guard before `READY_FOR_CLIENT_SWITCH`; if truncation is detected there, the client switch is blocked. ``` Validation FAILED: spot_check_records: 1 mismatched on orders:0 ``` **Cause:** A sampled record differs between source and target. This can happen with compacted topics where records are deleted between seed and tail. **Fix:** Check if the topic uses `cleanup.policy=compact`. If so, this is expected — the validation reports it as a warning, not a failure, for compacted topics. For non-compacted topics, investigate the specific offset in the validation report. ``` Validation WARNING: spot_check_records: 122 samples compared, 122 matched, 26 skipped (no record to compare) ``` **Cause:** Some partitions have no comparable record, usually because they are empty, zero-span, or Kafka Connect internal partitions with no retained payload at the sampled range. **Fix:** Treat this as an accepted warning only when `mismatches=[]`, `counts_and_offsets=PASSED`, `offset_floor_violations=0`, and every comparable sample matched. If any comparable sample mismatched, investigate as a real validation failure. ``` Validation FAILED: sentinel_presence: partition 3 sentinel not found ``` **Cause:** The cutover sentinel record was not found at the expected offset on the target. **Fix:** This is rare. Check the cutover log for the expected sentinel offsets. The sentinel may have been compacted away if the topic has aggressive compaction. Resume finalize — the check will retry. Yes. `finalize` writes evidence for the validation attempt. If validation surfaces a repairable issue, fix the affected topic or partition, rerun the verification, and finalize again. The successful terminal state is still `finalized`. Example repair/finalize journal shape: ``` awaiting_client_switch -> validating operator confirmed client cutover validating -> failed validation failed: counts_and_offsets: partition(s) exceed count_tolerance=1 (evidence at s3://///evidence.json) failed -> finalized evidence signed + uploaded to s3://///evidence.json (validation=PASSED) ``` The validation details identified one affected partition: ``` overall=FAILED counts_and_offsets: 1 partition(s) exceed count_tolerance=1 spot_check_records: 1050 samples compared, 1049 matched, 1 mismatched, 0 skipped cdc.mongo.catalog.reviews/2 source_span=330 target_span=295 diff=35 fetch cdc.mongo.catalog.reviews/2@0: Kafka error: Broker returned error code 1 ``` The source topic configuration showed finite delete retention: ``` cleanup.policy=delete message.timestamp.type=CreateTime retention.ms=604800000 segment.ms=604800000 ``` After repair, independent source/target comparisons passed: ``` topic_parity=PASSED counts_and_offsets=PASSED offset_floor_violations=0 spot_check_records=PASSED ``` ``` Error: resume fingerprint mismatch — config may have changed between runs ``` **Cause:** The migration config (source ARN, target ARN, bucket names) changed between the original execute and the resume attempt. **Fix:** - If the config change was intentional, use `--force-restart` to override (accepts risk of duplicates) - If unintentional, restore the original config and retry ``` Error: failed to parse journal entry at line 5 ``` **Cause:** The journal file was manually edited or partially written during a crash. **Fix:** If using local journal (`--journal-dir`), check the `journal.jsonl` file. Remove the incomplete last line if it was mid-write during a crash. The migration will resume from the last complete entry. ``` Error: PutObject failed: InvalidRequest — bucket does not have Object Lock enabled ``` **Cause:** The config specifies `evidence.retention` but the bucket was not created with Object Lock. **Fix:** Object Lock cannot be added retroactively. Either: 1. Create a new bucket with Object Lock enabled and update the config 2. Remove `evidence.retention` from the config (evidence uploads without retention) ``` Error: signing key not found at path: /etc/kafka-backup/evidence-signing.key ``` **Cause:** `evidence.signing_key_path` points to a file that doesn't exist. **Fix:** Either provide the Ed25519 signing key at the configured path, or remove `signing_key_path` from the config to use the built-in demo key. Typical causes: - **IAM permissions**: Dev uses broad admin policies; prod has scoped permissions. Use the generated `iam-policy-concrete.json` from `plan --format iam-policy`. - **Security groups**: Dev runner is in the same VPC; prod runner needs VPC peering or PrivateLink to reach MSK brokers. - **Cross-account**: Source and target in different AWS accounts need cross-account IAM roles and resource policies. - **KMS encryption**: Prod S3 bucket uses KMS — the runner needs `kms:GenerateDataKey` and `kms:Decrypt`. The tail phase runs indefinitely until lag converges. If it seems stuck: 1. Check `status` output for per-partition lag 2. Identify the lagging partitions 3. Check if those partitions have high producer throughput 4. Consider reducing source throughput or increasing runner resources 5. If acceptable, increase `cutover.drain_max_partition_lag` to allow more lag before declaring drain-ready **Cause:** Client applications have cached the old bootstrap servers or are connecting to the source cluster via a DNS alias that hasn't been updated. **Fix:** Verify that application configs point to the target cluster's bootstrap servers. Check DNS TTLs if using CNAME-based routing. Force-restart consumer applications to pick up the new config. - [Precheck Codes Reference](https://kafkabackup.com/enterprise/msk-kraft-precheck-codes.md) — all precheck findings - [Configuration Reference](https://kafkabackup.com/enterprise/msk-kraft-config-reference.md) — tuning parameters - [Monitoring Guide](https://kafkabackup.com/guides/msk-kraft-monitoring.md) — what to watch during migration --- title: Kafka Streams Example description: Backup and restore Kafka Streams applications source_url: html: https://kafkabackup.com/examples/kafka-streams md: https://kafkabackup.com/examples/kafka-streams.md --- # Kafka Streams Example This guide shows how to backup and restore data for Kafka Streams applications, including state stores and changelog topics. Kafka Streams applications create internal topics: ``` ┌────────────────────────────────────────────────────────────────────┐ │ Kafka Streams Application │ ├────────────────────────────────────────────────────────────────────┤ │ │ │ Input Topics State Stores Output Topics │ │ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ │ │ orders │ ───▶ │ order-store │ ───▶ │ enriched- │ │ │ │ │ │ (RocksDB) │ │ orders │ │ │ └─────────────┘ └──────┬──────┘ └─────────────┘ │ │ │ │ │ ▼ │ │ ┌─────────────────────┐ │ │ │ app-id-order-store- │ (Changelog topic) │ │ │ changelog │ │ │ └─────────────────────┘ │ │ │ │ ┌─────────────────────┐ │ │ │ app-id-KSTREAM- │ (Repartition topic) │ │ │ REPARTITION-0000 │ │ │ └─────────────────────┘ │ │ │ └────────────────────────────────────────────────────────────────────┘ ``` For complete Kafka Streams recovery: | Topic Type | Backup | Why | | --- | --- | --- | | Input topics | Yes | Source data | | Output topics | Yes | Processed results | | Changelog topics | Yes | State store data | | Repartition topics | Optional | Can be recreated | streams-backup.yaml ``` mode: backup backup_id: "streams-backup-${TIMESTAMP}" source: bootstrap_servers: - kafka:9092 topics: include: # Input topics - orders - inventory - customers # Output topics - enriched-orders - order-totals - alerts # Internal topics (by application.id) - "order-processor-*" # Captures changelog and repartition exclude: # Skip repartition topics if desired (can be recreated) - "*-repartition-*" storage: backend: s3 bucket: kafka-backups prefix: streams/order-processor backup: compression: zstd compression_level: 3 include_offset_headers: true source_cluster_id: "production" ``` Backup only input and output topics: ``` source: topics: include: - orders - inventory - customers - enriched-orders - order-totals # Skip internal topics (state will be rebuilt from input) exclude: - "order-processor-*" ``` Restore everything including state stores: streams-restore-full.yaml ``` mode: restore backup_id: "streams-backup-20241201" target: bootstrap_servers: - target-kafka:9092 storage: backend: s3 bucket: kafka-backups prefix: streams/order-processor restore: # Restore all topics including changelog topics: - orders - inventory - customers - enriched-orders - order-totals - "order-processor-*" include_original_offset_header: true ``` After restore: ``` # Restart Streams application # It will load state from changelog topics java -jar order-processor.jar ``` Restore only input topics, let Streams rebuild state: streams-restore-rebuild.yaml ``` mode: restore backup_id: "streams-backup-20241201" target: bootstrap_servers: - target-kafka:9092 restore: topics: - orders - inventory - customers # Reset to beginning to reprocess all data consumer_group_strategy: earliest ``` After restore: ``` # Delete state store directory rm -rf /var/kafka-streams/order-processor # Reset application to reprocess kafka-streams-application-reset \ --bootstrap-servers target-kafka:9092 \ --application-id order-processor \ --input-topics orders,inventory,customers # Restart application java -jar order-processor.jar ``` For point-in-time recovery, ensure state matches: streams-restore-pitr.yaml ``` mode: restore backup_id: "streams-backup-20241201" restore: # Restore to specific point time_window_end: 1701450000000 topics: # Input topics up to PITR point - orders - inventory - customers # DON'T restore changelog - state won't match PITR point # State will be rebuilt from input ``` Kafka Streams uses consumer groups named `{application.id}`: ``` # View Streams consumer group kafka-consumer-groups \ --bootstrap-server kafka:9092 \ --group order-processor \ --describe ``` Using Kafka Streams reset tool: ``` kafka-streams-application-reset \ --bootstrap-servers target-kafka:9092 \ --application-id order-processor \ --input-topics orders,inventory,customers \ --to-earliest ``` Using OSO Kafka Backup: ``` offset_reset: groups: - order-processor strategy: earliest ``` If restoring with changelog: ``` offset_reset: groups: - order-processor strategy: header-based source_cluster: "production" ``` ``` // Order processing Kafka Streams application Properties props = new Properties(); props.put(StreamsConfig.APPLICATION_ID_CONFIG, "order-processor"); props.put(StreamsConfig.BOOTSTRAP_SERVERS_CONFIG, "kafka:9092"); StreamsBuilder builder = new StreamsBuilder(); // State store for order aggregation builder.addStateStore( Stores.keyValueStoreBuilder( Stores.persistentKeyValueStore("order-totals"), Serdes.String(), Serdes.Double() ) ); // Process orders builder.stream("orders", Consumed.with(Serdes.String(), orderSerde)) .groupByKey() .aggregate( () -> 0.0, (key, order, total) -> total + order.getAmount(), Materialized.as("order-totals") ) .toStream() .to("customer-totals"); ``` ``` order-processor-order-totals-changelog order-processor-order-totals-repartition ``` ``` #!/bin/bash # backup-streams.sh APP_ID="order-processor" TIMESTAMP=$(date +%Y%m%d-%H%M%S) cat > /tmp/streams-backup.yaml << EOF mode: backup backup_id: "${APP_ID}-${TIMESTAMP}" source: bootstrap_servers: - kafka:9092 topics: include: - orders - customer-totals - "${APP_ID}-*" storage: backend: s3 bucket: kafka-backups prefix: streams/${APP_ID} backup: compression: zstd include_offset_headers: true source_cluster_id: "production" EOF kafka-backup backup --config /tmp/streams-backup.yaml ``` ``` #!/bin/bash # restore-streams.sh APP_ID="order-processor" BACKUP_ID="$1" # Step 1: Stop the Streams application kubectl scale deployment ${APP_ID} --replicas=0 # Step 2: Restore data cat > /tmp/streams-restore.yaml << EOF mode: restore backup_id: "${BACKUP_ID}" target: bootstrap_servers: - target-kafka:9092 storage: backend: s3 bucket: kafka-backups prefix: streams/${APP_ID} restore: topics: - orders - customer-totals - "${APP_ID}-*" include_original_offset_header: true EOF kafka-backup three-phase-restore --config /tmp/streams-restore.yaml # Step 3: Restart Streams application kubectl scale deployment ${APP_ID} --replicas=3 ``` Kafka Streams stores state locally: ``` /var/kafka-streams/{application.id}/{task.id}/ ├── rocksdb/ │ └── order-totals/ │ ├── CURRENT │ ├── MANIFEST-000001 │ ├── OPTIONS-000001 │ └── 000001.sst └── .checkpoint ``` | Option | Method | Time | Consistency | | --- | --- | --- | --- | | From changelog | Restore changelog topics | Fast | Exact state | | Rebuild | Reprocess input topics | Slow | Eventually consistent | | Standby replicas | Use standby task | Instant | Exact state | Configure standby replicas for faster recovery: ``` props.put(StreamsConfig.NUM_STANDBY_REPLICAS_CONFIG, 1); ``` ``` # Check topic exists with data kafka-console-consumer \ --bootstrap-server target-kafka:9092 \ --topic orders \ --from-beginning \ --max-messages 10 ``` ``` # Check changelog has state kafka-console-consumer \ --bootstrap-server target-kafka:9092 \ --topic order-processor-order-totals-changelog \ --from-beginning \ --max-messages 10 \ --property print.key=true ``` ``` # Query state store via interactive query curl http://streams-app:8080/state/order-totals/customer-123 ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: streams-backup namespace: kafka-backup spec: schedule: "0 * * * *" # Hourly kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders - customer-totals - "order-processor-*" storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 prefix: streams/order-processor compression: zstd ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaRestore metadata: name: streams-restore namespace: kafka-backup spec: backupId: "order-processor-20241201-100000" targetCluster: bootstrapServers: - target-kafka:9092 storage: storageType: s3 s3: bucket: kafka-backups region: us-west-2 prefix: streams/order-processor offsetReset: strategy: headerBased consumerGroups: - order-processor ``` 1. **Include changelog topics** for fast state recovery 2. **Use application.id prefix** to capture all internal topics 3. **Test state consistency** after restore 4. **Consider standby replicas** for production deployments 5. **Document internal topic naming** for your applications - [Spring Boot Example](https://kafkabackup.com/examples/spring-boot.md) - Spring Kafka integration - [Offset Management](https://kafkabackup.com/guides/offset-management.md) - Consumer offset handling - [Disaster Recovery](https://kafkabackup.com/use-cases/disaster-recovery.md) - DR planning --- title: Spring Boot Example description: Backup and restore Kafka for Spring Boot applications source_url: html: https://kafkabackup.com/examples/spring-boot md: https://kafkabackup.com/examples/spring-boot.md --- # Spring Boot Example This guide shows how to backup and restore Kafka data for Spring Boot applications using Spring Kafka. Typical Spring Boot application with Kafka: ``` ┌────────────────────────────────────────────────────────────────────┐ │ Spring Boot Application │ ├────────────────────────────────────────────────────────────────────┤ │ │ │ ┌──────────────────┐ ┌──────────────────┐ │ │ │ KafkaTemplate │ │ @KafkaListener │ │ │ │ (Producer) │ │ (Consumer) │ │ │ └────────┬─────────┘ └────────┬─────────┘ │ │ │ │ │ │ │ │ │ │ ▼ ▼ │ │ ┌──────────────────┐ ┌──────────────────┐ │ │ │ Output Topic │ │ Input Topic │ │ │ │ (orders-out) │ │ (orders-in) │ │ │ └──────────────────┘ └──────────────────┘ │ │ │ │ Consumer Group: ${spring.kafka.consumer.group-id} │ │ │ └────────────────────────────────────────────────────────────────────┘ ``` application.yml ``` spring: kafka: bootstrap-servers: kafka:9092 consumer: group-id: order-service auto-offset-reset: earliest enable-auto-commit: false key-deserializer: org.apache.kafka.common.serialization.StringDeserializer value-deserializer: org.springframework.kafka.support.serializer.JsonDeserializer producer: key-serializer: org.apache.kafka.common.serialization.StringSerializer value-serializer: org.springframework.kafka.support.serializer.JsonSerializer listener: ack-mode: MANUAL ``` ``` @Service public class OrderConsumer { @KafkaListener(topics = "orders-in", groupId = "order-service") public void processOrder( @Payload Order order, @Header(KafkaHeaders.OFFSET) long offset, Acknowledgment ack) { log.info("Processing order {} at offset {}", order.getId(), offset); // Process order orderService.process(order); // Manual commit ack.acknowledge(); } } ``` ``` @Service public class OrderProducer { private final KafkaTemplate kafkaTemplate; public void sendOrder(Order order) { kafkaTemplate.send("orders-out", order.getId(), order) .whenComplete((result, ex) -> { if (ex == null) { log.info("Sent order {} to partition {} offset {}", order.getId(), result.getRecordMetadata().partition(), result.getRecordMetadata().offset()); } }); } } ``` spring-backup.yaml ``` mode: backup backup_id: "order-service-${TIMESTAMP}" source: bootstrap_servers: - kafka:9092 topics: include: - orders-in - orders-out # DLQ topics - orders-in.DLT # Retry topics (if using Spring Retry) - orders-in-retry-0 - orders-in-retry-1 - orders-in-retry-2 storage: backend: s3 bucket: kafka-backups prefix: spring/order-service backup: compression: zstd include_offset_headers: true source_cluster_id: "production" ``` If using Spring Cloud Stream: ``` source: topics: include: # Input bindings - orders-in # Output bindings - orders-out # Error channel - orders-error # Spring Cloud Stream internal topics - "springCloudBus.*" ``` Restore and continue from where the application left off: spring-restore-resume.yaml ``` mode: restore backup_id: "order-service-20241201" target: bootstrap_servers: - target-kafka:9092 restore: include_original_offset_header: true consumer_group_strategy: header-based reset_consumer_offsets: true consumer_groups: - order-service storage: backend: s3 bucket: kafka-backups prefix: spring/order-service ``` ``` # Execute three-phase restore kafka-backup three-phase-restore --config spring-restore-resume.yaml ``` Restore and reprocess everything from the beginning: spring-restore-reprocess.yaml ``` mode: restore backup_id: "order-service-20241201" target: bootstrap_servers: - target-kafka:9092 restore: consumer_group_strategy: earliest consumer_groups: - order-service ``` Restore to a specific point in time: spring-restore-pitr.yaml ``` mode: restore backup_id: "order-service-20241201" target: bootstrap_servers: - target-kafka:9092 restore: time_window_end: 1701450000000 # Before the incident include_original_offset_header: true consumer_group_strategy: timestamp consumer_group_timestamp: 1701450000000 consumer_groups: - order-service ``` Spring Kafka's `@RetryableTopic` creates DLT topics: ``` source: topics: include: - orders-in - orders-in.DLT # Dead Letter Topic - orders-in-retry-0 # Retry topics - orders-in-retry-1 - orders-in-retry-2 ``` ``` restore: topics: - orders-in.DLT topic_mapping: orders-in.DLT: investigation-orders-dlt ``` After fixing the issue, reprocess DLT messages: ``` @KafkaListener(topics = "orders-in.DLT", groupId = "dlt-processor") public void reprocessDlt( @Payload Order order, @Header(KafkaHeaders.DLT_ORIGINAL_TOPIC) String originalTopic, @Header(KafkaHeaders.DLT_ORIGINAL_OFFSET) long originalOffset) { log.info("Reprocessing DLT message from {} offset {}", originalTopic, originalOffset); // Resend to original topic kafkaTemplate.send(originalTopic, order.getId(), order); } ``` ``` # Before restore kafka-consumer-groups \ --bootstrap-server kafka:9092 \ --group order-service \ --describe # Output: # TOPIC PARTITION CURRENT-OFFSET LOG-END-OFFSET LAG # orders-in 0 1000 1500 500 # orders-in 1 800 1200 400 ``` Configure Spring to handle reset: application.yml ``` spring: kafka: consumer: auto-offset-reset: earliest # or latest ``` For fine-grained control: ``` @Component public class OffsetManager implements ConsumerSeekAware { @Override public void onPartitionsAssigned( Map assignments, ConsumerSeekCallback callback) { // Seek to specific offsets from backup mapping assignments.forEach((tp, offset) -> { Long targetOffset = getOffsetFromMapping(tp); if (targetOffset != null) { callback.seek(tp.topic(), tp.partition(), targetOffset); } }); } } ``` ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: order-service-backup spec: schedule: "0 */2 * * *" # Every 2 hours kafkaCluster: bootstrapServers: - kafka:9092 topics: - orders-in - orders-out - orders-in.DLT storage: storageType: s3 s3: bucket: kafka-backups prefix: spring/order-service ``` ``` #!/bin/bash # restore-order-service.sh BACKUP_ID="$1" NAMESPACE="production" DEPLOYMENT="order-service" # Step 1: Scale down application echo "Scaling down $DEPLOYMENT..." kubectl scale deployment $DEPLOYMENT -n $NAMESPACE --replicas=0 kubectl wait --for=condition=available=false deployment/$DEPLOYMENT -n $NAMESPACE # Step 2: Run restore echo "Restoring from backup $BACKUP_ID..." kubectl apply -f - < kafkaTemplate; @Test void shouldRestoreMessagesCorrectly() { // Send test messages for (int i = 0; i < 100; i++) { kafkaTemplate.send("orders-in", "order-" + i, new Order("order-" + i, 100.0)); } // Run backup (use test configuration) // ... // Clear topic // ... // Run restore // ... // Verify messages restored Consumer consumer = createConsumer(); consumer.subscribe(List.of("orders-in")); ConsumerRecords records = consumer.poll(Duration.ofSeconds(10)); assertThat(records.count()).isEqualTo(100); } } ``` application-prod.yml ``` spring: kafka: bootstrap-servers: ${KAFKA_BOOTSTRAP_SERVERS} consumer: group-id: order-service auto-offset-reset: none # Fail if no offset found enable-auto-commit: false isolation-level: read_committed producer: acks: all retries: 3 properties: enable.idempotence: true listener: ack-mode: MANUAL concurrency: 3 # Custom backup metadata backup: cluster-id: ${KAFKA_CLUSTER_ID:production} app-id: order-service ``` ``` @Configuration public class BackupMetadataConfig { @Value("${backup.cluster-id}") private String clusterId; @Value("${backup.app-id}") private String appId; // Use in offset header matching } ``` 1. **Use manual commits** (`enable-auto-commit: false`) for reliable processing 2. **Backup DLT topics** to investigate failures 3. **Test restore process** in staging environment 4. **Scale down before restore** to avoid conflicts 5. **Use idempotent producers** to handle duplicates after restore 6. **Monitor consumer lag** after restore - [Kafka Streams Example](https://kafkabackup.com/examples/kafka-streams.md) - Stateful processing - [Offset Management](https://kafkabackup.com/guides/offset-management.md) - Consumer offsets - [Kubernetes Deployment](https://kafkabackup.com/deployment/kubernetes.md) - K8s setup --- title: AWS Lambda Example description: Backup and restore Kafka for AWS Lambda consumers source_url: html: https://kafkabackup.com/examples/aws-lambda md: https://kafkabackup.com/examples/aws-lambda.md --- # AWS Lambda Example This guide shows how to backup and restore Kafka data consumed by AWS Lambda functions using Amazon MSK as an event source. ``` ┌─────────────────────────────────────────────────────────────────────┐ │ AWS Account │ ├─────────────────────────────────────────────────────────────────────┤ │ │ │ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ │ │ Amazon │ │ Lambda │ │ Other │ │ │ │ MSK │────▶│ Function │────▶│ Services │ │ │ │ │ │ │ │ (DynamoDB, │ │ │ │ ┌────────┐ │ │ │ │ S3, etc) │ │ │ │ │ orders │ │ │ │ │ │ │ │ │ └────────┘ │ └──────────────┘ └──────────────┘ │ │ └──────────────┘ │ │ │ │ │ │ Backup │ │ ▼ │ │ ┌──────────────┐ │ │ │ S3 Bucket │ │ │ │ (Backups) │ │ │ └──────────────┘ │ │ │ └─────────────────────────────────────────────────────────────────────┘ ``` AWS Lambda can consume from MSK/Kafka using Event Source Mapping: ``` // Lambda function receiving Kafka events exports.handler = async (event) => { for (const record of event.records) { for (const topicRecord of record.value) { const key = Buffer.from(topicRecord.key, 'base64').toString(); const value = Buffer.from(topicRecord.value, 'base64').toString(); const offset = topicRecord.offset; console.log(`Processing record: key=${key}, offset=${offset}`); // Process the message await processOrder(JSON.parse(value)); } } return { statusCode: 200 }; }; ``` ``` { "EventSourceArn": "arn:aws:kafka:us-west-2:123456789:cluster/my-cluster/abc123", "FunctionName": "order-processor", "Topics": ["orders"], "StartingPosition": "LATEST", "BatchSize": 100, "MaximumBatchingWindowInSeconds": 5, "DestinationConfig": { "OnFailure": { "Destination": "arn:aws:sqs:us-west-2:123456789:orders-dlq" } } } ``` msk-backup.yaml ``` mode: backup backup_id: "msk-orders-${TIMESTAMP}" source: bootstrap_servers: - b-1.my-cluster.abc123.kafka.us-west-2.amazonaws.com:9092 - b-2.my-cluster.abc123.kafka.us-west-2.amazonaws.com:9092 security: security_protocol: SASL_SSL sasl_mechanism: AWS_MSK_IAM topics: include: - orders - order-events - notifications storage: backend: s3 bucket: kafka-backups region: us-west-2 prefix: msk/orders backup: compression: zstd include_offset_headers: true source_cluster_id: "msk-production" ``` For MSK with IAM authentication: ``` source: security: security_protocol: SASL_SSL sasl_mechanism: AWS_MSK_IAM # Uses default credential chain (IAM role, env vars, etc.) ``` IAM Policy for backup: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "kafka-cluster:Connect", "kafka-cluster:DescribeTopic", "kafka-cluster:ReadData", "kafka-cluster:DescribeGroup" ], "Resource": [ "arn:aws:kafka:us-west-2:123456789:cluster/my-cluster/*", "arn:aws:kafka:us-west-2:123456789:topic/my-cluster/*", "arn:aws:kafka:us-west-2:123456789:group/my-cluster/*" ] }, { "Effect": "Allow", "Action": [ "s3:PutObject", "s3:GetObject", "s3:ListBucket" ], "Resource": [ "arn:aws:s3:::kafka-backups", "arn:aws:s3:::kafka-backups/*" ] } ] } ``` Lambda uses auto-generated consumer group IDs: ``` Consumer Group ID Pattern: amazon.lambda.{event-source-uuid} Example: amazon.lambda.12345678-1234-1234-1234-123456789012 ``` ``` # List consumer groups aws kafka list-groups \ --cluster-arn arn:aws:kafka:us-west-2:123456789:cluster/my-cluster/abc123 # Or via Kafka tools kafka-consumer-groups \ --bootstrap-server b-1.my-cluster.abc123.kafka.us-west-2.amazonaws.com:9092 \ --list \ --command-config msk.properties ``` ``` backup: include_offset_headers: true source_cluster_id: "msk-production" # Track Lambda consumer position consumer_groups: - "amazon.lambda.12345678-1234-1234-1234-123456789012" ``` msk-restore.yaml ``` mode: restore backup_id: "msk-orders-20241201" target: bootstrap_servers: - b-1.my-cluster.abc123.kafka.us-west-2.amazonaws.com:9092 security: security_protocol: SASL_SSL sasl_mechanism: AWS_MSK_IAM storage: backend: s3 bucket: kafka-backups prefix: msk/orders restore: include_original_offset_header: true ``` After restore, reset Lambda event source: ``` # Delete and recreate event source mapping aws lambda delete-event-source-mapping \ --uuid 12345678-1234-1234-1234-123456789012 aws lambda create-event-source-mapping \ --function-name order-processor \ --event-source-arn arn:aws:kafka:us-west-2:123456789:cluster/my-cluster/abc123 \ --topics orders \ --starting-position TRIM_HORIZON # Start from beginning ``` ``` restore: time_window_end: 1701450000000 # Lambda will need to start from TRIM_HORIZON # to process restored messages ``` For resuming exactly where Lambda left off: ``` # Get current Lambda offset position OFFSET=$(kafka-consumer-groups \ --bootstrap-server b-1.my-cluster.abc123.kafka.us-west-2.amazonaws.com:9092 \ --group amazon.lambda.12345678-1234-1234-1234-123456789012 \ --describe \ --command-config msk.properties | grep orders | awk '{print $4}') # Note: Lambda event source mapping doesn't support custom offset # Must use TRIM_HORIZON or LATEST ``` Run backups using Lambda: ``` // backup-lambda.js const { execSync } = require('child_process'); const path = require('path'); exports.handler = async (event) => { const backupId = `msk-backup-${Date.now()}`; const config = { mode: 'backup', backup_id: backupId, source: { bootstrap_servers: [process.env.MSK_BOOTSTRAP_SERVERS], security: { security_protocol: 'SASL_SSL', sasl_mechanism: 'AWS_MSK_IAM' }, topics: { include: event.topics || ['orders'] } }, storage: { backend: 's3', bucket: process.env.BACKUP_BUCKET, prefix: 'msk-backups' }, backup: { compression: 'zstd', include_offset_headers: true } }; // Write config const configPath = '/tmp/backup-config.yaml'; require('fs').writeFileSync(configPath, require('yaml').stringify(config)); // Execute backup try { execSync(`/opt/kafka-backup backup --config ${configPath}`, { stdio: 'inherit' }); return { statusCode: 200, body: JSON.stringify({ backup_id: backupId }) }; } catch (error) { console.error('Backup failed:', error); throw error; } }; ``` ``` # Dockerfile for Lambda layer FROM amazonlinux:2 RUN yum install -y curl tar gzip # Download kafka-backup binary RUN curl -L https://github.com/osodevops/kafka-backup/releases/latest/download/kafka-backup-linux-amd64.tar.gz | \ tar xz -C /opt # Create layer structure RUN mkdir -p /opt/layer && \ cp /opt/kafka-backup /opt/layer/ # Package layer WORKDIR /opt/layer RUN zip -r /opt/layer.zip . ``` ``` { "Name": "msk-backup-schedule", "ScheduleExpression": "rate(1 hour)", "State": "ENABLED", "Targets": [ { "Id": "backup-lambda", "Arn": "arn:aws:lambda:us-west-2:123456789:function:msk-backup", "Input": "{\"topics\": [\"orders\", \"order-events\"]}" } ] } ``` ``` resource "aws_cloudwatch_event_rule" "backup_schedule" { name = "msk-backup-schedule" description = "Trigger MSK backup every hour" schedule_expression = "rate(1 hour)" } resource "aws_cloudwatch_event_target" "backup_lambda" { rule = aws_cloudwatch_event_rule.backup_schedule.name target_id = "backup-lambda" arn = aws_lambda_function.backup.arn input = jsonencode({ topics = ["orders", "order-events"] }) } resource "aws_lambda_permission" "allow_eventbridge" { statement_id = "AllowExecutionFromEventBridge" action = "lambda:InvokeFunction" function_name = aws_lambda_function.backup.function_name principal = "events.amazonaws.com" source_arn = aws_cloudwatch_event_rule.backup_schedule.arn } ``` When Lambda fails, messages go to SQS DLQ. Backup both: ``` source: topics: include: - orders # Also backup SQS DLQ separately (different tool needed) ``` ``` // Lambda to move DLQ messages back to Kafka const AWS = require('aws-sdk'); const { Kafka } = require('kafkajs'); exports.handler = async (event) => { const sqs = new AWS.SQS(); const kafka = new Kafka({ clientId: 'dlq-replayer', brokers: [process.env.MSK_BOOTSTRAP_SERVERS], ssl: true, sasl: { mechanism: 'aws', authorizationIdentity: 'msk-admin', accessKeyId: process.env.AWS_ACCESS_KEY_ID, secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY, } }); const producer = kafka.producer(); await producer.connect(); // Receive messages from DLQ const response = await sqs.receiveMessage({ QueueUrl: process.env.DLQ_URL, MaxNumberOfMessages: 10 }).promise(); for (const message of response.Messages || []) { const kafkaMessage = JSON.parse(message.Body); // Resend to Kafka await producer.send({ topic: 'orders', messages: [{ key: kafkaMessage.key, value: kafkaMessage.value }] }); // Delete from DLQ await sqs.deleteMessage({ QueueUrl: process.env.DLQ_URL, ReceiptHandle: message.ReceiptHandle }).promise(); } await producer.disconnect(); }; ``` backup-us-west-2.yaml ``` source: bootstrap_servers: - b-1.primary-cluster.kafka.us-west-2.amazonaws.com:9092 storage: backend: s3 bucket: kafka-backups-us-west-2 region: us-west-2 prefix: primary # Enable cross-region replication in S3 ``` restore-us-east-1.yaml ``` target: bootstrap_servers: - b-1.dr-cluster.kafka.us-east-1.amazonaws.com:9092 storage: backend: s3 bucket: kafka-backups-us-east-1 # Replicated bucket region: us-east-1 prefix: primary ``` ``` // Lambda backup function with metrics const AWS = require('aws-sdk'); const cloudwatch = new AWS.CloudWatch(); async function reportMetrics(backupStats) { await cloudwatch.putMetricData({ Namespace: 'KafkaBackup', MetricData: [ { MetricName: 'RecordsBackedUp', Value: backupStats.records, Unit: 'Count' }, { MetricName: 'BackupDurationSeconds', Value: backupStats.duration, Unit: 'Seconds' }, { MetricName: 'BackupSizeBytes', Value: backupStats.bytes, Unit: 'Bytes' } ] }).promise(); } ``` ``` { "AlarmName": "KafkaBackupFailed", "MetricName": "Errors", "Namespace": "AWS/Lambda", "Dimensions": [ { "Name": "FunctionName", "Value": "msk-backup" } ], "Statistic": "Sum", "Period": 3600, "EvaluationPeriods": 1, "Threshold": 1, "ComparisonOperator": "GreaterThanOrEqualToThreshold", "AlarmActions": ["arn:aws:sns:us-west-2:123456789:alerts"] } ``` 1. **Use IAM authentication** for MSK when possible 2. **Schedule backups** via EventBridge 3. **Enable S3 cross-region replication** for DR 4. **Monitor Lambda errors** with CloudWatch 5. **Handle DLQ messages** separately 6. **Test restore in DR region** regularly - [Multi-Cluster DR](https://kafkabackup.com/examples/multi-cluster-dr.md) - Complex DR scenarios - [AWS S3 Setup](https://kafkabackup.com/deployment/cloud-setup/aws-s3.md) - S3 configuration - [Disaster Recovery](https://kafkabackup.com/use-cases/disaster-recovery.md) - DR planning --- title: Multi-Cluster Disaster Recovery description: Disaster recovery across multiple Kafka clusters source_url: html: https://kafkabackup.com/examples/multi-cluster-dr md: https://kafkabackup.com/examples/multi-cluster-dr.md --- # Multi-Cluster Disaster Recovery This guide demonstrates a comprehensive disaster recovery setup across multiple Kafka clusters in different regions. ``` ┌─────────────────────────────────────────────────────────────────────────────┐ │ Multi-Cluster DR Architecture │ ├─────────────────────────────────────────────────────────────────────────────┤ │ │ │ US-WEST-2 (Primary) US-EAST-1 (DR) │ │ ┌─────────────────────┐ ┌─────────────────────┐ │ │ │ Kafka Cluster │ │ Kafka Cluster │ │ │ │ (Production) │ │ (Standby) │ │ │ │ │ │ │ │ │ │ ┌──────────────┐ │ │ ┌──────────────┐ │ │ │ │ │ orders │ │ │ │ orders │ │ │ │ │ │ payments │ │ │ │ payments │ │ │ │ │ │ inventory │ │ │ │ inventory │ │ │ │ │ └──────────────┘ │ │ └──────────────┘ │ │ │ └──────────┬──────────┘ └──────────▲──────────┘ │ │ │ │ │ │ │ Backup │ Restore │ │ ▼ │ │ │ ┌─────────────────────────────────────────────────────────┐ │ │ │ S3 Backup Storage │ │ │ │ │ │ │ │ s3://kafka-backups/ │ │ │ │ ├── us-west-2/ │ │ │ │ │ ├── hourly/ │ │ │ │ │ └── daily/ │ │ │ │ └── cross-region-replica (us-east-1) │ │ │ │ │ │ │ └─────────────────────────────────────────────────────────┘ │ │ │ │ EU-WEST-1 (Analytics) AP-SOUTHEAST-1 (APAC) │ │ ┌─────────────────────┐ ┌─────────────────────┐ │ │ │ Kafka Cluster │ │ Kafka Cluster │ │ │ │ (Read Replica) │ │ (Regional) │ │ │ └─────────────────────┘ └─────────────────────┘ │ │ │ └─────────────────────────────────────────────────────────────────────────────┘ ``` primary-backup.yaml ``` mode: backup backup_id: "primary-${TIMESTAMP}" source: bootstrap_servers: - kafka-0.us-west-2.example.com:9092 - kafka-1.us-west-2.example.com:9092 - kafka-2.us-west-2.example.com:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA256 sasl_username: backup-service sasl_password: ${KAFKA_PASSWORD} ssl_ca_location: /certs/ca.crt topics: include: - orders - payments - inventory - customers - "events-*" exclude: - "__consumer_offsets" - "_schemas" storage: backend: s3 bucket: kafka-backups-primary region: us-west-2 prefix: us-west-2/hourly backup: compression: zstd compression_level: 3 checkpoint_interval_secs: 30 include_offset_headers: true source_cluster_id: "prod-us-west-2" ``` dr-restore.yaml ``` mode: restore backup_id: "${BACKUP_ID}" target: bootstrap_servers: - kafka-0.us-east-1.example.com:9092 - kafka-1.us-east-1.example.com:9092 - kafka-2.us-east-1.example.com:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA256 sasl_username: restore-service sasl_password: ${KAFKA_PASSWORD} ssl_ca_location: /certs/ca.crt storage: backend: s3 bucket: kafka-backups-dr region: us-east-1 prefix: us-west-2/hourly restore: include_original_offset_header: true consumer_group_strategy: header-based reset_consumer_offsets: true consumer_groups: - order-service - payment-processor - inventory-manager - notification-service ``` ``` # Create replication IAM role aws iam create-role \ --role-name S3ReplicationRole \ --assume-role-policy-document file://trust-policy.json # Attach replication policy aws iam put-role-policy \ --role-name S3ReplicationRole \ --policy-name S3Replication \ --policy-document file://replication-policy.json # Enable versioning on both buckets aws s3api put-bucket-versioning \ --bucket kafka-backups-primary \ --versioning-configuration Status=Enabled aws s3api put-bucket-versioning \ --bucket kafka-backups-dr \ --versioning-configuration Status=Enabled # Configure replication aws s3api put-bucket-replication \ --bucket kafka-backups-primary \ --replication-configuration file://replication-config.json ``` replication-config.json ``` { "Role": "arn:aws:iam::123456789:role/S3ReplicationRole", "Rules": [ { "ID": "KafkaBackupReplication", "Status": "Enabled", "Priority": 1, "Filter": { "Prefix": "us-west-2/" }, "Destination": { "Bucket": "arn:aws:s3:::kafka-backups-dr", "ReplicationTime": { "Status": "Enabled", "Time": { "Minutes": 15 } }, "Metrics": { "Status": "Enabled", "EventThreshold": { "Minutes": 15 } } }, "DeleteMarkerReplication": { "Status": "Disabled" } } ] } ``` primary-operator.yaml ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: primary-hourly-backup namespace: kafka-backup spec: schedule: "0 * * * *" # Hourly kafkaCluster: bootstrapServers: - kafka-0.us-west-2.example.com:9092 securityProtocol: SASL_SSL saslSecret: name: kafka-credentials mechanism: SCRAM-SHA256 tlsSecret: name: kafka-tls topics: - orders - payments - inventory - customers - "events-*" storage: storageType: s3 s3: bucket: kafka-backups-primary region: us-west-2 prefix: us-west-2/hourly compression: zstd compressionLevel: 3 includeOffsetHeaders: true sourceClusterId: "prod-us-west-2" # Configure an object-storage lifecycle policy on the us-west-2/hourly # prefix to enforce 7 days of hourly backup retention. --- apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaBackup metadata: name: primary-daily-backup namespace: kafka-backup spec: schedule: "0 2 * * *" # Daily at 2 AM kafkaCluster: bootstrapServers: - kafka-0.us-west-2.example.com:9092 securityProtocol: SASL_SSL saslSecret: name: kafka-credentials mechanism: SCRAM-SHA256 topics: - "*" excludeTopics: - "__*" - "_*" storage: storageType: s3 s3: bucket: kafka-backups-primary region: us-west-2 prefix: us-west-2/daily compression: zstd compressionLevel: 6 # Higher compression for archival includeOffsetHeaders: true sourceClusterId: "prod-us-west-2" # Configure an object-storage lifecycle policy on the us-west-2/daily # prefix to enforce 90 days of daily backup retention. ``` dr-restore-operator.yaml ``` apiVersion: kafka.oso.sh/v1alpha1 kind: KafkaRestore metadata: name: dr-restore namespace: kafka-backup spec: # Triggered manually or by automation backupId: "primary-20241201-120000" targetCluster: bootstrapServers: - kafka-0.us-east-1.example.com:9092 securityProtocol: SASL_SSL saslSecret: name: kafka-credentials mechanism: SCRAM-SHA256 storage: storageType: s3 s3: bucket: kafka-backups-dr region: us-east-1 prefix: us-west-2/hourly offsetReset: strategy: headerBased sourceCluster: "prod-us-west-2" consumerGroups: - order-service - payment-processor - inventory-manager ``` ``` #!/bin/bash # failover.sh - Execute DR failover set -e BACKUP_PATH="s3://kafka-backups-dr/us-west-2/hourly" DR_CLUSTER="kafka-0.us-east-1.example.com:9092" NAMESPACE="kafka-backup" : "${LATEST_BACKUP:?Set LATEST_BACKUP to the backup ID approved for failover}" echo "==========================================" echo " KAFKA DR FAILOVER" echo " $(date)" echo "==========================================" # Step 1: Confirm the selected backup. `list` intentionally has no JSON mode; # do not choose a disaster-recovery restore point implicitly. echo "[1/6] Confirming backup $LATEST_BACKUP..." kafka-backup list --path "$BACKUP_PATH" --backup-id "$LATEST_BACKUP" # Step 2: Validate backup echo "[2/6] Validating backup..." kafka-backup validate --path "$BACKUP_PATH" --backup-id "$LATEST_BACKUP" --deep # Step 3: Scale down applications (optional, if running in DR) echo "[3/6] Scaling down DR applications..." kubectl scale deployment -n production -l tier=kafka-consumer --replicas=0 # Step 4: Execute restore echo "[4/6] Restoring data to DR cluster..." cat > /tmp/dr-restore.yaml << EOF mode: restore backup_id: "$LATEST_BACKUP" target: bootstrap_servers: - $DR_CLUSTER security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA256 sasl_username: restore-service sasl_password: \${KAFKA_PASSWORD} storage: backend: s3 bucket: kafka-backups-dr region: us-east-1 prefix: us-west-2/hourly restore: include_original_offset_header: true consumer_group_strategy: header-based reset_consumer_offsets: true consumer_groups: - order-service - payment-processor - inventory-manager - notification-service EOF kafka-backup three-phase-restore --config /tmp/dr-restore.yaml # Step 5: Verify restore echo "[5/6] Verifying restore..." kafka-topics --bootstrap-server "$DR_CLUSTER" --list kafka-consumer-groups --bootstrap-server "$DR_CLUSTER" --list # Step 6: Scale up applications echo "[6/6] Scaling up DR applications..." kubectl scale deployment -n production -l tier=kafka-consumer --replicas=3 echo "==========================================" echo " FAILOVER COMPLETE" echo " DR Cluster: $DR_CLUSTER" echo " Backup Used: $LATEST_BACKUP" echo "==========================================" ``` ``` #!/bin/bash # failback.sh - Return to primary after DR event set -e PRIMARY_CLUSTER="kafka-0.us-west-2.example.com:9092" DR_CLUSTER="kafka-0.us-east-1.example.com:9092" BACKUP_PATH="s3://kafka-backups-dr/us-east-1/failback" : "${FAILBACK_BACKUP_ID:?Set FAILBACK_BACKUP_ID to the restored backup ID}" echo "==========================================" echo " KAFKA DR FAILBACK" echo " $(date)" echo "==========================================" # Step 1: Backup DR cluster (capture changes made during DR) echo "[1/5] Backing up DR cluster changes..." kafka-backup backup --config dr-backup.yaml # Step 2: Verify primary is healthy echo "[2/5] Verifying primary cluster health..." kafka-broker-api-versions --bootstrap-server "$PRIMARY_CLUSTER" # Step 3: Sync DR changes to primary echo "[3/5] Syncing DR changes to primary..." # This restores only the delta (changes made during DR event) kafka-backup restore --config failback-restore.yaml # Step 4: Reset consumer offsets on primary echo "[4/5] Resetting consumer offsets..." kafka-backup offset-reset execute \ --path "$BACKUP_PATH" \ --backup-id "$FAILBACK_BACKUP_ID" \ --bootstrap-servers "$PRIMARY_CLUSTER" \ --groups order-service,payment-processor,inventory-manager # Step 5: Redirect traffic back to primary echo "[5/5] Redirecting traffic to primary..." # Update DNS, load balancer, or service mesh configuration kubectl patch configmap kafka-config -n production \ --patch '{"data":{"bootstrap.servers":"kafka-0.us-west-2.example.com:9092"}}' # Restart applications to pick up new config kubectl rollout restart deployment -n production -l tier=kafka-consumer echo "==========================================" echo " FAILBACK COMPLETE" echo " Primary Cluster: $PRIMARY_CLUSTER" echo "==========================================" ``` prometheus-rules.yaml ``` apiVersion: monitoring.coreos.com/v1 kind: PrometheusRule metadata: name: kafka-backup-alerts namespace: monitoring spec: groups: - name: kafka-backup rules: # Backup failure alert - alert: KafkaBackupFailed expr: increase(kafka_backup_operator_backups_total{outcome="failure"}[10m]) > 0 for: 5m labels: severity: critical annotations: summary: "Kafka backup failed" description: "Backup {{ $labels.namespace }}/{{ $labels.name }} failed" # No successful completion in the expected two-hour window - alert: KafkaBackupMissing expr: sum(increase(kafka_backup_operator_backups_total{outcome="success"}[2h])) < 1 for: 5m labels: severity: warning annotations: summary: "Kafka backup completion is missing" description: "No successful backup in 2 hours" # S3 replication lag - alert: S3ReplicationLag expr: aws_s3_replication_latency_seconds > 1800 for: 5m labels: severity: warning annotations: summary: "S3 cross-region replication lag" description: "Replication lag exceeds 30 minutes" ``` `aws_s3_replication_latency_seconds` must come from your AWS/S3 exporter; it is not emitted by Kafka Backup. dr-dashboard.json ``` { "dashboard": { "title": "Multi-Cluster DR Status", "panels": [ { "title": "Successful Backups (24h)", "type": "stat", "targets": [ { "expr": "sum by (name) (increase(kafka_backup_operator_backups_total{outcome=\"success\"}[24h]))" } ] }, { "title": "Backup Failures (24h)", "type": "gauge", "targets": [ { "expr": "sum by (name) (increase(kafka_backup_operator_backups_total{outcome=\"failure\"}[24h]))" } ], "thresholds": { "steps": [ { "value": 0, "color": "green" }, { "value": 1, "color": "yellow" }, { "value": 2, "color": "red" } ] } }, { "title": "Backup Size Trend", "type": "graph", "targets": [ { "expr": "kafka_backup_operator_backup_size_bytes" } ] }, { "title": "Records Backed Up", "type": "graph", "targets": [ { "expr": "kafka_backup_operator_backup_records_total" } ] } ] } } ``` ``` #!/bin/bash # dr-drill.sh - Monthly DR test echo "Starting DR Drill..." # 1. Create test topic in primary kafka-topics --create --topic dr-test-$(date +%Y%m%d) \ --bootstrap-server primary-kafka:9092 # 2. Produce test messages kafka-producer-perf-test \ --topic dr-test-$(date +%Y%m%d) \ --num-records 10000 \ --record-size 1000 \ --throughput -1 \ --producer-props bootstrap.servers=primary-kafka:9092 # 3. Trigger backup kafka-backup backup --config primary-backup.yaml # 4. Wait for S3 replication sleep 300 # 5. Restore to DR kafka-backup restore --config dr-restore.yaml # 6. Verify data in DR COUNT=$(kafka-run-class kafka.tools.GetOffsetShell \ --broker-list dr-kafka:9092 \ --topic dr-test-$(date +%Y%m%d) | awk -F: '{sum+=$3} END {print sum}') if [ "$COUNT" -eq "10000" ]; then echo "DR Drill PASSED: All records restored" else echo "DR Drill FAILED: Expected 10000, got $COUNT" exit 1 fi # 7. Cleanup kafka-topics --delete --topic dr-test-$(date +%Y%m%d) \ --bootstrap-server primary-kafka:9092 kafka-topics --delete --topic dr-test-$(date +%Y%m%d) \ --bootstrap-server dr-kafka:9092 echo "DR Drill Complete" ``` 1. **Automate backups** - Use scheduled backups via operator or cron 2. **Enable S3 replication** - Cross-region replication for DR bucket 3. **Test DR regularly** - Monthly DR drills 4. **Monitor replication lag** - Alert on backup and replication delays 5. **Document runbooks** - Clear procedures for failover/failback 6. **Secure credentials** - Use secrets management for all clusters 7. **Version control configs** - GitOps for backup/restore configurations - [Disaster Recovery Guide](https://kafkabackup.com/use-cases/disaster-recovery.md) - DR planning - [Kubernetes Operator](https://kafkabackup.com/operator.md) - Operator setup - [Compliance](https://kafkabackup.com/use-cases/compliance-audit.md) - Meeting compliance requirements --- title: Example: SOX Compliance Evidence description: End-to-end example of generating signed compliance evidence for SOX ITGC audits source_url: html: https://kafkabackup.com/examples/compliance-evidence-sox md: https://kafkabackup.com/examples/compliance-evidence-sox.md --- # Example: SOX Compliance Evidence This example walks through a complete SOX Section 404 (IT General Controls) compliance scenario — from backup to signed evidence report. You're a platform engineer at a financial services company. Your SOX auditor requires: - Weekly proof that Kafka backups containing financial transaction data can be restored - SHA-256 checksums as tamper-evident documentation - Evidence retained for 7 years in write-once storage - A machine-readable report they can feed into their GRC platform sox-backup.yaml ``` mode: backup backup_id: "financial-data-weekly" source: bootstrap_servers: - kafka-prod-0.internal:9092 - kafka-prod-1.internal:9092 - kafka-prod-2.internal:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA-512 sasl_username: backup-service sasl_password: "${KAFKA_BACKUP_PASSWORD}" topics: include: - transactions - settlements - audit-log - "ledger-*" storage: backend: s3 bucket: sox-compliance-backups region: us-east-1 prefix: production/weekly backup: compression: zstd compression_level: 9 # Maximum compression for archival include_offset_headers: true # Required for offset restoration source_cluster_id: "prod-us-east-1" stop_at_current_offsets: true # Snapshot mode ``` Run the backup: ``` $ kafka-backup backup --config sox-backup.yaml ``` sox-restore.yaml ``` mode: restore backup_id: "financial-data-weekly" target: bootstrap_servers: - validation-kafka:9092 storage: backend: s3 bucket: sox-compliance-backups region: us-east-1 prefix: production/weekly restore: create_topics: true topic_mapping: transactions: validation-transactions settlements: validation-settlements audit-log: validation-audit-log ``` ``` $ kafka-backup restore --config sox-restore.yaml ``` ``` # One-time setup — store the private key securely $ openssl ecparam -genkey -name prime256v1 -noout | \ openssl pkcs8 -topk8 -nocrypt -out /etc/kafka-backup/sox-signing-key.pem $ openssl ec -in /etc/kafka-backup/sox-signing-key.pem \ -pubout -out /etc/kafka-backup/sox-signing-key-pub.pem # Restrict permissions $ chmod 600 /etc/kafka-backup/sox-signing-key.pem ``` sox-validation.yaml ``` backup_id: "financial-data-weekly" storage: backend: s3 bucket: sox-compliance-backups region: us-east-1 prefix: production/weekly target: bootstrap_servers: - validation-kafka:9092 checks: message_count: enabled: true mode: exact topics: - transactions - settlements - audit-log fail_threshold: 0 offset_range: enabled: true consumer_group_offsets: enabled: false # Not restoring consumer groups in this scenario evidence: formats: [json, pdf] signing: enabled: true private_key_path: "/etc/kafka-backup/sox-signing-key.pem" storage: prefix: "evidence-reports/sox/" retention_days: 2555 # 7 years (SOX requirement) notifications: slack: webhook_url: "${SLACK_SOX_CHANNEL_WEBHOOK}" triggered_by: "weekly-sox-validation-cron" ``` ``` $ kafka-backup validation run --config sox-validation.yaml ``` Expected output: ``` === Validation Results === Overall: PASSED Checks: 2/2 passed, 0 failed, 0 skipped Duration: 45ms [PASSED] MessageCountCheck — 3 topics; 2,841,293 messages expected, 2,841,293 restored; 0 discrepancies [PASSED] OffsetRangeCheck — 9 partitions checked; 9 passed; 0 issues JSON evidence uploaded: evidence-reports/sox/validation-a1b2c3d4/2026/04/validation-a1b2c3d4.json PDF evidence uploaded: evidence-reports/sox/validation-a1b2c3d4/2026/04/validation-a1b2c3d4.pdf Signature uploaded: evidence-reports/sox/validation-a1b2c3d4/2026/04/validation-a1b2c3d4.sig ``` /etc/cron.d/sox-validation ``` # Run every Sunday at 02:00 UTC 0 2 * * 0 kafka-backup validation run --config /etc/kafka-backup/sox-validation.yaml ``` After 52 weeks, your auditor has a URL to a year of uninterrupted weekly evidence — no manual intervention required. Your auditor downloads the evidence and verifies independently: ``` # Download the report and signature $ kafka-backup validation evidence-get \ --path s3://sox-compliance-backups \ --report-id validation-a1b2c3d4 \ --format json --output report.json # Verify integrity $ kafka-backup validation evidence-verify \ --report report.json \ --signature report.sig \ --public-key sox-signing-key-pub.pem ``` ``` SHA-256 checksum: VALID ECDSA signature: VALID Evidence report integrity: VERIFIED ``` The JSON evidence report contains a `compliance_mappings.sox_itgc` section: ``` { "sox_itgc": { "control": "IT General Controls - Backup and Recovery", "satisfied_by": ["MessageCountCheck", "OffsetRangeCheck"], "evidence_retention_required_years": 7, "evidence_retention_configured_days": 2555 } } ``` The PDF report contains: - **Page 1**: Overall result badge (PASSED), report ID, generation timestamp - **Page 2**: Per-check results table with record counts and durations - **Page 3**: SHA-256 checksums, signature details, SOX ITGC control mapping with retention confirmation For WORM compliance (immutable evidence): ``` $ aws s3api put-object-lock-configuration \ --bucket sox-compliance-backups \ --object-lock-configuration '{ "ObjectLockEnabled": "Enabled", "Rule": { "DefaultRetention": { "Mode": "COMPLIANCE", "Years": 7 } } }' ``` - [GDPR Compliance Example](https://kafkabackup.com/examples/compliance-evidence-gdpr.md) — shorter retention, PITR focus - [Backup Validation Guide](https://kafkabackup.com/guides/validation-compliance.md) — complete feature guide - [Evidence Signing Guide](https://kafkabackup.com/guides/evidence-signing.md) — key management deep-dive --- title: Example: GDPR Compliance Evidence description: Proving Kafka restore capability for GDPR Article 32 with signed evidence reports source_url: html: https://kafkabackup.com/examples/compliance-evidence-gdpr md: https://kafkabackup.com/examples/compliance-evidence-gdpr.md --- # Example: GDPR Compliance Evidence Demonstrate compliance with GDPR Article 32 — "regularly testing, assessing, and evaluating the effectiveness of technical and organisational measures" — by generating signed evidence that your Kafka backups containing personal data can be restored. You process personal data (user events, consent records, DSAR requests) through Kafka. Your DPO needs: - Monthly proof that backups of personal data can be restored - Evidence of restore capability (RTO demonstration) - Shorter retention than SOX (1 year, not 7) - PITR validation to prove data can be recovered at any point in time gdpr-backup.yaml ``` mode: backup backup_id: "personal-data-monthly" source: bootstrap_servers: - kafka-eu:9092 security: security_protocol: SASL_SSL sasl_mechanism: SCRAM-SHA-512 sasl_username: backup-service sasl_password: "${KAFKA_PASSWORD}" topics: include: - user-events - consent-records - dsar-requests - "pii-*" storage: backend: gcs bucket: gdpr-kafka-backups prefix: eu-west-1/monthly backup: compression: zstd include_offset_headers: true stop_at_current_offsets: true ``` ``` $ kafka-backup backup --config gdpr-backup.yaml ``` The key GDPR differentiator: prove you can restore to a **specific point in time**. gdpr-restore.yaml ``` mode: restore backup_id: "personal-data-monthly" target: bootstrap_servers: - validation-kafka-eu:9092 storage: backend: gcs bucket: gdpr-kafka-backups prefix: eu-west-1/monthly restore: # Restore data from a specific 24-hour window time_window_start: 1711843200000 # 2026-03-31T00:00:00Z time_window_end: 1711929600000 # 2026-04-01T00:00:00Z create_topics: true ``` ``` $ kafka-backup restore --config gdpr-restore.yaml ``` gdpr-validation.yaml ``` backup_id: "personal-data-monthly" storage: backend: gcs bucket: gdpr-kafka-backups prefix: eu-west-1/monthly target: bootstrap_servers: - validation-kafka-eu:9092 pitr_timestamp: 1711929600000 # The PITR point we restored to checks: message_count: enabled: true mode: exact topics: - user-events - consent-records - dsar-requests offset_range: enabled: true evidence: formats: [json, pdf] signing: enabled: true private_key_path: "/etc/kafka-backup/gdpr-signing-key.pem" storage: prefix: "evidence-reports/gdpr/" retention_days: 365 # 1 year (GDPR typical) notifications: slack: webhook_url: "${SLACK_DPO_CHANNEL}" triggered_by: "monthly-gdpr-validation" ``` ``` $ kafka-backup validation run --config gdpr-validation.yaml ``` Expected output: ``` === Validation Results === Overall: PASSED Checks: 2/2 passed, 0 failed, 0 skipped Duration: 23ms [PASSED] MessageCountCheck — 3 topics; 142,891 messages expected, 142,891 restored; 0 discrepancies [PASSED] OffsetRangeCheck — 9 partitions checked; 9 passed; 0 issues ``` The `compliance_mappings.gdpr_art32` section demonstrates Article 32 compliance: ``` { "gdpr_art32": { "control": "Article 32 - Testing technical measures", "satisfied_by": ["MessageCountCheck", "OffsetRangeCheck"], "test_frequency": "on-demand", "rto_demonstrated_seconds": 47 } } ``` The `pitr_timestamp` in the backup section proves point-in-time recovery capability: ``` { "backup": { "pitr_timestamp": 1711929600000, "total_records": 142891 } } ``` /etc/cron.d/gdpr-validation ``` # Run on the 1st of every month at 03:00 UTC 0 3 1 * * kafka-backup validation run --config /etc/kafka-backup/gdpr-validation.yaml ``` When you receive a Data Subject Access Request (DSAR) and need to prove what data existed at a specific time: ``` # 1. Auditor specifies the exact point in time $ kafka-backup validation run \ --config gdpr-validation.yaml \ --pitr 1709251200000 \ --triggered-by "DSAR-2026-0142 - data subject request for user ID 12345" # 2. The evidence report includes the triggered_by field # proving chain of custody for the specific DSAR ``` GDPR requires data minimisation. Set evidence retention to match your data retention policy: ``` evidence: storage: retention_days: 365 # 1 year — matches your GDPR data retention ``` > [!TIP] > > [!NOTE] > > tip > > [!NOTE] > > For GDPR workloads, avoid storing personal data in the evidence report itself. The report contains only metadata (topic names, record counts, timestamps) — not the actual message content. - [SOX Compliance Example](https://kafkabackup.com/examples/compliance-evidence-sox.md) — longer retention, financial data focus - [Backup Validation Guide](https://kafkabackup.com/guides/validation-compliance.md) — complete feature guide - [Compliance Evidence Use Cases](https://kafkabackup.com/use-cases/compliance-evidence.md) — framework comparison --- title: Example: MSK KRaft Migration (IAM to IAM) description: Complete worked example migrating an AWS MSK cluster from ZooKeeper to KRaft using IAM authentication on both sides. source_url: html: https://kafkabackup.com/examples/msk-kraft-migration-iam md: https://kafkabackup.com/examples/msk-kraft-migration-iam.md --- # Example: MSK KRaft Migration (IAM to IAM) This example walks through a complete migration of a 3-broker MSK ZooKeeper cluster to a 3-broker MSK KRaft cluster, both using IAM authentication in `us-east-1`. > [!NOTE] > > [!NOTE] > > Release evidence > > [!NOTE] > > This page is a worked IAM configuration example. The May 8, 2026 full AWS release qualification proved the same migration flow on SCRAM-SHA-512 source and SCRAM-SHA-512 target across four KRaft targets. Run this IAM path in staging before using the result as production change evidence. | | Source | Target | | --- | --- | --- | | **Cluster type** | MSK Provisioned | MSK Provisioned | | **Metadata mode** | ZooKeeper | KRaft | | **Kafka version** | 3.6.0 | 3.9.0 | | **Authentication** | IAM | IAM | | **Brokers** | 3 (kafka.m5.large) | 3 (kafka.m5.large) | | **Topics** | 50 | 0 (empty target) | | **Data volume** | ~500 GB | — | | **Consumer groups** | 12 | — | ``` # Create MSK configuration aws kafka create-configuration \ --name "prod-kraft-config" \ --kafka-versions "3.9.0" \ --server-properties "$(cat <<'EOF' auto.create.topics.enable=false default.replication.factor=3 min.insync.replicas=2 num.partitions=6 EOF )" # Create the KRaft cluster aws kafka create-cluster-v2 \ --cluster-name "prod-kraft" \ --provisioned '{ "brokerNodeGroupInfo": { "instanceType": "kafka.m5.large", "clientSubnets": ["subnet-aaa", "subnet-bbb", "subnet-ccc"], "securityGroups": ["sg-migration"], "storageInfo": {"ebsStorageInfo": {"volumeSize": 1000}} }, "numberOfBrokerNodes": 3, "clientAuthentication": {"sasl": {"iam": {"enabled": true}}}, "encryptionInfo": { "encryptionInTransit": {"clientBroker": "TLS", "inCluster": true} }, "kafkaVersion": "3.9.0" }' ``` ``` aws s3 mb s3://prod-migration-segments --region us-east-1 aws s3 mb s3://prod-migration-evidence --region us-east-1 # Optional: enable Object Lock on the evidence bucket for compliance aws s3api put-object-lock-configuration \ --bucket prod-migration-evidence \ --object-lock-configuration '{ "ObjectLockEnabled": "Enabled", "Rule": {"DefaultRetention": {"Mode": "COMPLIANCE", "Years": 7}} }' ``` The migration runner's IAM role needs access to both clusters and both buckets. Use the generated IAM policy from the `plan` command (Step 2 below) for the exact permissions. migration.yaml ``` enterprise: msk_kraft_migration: source: cluster_arn: arn:aws:kafka:us-east-1:123456789012:cluster/prod-zk/a1b2c3d4-5678-90ab-cdef-111111111111 auth: mode: iam target: cluster_arn: arn:aws:kafka:us-east-1:123456789012:cluster/prod-kraft/a1b2c3d4-5678-90ab-cdef-222222222222 auth: mode: iam backup: s3_bucket: prod-migration-segments s3_prefix: zk-to-kraft/ evidence: s3_bucket: prod-migration-evidence s3_prefix: migrations/ retention: 7y cutover: drain_timeout: 30m drain_max_partition_lag: 100 max_producer_freeze: 60s producer_freeze_webhook: https://internal-api.example.com/kafka/freeze validation: count_tolerance: 1 spot_check_records_per_partition: 5 seed: max_concurrent_partitions: 8 acl: on_drift: merge ``` ``` kafka-backup migrate msk-kraft plan \ --config migration.yaml \ --format all \ --out-dir ./migration-plan ``` Review the generated artifacts: - `migration-plan/runbook.md` — customized step-by-step runbook - `migration-plan/iam-policy-concrete.json` — attach this to the runner's IAM role - `migration-plan/cost-estimate.json` — estimated S3 costs (~$12 for 500GB) ``` aws iam put-role-policy \ --role-name migration-runner-role \ --policy-name kafka-migration \ --policy-document file://migration-plan/iam-policy-concrete.json ``` ``` kafka-backup migrate msk-kraft precheck --config migration.yaml ``` Expected output: no blockers. You may see: - **I01**: Target is IAM-auth — ACLs emitted as access-map.json (expected for IAM targets) - **W04**: Target message-size floor could not be verified from dynamic broker config (manually verify `message.max.bytes` and `replica.fetch.max.bytes`) Example precheck output from an IAM migration: ``` W04 warn: could not verify target message-size floor (target broker DescribeConfigs returned no message.max.bytes or replica.fetch.max.bytes (dynamic-config only on this broker)) — ensure target `message.max.bytes` and `replica.fetch.max.bytes` ≥ largest source topic's effective max.message.bytes W03 info: KMS key ARN set on backup channel — CMK access is not verified by this precheck phase; ensure the caller has kms:Encrypt/Decrypt/GenerateDataKey I01 info: target is IAM-auth — ACLs will be emitted as access-map.json for customer IaC to translate to IAM policies (tool does not apply IAM) ``` ``` kafka-backup migrate msk-kraft execute \ --config migration.yaml \ --journal-dir ./journal ``` For 500GB, expect: - Seed phase: ~3-4 hours - Tail convergence: ~10-15 minutes - Total to drain-ready: ~4 hours Example drain-ready output: ``` topology_copy -> seed seed -> tail tail -> drain_ready drain ready: max_partition_lag= records_replayed= bytes_replayed= ``` Coordinate with application teams, then: ``` kafka-backup migrate msk-kraft cutover \ --config migration.yaml \ --migration-id \ --journal-dir ./journal ``` The webhook at `https://internal-api.example.com/kafka/freeze` receives a POST request. Your application pauses producers for ~30 seconds while cutover completes. Example cutover-ready output: ``` cutover -> awaiting_client_switch READY_FOR_CLIENT_SWITCH: groups_translated= offsets_committed= warnings=0 ``` Update application configs to point to the KRaft cluster bootstrap servers: ``` # Get new bootstrap servers aws kafka get-bootstrap-brokers \ --cluster-arn arn:aws:kafka:us-east-1:123456789012:cluster/prod-kraft/a1b2c3d4-5678-90ab-cdef-222222222222 ``` Roll your deployments. Consumers resume from the translated target offsets, preserving message continuity across the switch. ``` # Acknowledge client switch kafka-backup migrate msk-kraft cutover-ack \ --config migration.yaml \ --migration-id \ --journal-dir ./journal # Finalize (runs validation + uploads evidence) kafka-backup migrate msk-kraft finalize \ --config migration.yaml \ --migration-id \ --journal-dir ./journal ``` ``` aws s3 cp \ "s3://prod-migration-evidence/migrations//evidence.json" \ ./evidence.json # Check validation outcome cat evidence.json | jq -r '.bundle_json' | jq '.validation.overall' # Expected: "PASSED", or "WARNING" when the report explains an accepted warning such as empty partitions with no spot-check sample. # Check per-check results cat evidence.json | jq -r '.bundle_json' | jq '{ topic_parity: .validation.topic_parity.outcome, counts_and_offsets: .validation.counts_and_offsets.outcome, offset_floor_violations: .validation.counts_and_offsets.data.offset_floor_violations, spot_check_records: .validation.spot_check_records.outcome, sentinel_presence: .validation.sentinel_presence.outcome, consumer_group_reconciliation: .validation.consumer_group_reconciliation.outcome }' ``` Optional source/target comparison after finalize: ``` topic_parity=PASSED counts_and_offsets=PASSED offset_floor_violations=0 sentinel_presence=PASSED consumer_group_reconciliation=PASSED ``` After verifying the migration is successful and all applications are stable on the KRaft cluster: ``` # Remove migration segments from S3 aws s3 rm s3://prod-migration-segments/zk-to-kraft/ --recursive # Decommission the source ZK cluster (when confident) aws kafka delete-cluster \ --cluster-arn arn:aws:kafka:us-east-1:123456789012:cluster/prod-zk/a1b2c3d4-5678-90ab-cdef-111111111111 ``` > [!WARNING] > > [!NOTE] > > Keep the evidence bucket > > [!NOTE] > > Do not delete the evidence bucket. The signed evidence bundle is your compliance proof that the migration succeeded. With Object Lock enabled, it's retained for 7 years automatically. - [Cross-Auth Example](https://kafkabackup.com/examples/msk-kraft-migration-cross-auth.md) — SCRAM to IAM migration - [Monitoring Guide](https://kafkabackup.com/guides/msk-kraft-monitoring.md) — what to watch during migration - [Troubleshooting](https://kafkabackup.com/troubleshooting/msk-kraft-migration.md) — common migration errors --- title: Example: MSK KRaft Migration (SCRAM to IAM) description: Worked example migrating an AWS MSK cluster from ZooKeeper with SCRAM-SHA-512 authentication to KRaft with IAM — cross-auth migration with ACL translation. source_url: html: https://kafkabackup.com/examples/msk-kraft-migration-cross-auth md: https://kafkabackup.com/examples/msk-kraft-migration-cross-auth.md --- # Example: MSK KRaft Migration (SCRAM to IAM) This example covers a cross-auth migration: SCRAM-SHA-512 on the source ZooKeeper cluster to IAM on the target KRaft cluster. This is common when modernizing legacy authentication alongside the ZK→KRaft migration. > [!NOTE] > > [!NOTE] > > Release evidence > > [!NOTE] > > Cross-auth is implemented and covered by access-map, endpoint-selection, token, and smoke tests. The May 8, 2026 full AWS release qualification proved SCRAM-SHA-512 to SCRAM-SHA-512 across four KRaft targets; use this page to rehearse SCRAM-to-IAM in staging before production cutover. | | Source | Target | | --- | --- | --- | | **Metadata mode** | ZooKeeper | KRaft | | **Kafka version** | 3.5.0 | 3.9.0 | | **Authentication** | SCRAM-SHA-512 | IAM | | **SCRAM users** | admin, order-svc, analytics-svc, billing-svc | — (IAM roles) | | **ACLs** | Per-user topic/group ACLs | IAM policies (generated) | | **Topics** | 25 | 0 (empty target) | When migrating from SCRAM to IAM: - Kafka ACLs (based on SCRAM usernames) don't apply on IAM-auth clusters - Each SCRAM user's permissions must be translated to an IAM policy - The tool generates an `access-map.json` that maps SCRAM principals to required IAM actions migration-cross-auth.yaml ``` enterprise: msk_kraft_migration: source: cluster_arn: arn:aws:kafka:eu-west-1:123456789012:cluster/legacy-scram/abc-123 auth: mode: scram-sha-512 username: ${KAFKA_ADMIN_USER} password: ${KAFKA_ADMIN_PASS} target: cluster_arn: arn:aws:kafka:eu-west-1:123456789012:cluster/modern-kraft/def-456 auth: mode: iam backup: s3_bucket: migration-segments-eu s3_prefix: scram-to-iam/ evidence: s3_bucket: migration-evidence-eu s3_prefix: migrations/ retention: 7y cutover: drain_timeout: 30m max_producer_freeze: 120s acl: on_drift: merge ``` > [!NOTE] > > [!NOTE] > > Environment variable interpolation > > [!NOTE] > > SCRAM credentials support `${ENV_VAR}` interpolation. Set `KAFKA_ADMIN_USER` and `KAFKA_ADMIN_PASS` in your environment before running. ``` export KAFKA_ADMIN_USER=admin export KAFKA_ADMIN_PASS= kafka-backup migrate msk-kraft plan \ --config migration-cross-auth.yaml \ --format all \ --out-dir ./migration-plan kafka-backup migrate msk-kraft precheck --config migration-cross-auth.yaml ``` Expected precheck findings: - **I01**: Target is IAM-auth — ACLs emitted as access-map.json - **W09**: MSK internal ACLs will be filtered (User:ANONYMOUS, etc.) These are expected and do not block migration. ``` kafka-backup migrate msk-kraft execute \ --config migration-cross-auth.yaml \ --journal-dir ./journal ``` ``` kafka-backup migrate msk-kraft cutover \ --config migration-cross-auth.yaml \ --migration-id \ --journal-dir ./journal ``` After cutover, the evidence bundle includes an `access-map.json` that maps each SCRAM principal's Kafka ACLs to the IAM actions needed: ``` { "principals": { "User:order-svc": { "topics": { "orders": ["kafka-cluster:ReadData", "kafka-cluster:WriteData", "kafka-cluster:DescribeTopic"], "order-events": ["kafka-cluster:ReadData", "kafka-cluster:DescribeTopic"] }, "groups": { "order-processing-cg": ["kafka-cluster:ReadGroup", "kafka-cluster:DescribeGroup"] } }, "User:analytics-svc": { "topics": { "orders": ["kafka-cluster:ReadData", "kafka-cluster:DescribeTopic"], "events": ["kafka-cluster:ReadData", "kafka-cluster:DescribeTopic"] }, "groups": { "analytics-cg": ["kafka-cluster:ReadGroup", "kafka-cluster:DescribeGroup"] } } } } ``` Use the access map to create IAM policies for each service: order-svc-kafka-policy.json ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "kafka-cluster:Connect", "kafka-cluster:DescribeCluster" ], "Resource": "arn:aws:kafka:eu-west-1:123456789012:cluster/modern-kraft/*" }, { "Effect": "Allow", "Action": [ "kafka-cluster:ReadData", "kafka-cluster:WriteData", "kafka-cluster:DescribeTopic" ], "Resource": "arn:aws:kafka:eu-west-1:123456789012:topic/modern-kraft/*/orders" }, { "Effect": "Allow", "Action": [ "kafka-cluster:ReadData", "kafka-cluster:DescribeTopic" ], "Resource": "arn:aws:kafka:eu-west-1:123456789012:topic/modern-kraft/*/order-events" }, { "Effect": "Allow", "Action": [ "kafka-cluster:ReadGroup", "kafka-cluster:DescribeGroup" ], "Resource": "arn:aws:kafka:eu-west-1:123456789012:group/modern-kraft/*/order-processing-cg" } ] } ``` ``` # Attach to each service's IAM role aws iam put-role-policy \ --role-name order-svc-role \ --policy-name kafka-access \ --policy-document file://order-svc-kafka-policy.json ``` Each service needs to switch from SCRAM to IAM authentication: **Before (SCRAM):** ``` security.protocol=SASL_SSL sasl.mechanism=SCRAM-SHA-512 sasl.jaas.config=org.apache.kafka.common.security.scram.ScramLoginModule required username="order-svc" password="secret"; ``` **After (IAM):** ``` security.protocol=SASL_SSL sasl.mechanism=AWS_MSK_IAM sasl.jaas.config=software.amazon.msk.auth.iam.IAMLoginModule required; sasl.client.callback.handler.class=software.amazon.msk.auth.iam.IAMClientCallbackHandler ``` ``` kafka-backup migrate msk-kraft cutover-ack \ --config migration-cross-auth.yaml \ --migration-id \ --journal-dir ./journal kafka-backup migrate msk-kraft finalize \ --config migration-cross-auth.yaml \ --migration-id \ --journal-dir ./journal ``` - [IAM-to-IAM Example](https://kafkabackup.com/examples/msk-kraft-migration-iam.md) — same-auth migration - [MSK KRaft Migration Overview](https://kafkabackup.com/enterprise/msk-kraft-migration.md) — full feature documentation - [Configuration Reference](https://kafkabackup.com/enterprise/msk-kraft-config-reference.md) — all config options