Implementing Kafka Tiered Storage for Cost-Effective Long-Term Data Retention

Introduction

Apache Kafka has established itself as a cornerstone technology in modern data streaming architectures. Renowned for its high-throughput, fault-tolerant, and distributed nature, Kafka enables organizations to process and analyze vast streams of real-time data efficiently. However, as data volume grows exponentially, managing long-term data retention within Kafka clusters presents significant challenges.

Traditional Kafka storage approaches, primarily relying on local disk storage, can become cost-prohibitive and operationally cumbersome for retaining large datasets over extended periods. Long-term retention poses risks such as exorbitant storage costs, reduced performance due to disk saturation, and increased maintenance overhead.

Kafka tiered storage offers an innovative solution by abstracting storage into multiple cost and performance tiers. This architecture allows hot, frequently accessed data to reside on fast local disks, while cold, infrequently accessed data is offloaded to more affordable external storage, such as cloud object stores. This not only optimizes storage costs but also extends cluster scalability and simplifies data lifecycle management.

In this article, we'll explore the concept of Kafka tiered storage, its architecture, and provide a detailed, practical guide to implementing tiered storage in your Kafka environment, ensuring cost-effective and scalable long-term data retention.

Understanding Kafka Tiered Storage

What is Kafka Tiered Storage?

Kafka tiered storage is a storage architecture enhancement that decouples Kafka's data retention from local disk storage to a layered storage approach. It enables Kafka brokers to offload older segments of data from expensive local storage to cheaper, scalable external object storage services, such as AWS S3, Azure Blob Storage, or Google Cloud Storage.

This design allows Kafka to maintain high write and read throughput by keeping recent data locally, while still providing the capability to access archived historical data stored externally when needed.

Key Components and Architecture

  1. Local Storage Tier: Low-latency, high-throughput local SSDs or HDDs on Kafka brokers store recent and frequently accessed data segments.
  1. Remote/Object Storage Tier: An external durable storage layer (cloud or on-prem object storage) holds older data segments that are infrequently accessed but retained for compliance, analytics, or backup.
  1. Segment Offloader: A Kafka broker component responsible for asynchronously offloading local segment files to the remote tier once they reach a defined retention age or size.
  1. Segment Fetcher: Retrieves data from remote storage dynamically in response to consumer requests targeting older segments, providing transparent access to offloaded data.
  1. Metadata Manager: Tracks the lifecycle and location of segments to manage retrieval and deletion seamlessly.

How Tiered Storage Differs From Traditional Kafka Storage

Traditional Kafka clusters store all data locally on broker disks with retention policies that truncate data after set periods or sizes. This approach means backing up or long-term storage requires external processes.

Tiered storage, in contrast, integrates external object storage directly into the Kafka storage lifecycle, enabling brokers to manage offloading and retrieval seamlessly without manual intervention. It provides infinite retention possibilities constrained primarily by external storage capacity and budget, not the broker’s disk space.

Preparing Your Kafka Environment for Tiered Storage

Kafka Version and Compatibility Requirements

Tiered storage is a relatively recent addition to Kafka. Ensure you are running Kafka version 3.4.0 or later, as prior versions lack native tiered storage support. Many tiered storage features rely on KIP-405 and KIP-405 enhancements.

Apache Kafka distributions like Confluent Kafka or Redpanda may have additional proprietary implementations; always consult your vendor documentation.

Configuring Storage Tiers and Brokers

Prepare your Kafka brokers' hardware and configurations to accommodate tiered storage. Consider:

  • Local Disk: Use fast and reliable disks (NVMe SSDs recommended) for hot data and for caching retrieved remote segments.
  • Network Configuration: Optimize network throughput and latency since brokers will interact with remote storage frequently.
  • Broker Configuration: Enable and configure tiered storage settings such as segment rollover sizes and retention periods.

Setting Up Compatible Cloud/Object Storage

Choose a compatible cloud or on-premises object storage solution supported by Kafka, such as:

  • AWS S3: A popular choice offering durability, scalability, and integration with IAM for secure access.
  • Azure Blob Storage: Azure-native object storage with powerful lifecycle management.
  • Google Cloud Storage: High availability and integrated IAM controls.

Ensure you have service credentials ready, appropriate bucket/container permissions configured, and network connectivity tested.

Step-by-Step Guide to Implementing Tiered Storage

1. Configure Kafka Brokers for Tiered Storage

Update your server.properties to include tiered storage configurations:

# Enable tiered storage
log.storage.backend=remote

# Rollover settings
log.segment.bytes=1073741824  # 1 GiB segment size
log.roll.ms=3600000            # 1 hour segment rollover

# Configure offload initiation parameters
log.retention.bytes=10737418240 # 10 GiB retention

# Remote storage bucket details
log.remote.storage.bucket=your-bucket-name

log.remote.storage.region=us-east-1

# Credentials and endpoint
log.remote.storage.access.key.id=YOUR_ACCESS_KEY
log.remote.storage.secret.access.key=YOUR_SECRET_KEY
log.remote.storage.endpoint=s3.amazonaws.com

# For security
log.remote.storage.sse.enabled=true

# Set cache size for locally cached remote segments
log.local.cache.size.bytes=5368709120  # 5 GiB

2. Integrate External Object Storage for Offloading Data

  • Create or identify your cloud storage bucket/container for Kafka data.
  • Configure access policies allowing Kafka brokers to PUT/GET objects.
  • Validate connectivity by uploading and downloading test objects.

Many Kafka tiered storage implementations include offloader plugins or modules to automate segment transfer.

3. Manage Retention Policies and Segment Rollover Settings

  • Define segment size and rollover intervals that balance granular offloading and performance.
  • Set retention thresholds that trigger segment offloads to remote storage, controlling local disk usage.
  • Use log compaction in conjunction for use cases needing last-update semantics.

4. Monitor and Troubleshoot Common Issues

  • Monitor broker logs for offload/upload failures.
  • Track lag and consumer fetch metrics to ensure remote segment retrievals do not introduce persistent stalls.
  • Watch cloud storage metrics and costs to detect anomalies.
  • Validate data integrity with checksum or hash verification.

Code Example: Configuring Kafka Tiered Storage

Here is a sample Kafka broker configuration snippet for AWS S3 tiered storage:

# Enable tiered storage using remote backend
log.storage.backend=remote

# Segment size and rollover
log.segment.bytes=1073741824
log.roll.ms=3600000

# Remote storage bucket configuration
log.remote.storage.bucket=my-kafka-archive
log.remote.storage.region=us-east-1

# AWS Credentials (ensure these are secured and ideally injected via env variables or IAM roles)
log.remote.storage.access.key.id=AKIAEXAMPLEKEY
log.remote.storage.secret.access.key=EXAMPLETESECRETKEY

# Enable server-side encryption
log.remote.storage.sse.enabled=true

# Cache remote segments locally
log.local.cache.size.bytes=4294967296

# Offload parameters
log.retention.bytes=21474836480

Example Shell Script for Manual Segment Offload (Simplified)

#!/bin/bash

LOCAL_SEGMENT_PATH="/var/lib/kafka/data/topic-partition/"
S3_BUCKET="s3://my-kafka-archive"

# Offload segments older than 7 days
find "$LOCAL_SEGMENT_PATH" -type f -name '*.log' -mtime +7 | while read -r segment; do
  aws s3 cp "$segment" "$S3_BUCKET/$(basename "$segment")" && rm "$segment"
  echo "Offloaded and deleted $segment"
done

Automating Retention and Cleanup with Kafka Admin API (Java)

import org.apache.kafka.clients.admin.AdminClient;
import org.apache.kafka.clients.admin.AdminClientConfig;
import org.apache.kafka.clients.admin.Config;
import org.apache.kafka.clients.admin.ConfigEntry;
import org.apache.kafka.clients.admin.ConfigResource;

import java.util.Collections;
import java.util.Properties;
import java.util.concurrent.ExecutionException;

public class KafkaRetentionManager {
    public static void main(String[] args) throws ExecutionException, InterruptedException {
        Properties props = new Properties();
        props.put(AdminClientConfig.BOOTSTRAP_SERVERS_CONFIG, "localhost:9092");

        try (AdminClient adminClient = AdminClient.create(props)) {
            ConfigResource resource = new ConfigResource(ConfigResource.Type.TOPIC, "your-topic");

            ConfigEntry retentionEntry = new ConfigEntry("retention.ms", String.valueOf(604800000)); // 7 days

            Config newConfig = new Config(Collections.singleton(retentionEntry));

            adminClient.alterConfigs(Collections.singletonMap(resource, newConfig)).all().get();

            System.out.println("Retention policy updated to 7 days for 'your-topic'");
        }
    }
}

Best Practices and Performance Optimization

Cost Management Strategies Using Tiered Storage

  • Fine-tune segment rollover and retention periods: Smaller segments enable faster offloads but increase overhead; balance these according to workload.
  • Choose appropriate storage classes: Use infrequent access or archival classes (e.g., S3 Glacier) for very old data.
  • Monitor and automate lifecycle policies: Automate data tier transitions to reduce costs.

Balancing Latency Versus Storage Costs

  • Cache frequently accessed cold data locally to reduce retrieval latency.
  • Use fast object storage endpoints or regional buckets to lower network latency.
  • Tune client consumer fetch configurations to handle occasional higher latency when accessing remote tiers.

Security Considerations for Data Stored in Tiers

  • Enable encryption both at rest and in transit for remote storage.
  • Use fine-grained IAM roles and policies to restrict access.
  • Audit access logs regularly.
  • Use Kafka’s ACLs alongside storage permissions to enforce layered security.

Conclusion

Kafka tiered storage represents a significant advancement for managing long-term data retention affordably and scalably. By decoupling data storage from local disks and leveraging cost-efficient external object stores, organizations can maintain extensive datasets without compromising Kafka’s real-time streaming performance.

Implementation involves ensuring compatible Kafka versions, configuring brokers for tiered offloading, integrating supported cloud storage, and tuning retention and segment policies. Additionally, applying best practices around cost, latency, and security ensures the solution is robust and sustainable.

As Kafka and cloud storage technologies evolve, expect tiered storage capabilities to become more sophisticated, further extending Kafka’s scalability and economic viability for data-driven enterprises.

FAQ

Q1: Which Kafka versions support tiered storage?

A: Kafka 3.4.0 and later include native tiered storage capabilities, though some features continue to evolve.

Q2: Can tiered storage be used with any cloud storage?

A: Kafka supports integration with popular object storage providers like AWS S3, Azure Blob, and Google Cloud Storage. Compatibility depends on the offloader plugins and configurations.

Q3: Does tiered storage add latency to data consumption?

A: Accessing remote stored segments can add some latency compared to local reads, but local caching and network optimization can minimize this impact.

Q4: How does tiered storage affect Kafka’s fault tolerance?

A: Tiered storage maintains Kafka’s fault tolerance by offloading immutable segment files. Proper configuration ensures data availability even during failures.

Q5: Is data in tiered storage encrypted?

A: Encryption is supported and should be enabled both in transit and at rest on the remote storage side.

References and Further Reading

Related reading