Revision note (2026-09-15). The earlier version of this article assumed "Kafka 2.5+" and never mentioned KRaft, told readers to add controlled.shutdown.enable=true and restart to apply it although it has been the default for years, checked cluster health with kafka-topics.sh --describe | grep UnderReplicated, which prints nothing on any Kafka version and so reports "healthy" while partitions are under-replicated, confirmed a broker's return with grep Leader | grep broker1 although the output only ever shows numeric broker ids, waited a fixed 60 seconds instead of checking the ISR, showed a kafka-reassign-partitions.sh --throttle ... --execute call without --bootstrap-server, and presented --verify as a pre-execution check. An even earlier revision used --zookeeper flags. This version is a runbook that was executed, one broker at a time, on a 3-node Kafka 4.0.0 KRaft cluster while a producer was writing with acks=all. The lab, scripts, and raw logs are in examples/kafka-rolling-restart of the site repository.
ZooKeeper commands are gone; here is what replaced them
Until Kafka 2.x, operators described topics with kafka-topics.sh --zookeeper zk1:2181 --describe, moved replicas with kafka-reassign-partitions.sh --zookeeper ..., and read the active controller from the /controller znode. Kafka 3.0 removed the --zookeeper option from kafka-topics and kafka-reassign-partitions (KIP-604; see the 3.0 section of the upgrade notes). Kafka 4.0 removed ZooKeeper mode entirely: the upgrade guide states that "Apache Kafka 4.0 only supports KRaft mode – ZooKeeper mode has been removed", and a cluster still on ZooKeeper must migrate to KRaft on 3.9 before it can run 4.0.
On Kafka 4.0.0 the old flag fails immediately:
$ kafka-topics.sh --zookeeper zk:2181 --describe
zookeeper is not a recognized option
The current equivalents used in this runbook:
| Question | Command (Kafka 4.0.0) |
|---|---|
| Which partitions are under-replicated? | kafka-topics.sh --bootstrap-server ... --describe --under-replicated-partitions |
| Who is the active controller? | kafka-metadata-quorum.sh --bootstrap-server ... describe --status (the LeaderId line) |
| Are all controllers caught up? | kafka-metadata-quorum.sh --bootstrap-server ... describe --replication |
| Which brokers answer requests? | kafka-broker-api-versions.sh --bootstrap-server ... |
| Move leadership back to preferred replicas | kafka-leader-election.sh --bootstrap-server ... --election-type PREFERRED --all-topic-partitions |
| What are the topic's end offsets? | kafka-get-offsets.sh --bootstrap-server ... --topic ... |
Pass all brokers in --bootstrap-server, comma-separated. During a rolling restart one of them is always down, and a single-address bootstrap list points at the wrong broker one third of the time.
The lab
- Apache Kafka 4.0.0 from the
apache/kafka:4.0.0image, three nodes in KRaft combined mode (process.roles=broker,controller), Docker 27.4.0 on macOS. Combined mode keeps the lab at three containers; the Kafka documentation recommends dedicated controllers for production, and in combined mode every broker restart is also a controller restart. - Topic
maintenance-demo: 6 partitions, replication factor 3,min.insync.replicas=2. - Producer:
kafka-producer-perf-test.sh, 60,000 records of 200 bytes at 300 records per second,acks=all,enable.idempotence=true, running from before the first stop until after the last restart. - Stop and start through
docker stop(SIGTERM) anddocker start, withstop_grace_period: 60sso Docker does not send SIGKILL while the broker is still handing off leadership.
What was not measured: production-scale replication catch-up, consumers, dedicated controller nodes, SIGKILL or crash restarts, and software upgrades. The numbers below describe a nearly idle cluster on one machine.
The runbook, per broker
For each broker in turn:
- Refuse to start if any partition is under-replicated: the output of
--describe --under-replicated-partitionsmust be empty. Restarting a second broker while a partition already has only two in-sync replicas of three leaves it with one, andacks=allwrites withmin.insync.replicas=2start failing withNotEnoughReplicas. - Note the active controller (
LeaderId). - Stop the broker with SIGTERM. Controlled shutdown is on by default; the broker asks the controller to move its leaderships away, and exits when the controller says so.
- Do the maintenance.
- Start the broker. Wait until it answers
kafka-broker-api-versions.sh, then wait until--under-replicated-partitionsis empty again. Only then move to the next broker.
After the last broker: check the controller and the ISR once more, then run a preferred leader election, because every restarted broker comes back as a follower for all of its partitions.
The script that does this (scripts/rolling-restart.sh in the lab) is about 50 lines. Its core:
BOOTSTRAP="dd-kafka-b-1:9092,dd-kafka-b-2:9092,dd-kafka-b-3:9092"
TOPIC=maintenance-demo
urp_count() {
kafka-topics.sh --bootstrap-server "$BOOTSTRAP" --describe --topic "$TOPIC"
--under-replicated-partitions 2>/dev/null | grep -c 'Topic:' || true
}
for id in 1 2 3; do
c="dd-kafka-b-$id"
if [ "$(urp_count)" != "0" ]; then echo "ABORT: URP=$(urp_count) before stopping $c"; exit 1; fi
kafka-metadata-quorum.sh --bootstrap-server "$BOOTSTRAP" describe --status | grep LeaderId
docker stop "$c" # SIGTERM -> controlled shutdown; on a host: kafka-server-stop.sh or systemctl stop kafka
# ... maintenance ...
docker start "$c"
until kafka-broker-api-versions.sh --bootstrap-server "$BOOTSTRAP" 2>/dev/null | grep -q "^$c:"; do sleep 1; done
until [ "$(urp_count)" = "0" ]; do sleep 1; done
done
kafka-leader-election.sh --bootstrap-server "$BOOTSTRAP" --election-type PREFERRED --all-topic-partitions
The health check counts lines of --under-replicated-partitions output. The earlier version of this article piped --describe into grep UnderReplicated; run against the lab with one broker stopped, that pipeline printed 0 while six partitions were under-replicated, because --describe never prints that word. A health check that cannot fail is worse than none.
What happened
Times are local; the shell log, the producer log, and the broker logs are in the lab's results directory.
Stopping broker 1
docker stop dd-kafka-b-1 returned after 3 seconds. Broker 1's log shows the controlled shutdown taking 2.2 seconds end to end:
[2026-09-15 04:46:15,587] INFO Terminating process due to signal SIGTERM (org.apache.kafka.common.utils.LoggingSignalHandler)
[2026-09-15 04:46:15,599] INFO [BrokerLifecycleManager id=1] Beginning controlled shutdown. (kafka.server.BrokerLifecycleManager)
[2026-09-15 04:46:17,756] INFO [BrokerLifecycleManager id=1] The controller has asked us to exit controlled shutdown. (kafka.server.BrokerLifecycleManager)
Immediately after, all six partitions were under-replicated and the two partitions that broker 1 led had moved:
$ kafka-topics.sh --bootstrap-server ... --describe --topic maintenance-demo
Topic: maintenance-demo TopicId: aWTNU98TQO6kiWWlT9h4XQ PartitionCount: 6 ReplicationFactor: 3 Configs: min.insync.replicas=2
Topic: maintenance-demo Partition: 0 Leader: 2 Replicas: 2,3,1 Isr: 2,3 Elr: LastKnownElr:
Topic: maintenance-demo Partition: 1 Leader: 3 Replicas: 3,1,2 Isr: 3,2 Elr: LastKnownElr:
Topic: maintenance-demo Partition: 2 Leader: 2 Replicas: 1,2,3 Isr: 2,3 Elr: LastKnownElr:
Topic: maintenance-demo Partition: 3 Leader: 3 Replicas: 3,1,2 Isr: 3,2 Elr: LastKnownElr:
Topic: maintenance-demo Partition: 4 Leader: 2 Replicas: 1,2,3 Isr: 2,3 Elr: LastKnownElr:
Topic: maintenance-demo Partition: 5 Leader: 2 Replicas: 2,3,1 Isr: 2,3 Elr: LastKnownElr:
Six under-replicated partitions is the expected state with one of three brokers down; it is the number the pre-check must see as zero before the next stop, not a reason to panic during this one. With two replicas still in sync and min.insync.replicas=2, acks=all writes kept succeeding.
Because the nodes are combined broker plus controller, stopping node 1 also stopped the active controller. kafka-metadata-quorum.sh describe --status went from LeaderId: 1 (epoch 1) to LeaderId: 2 (epoch 2). With dedicated controllers this line would not change during broker restarts.
Starting broker 1
After docker start, broker 1 answered kafka-broker-api-versions.sh 3 seconds later, and one second after that every ISR was back to three replicas:
Topic: maintenance-demo Partition: 2 Leader: 2 Replicas: 1,2,3 Isr: 2,3,1 Elr: LastKnownElr:
Topic: maintenance-demo Partition: 4 Leader: 2 Replicas: 1,2,3 Isr: 2,3,1 Elr: LastKnownElr:
Broker 1 is back in the ISR of partitions 2 and 4 but is no longer their leader. A restarted broker returns as a follower for everything it hosts. Kafka moves leadership back on its own when auto.leader.rebalance.enable is true (the default) and the imbalance check runs (every 300 seconds by default); the runbook does it explicitly at the end instead of waiting.
Brokers 2 and 3, and the end state
The same pattern repeated: docker stop returned in 3 seconds each time, URP went to 6 and the ISR shrank to two replicas per partition, the restarted broker answered in 4 seconds (broker 2) and 5 seconds (broker 3), and URP was back to 0 within 2 seconds of that. The controller moved to node 1 (epoch 3) when node 2 stopped and stayed there when node 3 stopped.
After the third restart the leaders were 2, 3, 1, 3, 1, 2 for partitions 0 to 5, and partitions 1 and 3 were led by a non-preferred replica. The election fixed them:
$ kafka-leader-election.sh --bootstrap-server ... --election-type PREFERRED --all-topic-partitions
Successfully completed leader election (PREFERRED) for partitions maintenance-demo-1, maintenance-demo-3
A note on wall-clock time: the script printed 37 to 39 seconds of "downtime" per broker, but most of that is the script itself. Every kafka-*.sh invocation starts a JVM and takes 2 to 5 seconds, and the script runs several between stop and start to record what is shown here. The broker-side figures are the ones that matter: about 2 seconds to hand off leadership, 3 to 5 seconds to answer requests after start, 1 to 2 seconds more to rejoin every ISR, on an almost idle cluster. On a busy broker the last number grows with the amount of log it has to fetch to catch up, which is exactly why the runbook waits for URP to reach zero instead of sleeping for 60 seconds.
What the producer saw
The producer's final line:
60000 records sent, 300.0 records/sec (0.06 MB/sec), 13.23 ms avg latency, 1373.00 ms max latency, 7 ms 50th, 12 ms 95th, 139 ms 99th, 1204 ms 99.9th.
Between restarts the 5-second windows reported 5 to 7 ms average latency; the windows that overlapped a broker stop reported up to 103 ms average and 985 ms maximum. No send failed. The producer logged bursts of
WARN [Producer clientId=perf-producer-client] Got error produce response with correlation id 5790 on topic-partition maintenance-demo-5, retrying (2147483646 attempts left). Error: NOT_LEADER_OR_FOLLOWER
at four moments, each matching a leadership change: the three broker stops and the preferred leader election at the end. Each burst was one to eight retries before the producer refreshed its metadata. The election is a leadership change too, so expect the same warnings from it and schedule it inside the maintenance window, not after it.
End offsets after the run:
maintenance-demo:0:10246
maintenance-demo:1:9037
maintenance-demo:2:11839
maintenance-demo:3:10416
maintenance-demo:4:8545
maintenance-demo:5:9917
They sum to 60,000, the number of records the producer reported as sent: nothing lost, nothing duplicated. The three settings that made that true are replication factor 3, min.insync.replicas=2, and acks=all with idempotence on the producer. A topic with replication factor 1 has no replica to fail over to, and a producer with acks=1 can lose the records the old leader acknowledged but had not yet replicated.
Three more things the earlier version got wrong
controlled.shutdown.enable does not need to be added or applied. On 4.0.0, kafka-configs.sh --bootstrap-server ... --describe --entity-type brokers --entity-name 1 --all shows controlled.shutdown.enable=true with DEFAULT_CONFIG:controlled.shutdown.enable=true. The setting only needs attention if someone turned it off. What can defeat it is a stop that does not send SIGTERM (a kill -9, or a container runtime that gives up before the handoff completes), and a partition with no other live replica: the operations guide notes that controlled shutdown only succeeds when every partition on the broker has another live replica, which is one more reason to keep replication factor above 1.
kafka-reassign-partitions.sh --verify is not a dry run. Run against a plan that had never been executed, it printed Reassignment of partition maintenance-demo-0 is completed. and then Clearing broker-level throttles on brokers 1,2,3. It reports completion of an executed reassignment and removes the throttles that --execute --throttle set; running it "to check the plan first" clears throttles on any reassignment that is in progress. The --execute call also needs --bootstrap-server; without it, 4.0.0 answers Please specify either --bootstrap-server or --bootstrap-controller.
Rebalancing after maintenance is usually not needed. A rolling restart does not move replicas, only leaders, and a preferred leader election (or the automatic one) restores the leader distribution. Replica reassignment is for capacity changes, and Cruise Control is a separate project with its own operational cost; neither was run for this article, so neither is described here.
Sources
- Apache Kafka 4.0 upgrade guide: ZooKeeper mode removed
- Apache Kafka 3.9 upgrade notes, section for 3.0: –zookeeper removed from kafka-topics and kafka-reassign-partitions
- KIP-604: Remove ZooKeeper Flags from the Administrative Tools
- Apache Kafka 4.0 operations: graceful shutdown and balancing leadership
- Apache Kafka 4.0 KRaft operations: process roles and combined mode
- Apache Kafka 4.0 broker configuration reference: controlled.shutdown.enable
