Revision note (2026-09-15). The earlier version of this article called admin.alterClientQuotas(Map<ClientQuotaEntity, Map<String, Double>>) and admin.describeClientQuotas(Collections.singleton(entity)).values(), neither of which exists in any Kafka client (the API takes a Collection<ClientQuotaAlteration> and a ClientQuotaFilter; both snippets fail to compile against kafka-clients 4.3.1). It told readers to watch kafka.network:type=Quota,...,name=QuotaThrottleTimeMs, an MBean that does not exist, said throttled clients "retry with exponential backoff" (they do not; the broker delays the response and the client waits), described request quotas as "not CPU" when they are defined as a percentage of broker thread time, required "Kafka 2.3+" and "Java 8+" for an API added in 2.6 on a client that needs Java 11, and cited a Confluent blog post and a KIP page that both return 404. This version replaces all of it with what a Kafka 4.3.1 broker actually did in a lab: nine integration tests, a kafka-producer-perf-test.sh run and a JMX dump, all pasted as printed.
What a quota is, and what it is not
A Kafka client quota is a per-broker limit on one client group. The Kafka design documentation defines two kinds: network bandwidth quotas, a byte-rate threshold for produce and fetch traffic, and request rate quotas, "the percentage of time a client can utilize on request handler I/O threads and network threads of each broker within a quota window". A client group is a user principal, a client-id, or the pair. Quotas are not per topic and not per partition, and they are enforced by each broker independently, so a 2 MB/s quota on a three-broker cluster is 2 MB/s on each broker.
Two consequences follow that the earlier version got wrong. First, quotas are dynamic: the design docs say the overrides "are written to the metadata log … and are effective immediately", and in the lab every quota took effect for the next run without touching the broker. Second, throttling is not an error path. The docs describe the mechanism: "the broker first computes the amount of delay needed to bring the violating client under its quota and returns a response with the delay immediately … Then, the broker mutes the channel to the client." The client sees a slow response, not a failure, and nothing is retried.
Tested versions: apache/kafka:4.3.1 Docker image running a single KRaft node with a PLAINTEXT listener and a SASL/PLAIN listener, kafka-clients 4.3.1, Java 17. Broker defaults were left alone (quota.window.num=11, quota.window.size.seconds=1).
The three quota keys and where they can be set
For the users and clients entity types, kafka-configs.sh --help on 4.3.1 lists four keys: producer_byte_rate, consumer_byte_rate, request_percentage, and controller_mutation_rate (topic create/delete/alter rate, KIP-599). The ips entity type has one key, connection_creation_rate, and there is no --zookeeper option anywhere:
$ kafka-configs.sh --zookeeper localhost:2181 --describe --entity-type users
zookeeper is not a recognized option
Quotas can be set for a named user, a named client-id, a (user, client-id) pair, or the default of any of those levels. Each of these commands ran against the lab broker with the output shown:
$ kafka-configs.sh --bootstrap-server localhost:9098 --alter --add-config producer_byte_rate=2000000,consumer_byte_rate=1000000 --entity-type users --entity-name tenantUserA
Completed updating config for user tenantUserA.
$ kafka-configs.sh --bootstrap-server localhost:9098 --alter --add-config request_percentage=50 --entity-type clients --entity-name reporting-job
Completed updating config for client reporting-job.
$ kafka-configs.sh --bootstrap-server localhost:9098 --alter --add-config producer_byte_rate=5000000 --entity-type users --entity-default
Completed updating default config for users in the cluster.
$ kafka-configs.sh --bootstrap-server localhost:9098 --alter --add-config producer_byte_rate=500000 --entity-type users --entity-name tenantUserA --entity-type clients --entity-name batch-loader
Completed updating config for user tenantUserA.
$ kafka-configs.sh --bootstrap-server localhost:9098 --describe --entity-type users --entity-name tenantUserA
Quota configs for user-principal 'tenantUserA' are consumer_byte_rate=1000000.0, producer_byte_rate=2000000.0
$ kafka-configs.sh --bootstrap-server localhost:9098 --describe --entity-type users --entity-name tenantUserA --entity-type clients --entity-name batch-loader
Quota configs for user-principal 'tenantUserA', client-id 'batch-loader' are producer_byte_rate=500000.0
$ kafka-configs.sh --bootstrap-server localhost:9098 --alter --delete-config producer_byte_rate --entity-type users --entity-name tenantUserA
Completed updating config for user tenantUserA.
When more than one level matches a connection, the most specific wins. The design docs list the order: (user, client-id), then (user, default client-id), then user, then (default user, client-id), then (default user, default client-id), then default user, then client-id, then default client-id. The lab checked one step of that ladder: with a default user quota of 2 MB/s, tenant-c was throttled (12,117 ms for 30 MB); after adding a named tenant-c quota of 200 MB/s, the same run took 644 ms with zero throttle time.
The Admin API as it exists
The Java API for quotas arrived with KIP-546 in Kafka 2.6.0 and has two methods. alterClientQuotas takes a collection of ClientQuotaAlteration, each an entity plus a list of Op(key, value); a null value removes that key. describeClientQuotas takes a ClientQuotaFilter and returns a result whose entities() future resolves to a map from entity to quota values. A null entity name means the default for that entity type.
Map<String, String> who = new HashMap<>();
who.put(ClientQuotaEntity.USER, "tenantUserA"); // null value here would mean <default user>
ClientQuotaEntity user = new ClientQuotaEntity(who);
List<ClientQuotaAlteration.Op> ops = List.of(
new ClientQuotaAlteration.Op("producer_byte_rate", 2_000_000.0),
new ClientQuotaAlteration.Op("consumer_byte_rate", 1_000_000.0));
admin.alterClientQuotas(List.of(new ClientQuotaAlteration(user, ops))).all().get();
ClientQuotaFilter filter = ClientQuotaFilter.containsOnly(
List.of(ClientQuotaFilterComponent.ofEntity(ClientQuotaEntity.USER, "tenantUserA")));
Map<ClientQuotaEntity, Map<String, Double>> described =
admin.describeClientQuotas(filter).entities().get();
The lab printed the result of that describe call and then read the same entity with the CLI:
[lab] describeClientQuotas(user=tenantUserA) -> {ClientQuotaEntity(entries={user=tenantUserA})={consumer_byte_rate=1000000.0, producer_byte_rate=2000000.0}}
Quota configs for user-principal 'tenantUserA' are consumer_byte_rate=1000000.0, producer_byte_rate=2000000.0
The earlier version's two snippets were copied verbatim and compiled against kafka-clients 4.3.1:
KafkaQuotaSetter.java:22: error: incompatible types: Map<ClientQuotaEntity,Map<String,Double>> cannot be converted to Collection<ClientQuotaAlteration>
KafkaQuotaDescriber.java:13: error: incompatible types: no instance(s) of type variable(s) T exist so that Set<T> conforms to ClientQuotaFilter
A reflection test in the lab lists the only overloads: alterClientQuotas(Collection) and alterClientQuotas(Collection, AlterClientQuotasOptions); describeClientQuotas(ClientQuotaFilter) and describeClientQuotas(ClientQuotaFilter, DescribeClientQuotasOptions).
What throttling does to a producer
Each measured run sends 30,000 records of 1,000 bytes (about 30 MB) with acks=1, linger.ms=5, batch.size=65536, waits for every acknowledgement, and then reads the producer's own metrics. With no quota, and then with producer_byte_rate=2000000 on the client-id:
[lab] no quota, client-id=unthrottled-producer-72efe79f: 357 ms, throttle-time max=0 ms avg=0.0 ms, retries=0, errors=0
[lab] producer_byte_rate=2000000 on client-id=throttled-producer-9b7ca577: 12100 ms, throttle-time max=1031 ms avg=24.7 ms, retries=0, errors=0
Three things to read off that. The run took 34 times longer. produce-throttle-time-max reached 1,031 ms, which is the broker's delay on a single response. And record-retry-total and record-error-total stayed at zero: the producer was never told to retry anything, it was simply not answered for a while. The same shape appears on the consumer side with consumer_byte_rate=2000000 and a consumer reading the same 30 MB with assign():
[lab] no quota, client-id=unthrottled-consumer-77d31d98: 243 ms, throttle-time max=0 ms avg=0.0 ms, retries=0, errors=0
[lab] consumer_byte_rate=2000000 on client-id=throttled-consumer-59aa88b0: 11670 ms, throttle-time max=996 ms avg=241.3 ms, retries=0, errors=0
kafka-producer-perf-test.sh shows where the time goes. Unthrottled, 30,000 records went out at 81.51 MB/s with a 99th percentile latency of 86 ms. With the 2 MB/s quota:
22017 records sent, 4183.4 records/sec (3.99 MB/sec), 118.3 ms avg latency, 5017.0 ms max latency.
320 records sent, 62.0 records/sec (0.06 MB/sec), 7100.2 ms avg latency, 10180.0 ms max latency.
30000 records sent, 2465.078061 records/sec (2.35 MB/sec), 3176.78 ms avg latency, 11892.00 ms max latency, 104 ms 50th, 11882 ms 95th, 11890 ms 99th, 11892 ms 99.9th.
producer-metrics:produce-throttle-time-avg:{client-id=perf-throttled-20814} : 24.464
producer-metrics:produce-throttle-time-max:{client-id=perf-throttled-20814} : 1032.000
producer-metrics:record-retry-total:{client-id=perf-throttled-20814} : 0.000
The 99th percentile latency went from 86 ms to 11.9 s. That is what a quota looks like from an application: send() callbacks arrive late, the accumulator fills, and eventually buffer.memory is exhausted and send() blocks for up to max.block.ms. There is no exception to catch and no backoff to tune. The advice in the earlier version, "clients aggressively retry on throttle; use exponential backoff", would have sent readers looking for a knob that does nothing here.
The burst allowance
The perf run above also shows something the earlier version did not mention: 22,017 records, about 22 MB, went through in the first five seconds at 4 MB/s before the broker started delaying a 2 MB/s client. Byte rates are measured over a sliding set of windows (default 11 windows of 1 second), and the rate is computed over the whole span, so a fresh client can move roughly ten times its per-second quota before the measured rate crosses the line. A separate test made this explicit:
[lab] producer_byte_rate=1000000, 6 MB in one burst on client-id=bursty-47610705: 112 ms, throttle-time max=0 ms avg=0.0 ms, retries=0, errors=0
Six megabytes in one burst under a 1 MB/s quota was not throttled at all. If you set quotas to protect a broker from a short spike, the default windows will let the spike through; the design docs describe the trade-off ("large measurement windows … leads to large bursts of traffic followed by long delays") and the window settings are quota.window.num and quota.window.size.seconds, both read-only broker properties. The lab did not sweep these; the ten-times figure is what the defaults produced here.
A user quota needs an authenticated user
The users entity type matches the authenticated principal, not anything the client sends in its configuration. On a PLAINTEXT listener every connection is User:ANONYMOUS. The lab set producer_byte_rate=2000000 for user tenant-a and ran three producers:
[lab] user quota tenant-a; PLAINTEXT listener, client-id=tenant-a (principal User:ANONYMOUS): 828 ms, throttle-time max=0 ms avg=0.0 ms, retries=0, errors=0
[lab] user quota tenant-a; SASL as tenant-a, client-id=app-1: 12161 ms, throttle-time max=1031 ms avg=24.8 ms, retries=0, errors=0
[lab] user quota tenant-a; SASL as tenant-b, client-id=app-1: 505 ms, throttle-time max=0 ms avg=0.0 ms, retries=0, errors=0
A client that merely names itself tenant-a in client.id over an unauthenticated listener is untouched. Only the SASL connection authenticated as tenant-a was throttled, and tenant-b on the same client-id was not. If your cluster has an unauthenticated listener, user quotas do not apply to traffic on it; use clients quotas there, or the default user quota, and know that a client-id is chosen by the client and can be anything.
Request quotas are CPU quotas
request_percentage is not a request count. The design docs define it as "the percentage of time a client can utilize on request handler I/O threads and network threads of each broker within a quota window", out of a total of (num.io.threads + num.network.threads) * 100 percent, and state that "request rate quotas represent the total percentage of CPU that may be used by each group of clients sharing the quota." The earlier version's claim that quotas "do not directly manage CPU" is the opposite of the definition. With request_percentage=1 on a client-id, 10 MB of produce traffic took 12.2 s with a maximum throttle of 1,000 ms and no errors:
[lab] request_percentage=1, 10 MB on client-id=cpu-bound-2e81cb79: 12218 ms, throttle-time max=1000 ms avg=74.6 ms, retries=0, errors=0
That run is not the whole story. Repeating it eight times on the same machine, the batched 10 MB producer was throttled four times (11.7 to 37.4 s) and finished untouched four times (68 to 204 ms, throttle-time 0). The quota meters handler and network thread time per 1 s window; a producer that packs 10 MB into a handful of 64 KB batches costs the broker very little of a thread, and whether it crosses 1 % depends on how the batches land in the windows. Sending 300 records one per request (linger.ms=0, batch.size=1) throttled in three of four runs (23.3 s, 12.0 s, 0.2 s with throttle-time max 1000, 1000 and 23 ms) and not in the fourth (45 ms). So on a single laptop a 1 % request quota is a coin flip for small traffic, and the lab test asserts only what is deterministic: the quota is stored, requests are delayed rather than failed, and the client never retries. A first attempt with 20,000 one-per-request records was throttled every time but expired thousands of records against the client's 120 s delivery timeout: a request quota this tight turns a chatty producer's backlog into TimeoutExceptions, a failure mode to plan for rather than a bug. In production the quota is meant for clients that keep a broker's request handlers busy for seconds at a time, where the thread-time accounting is not noisy.
The Kafka multi-tenancy guide argues that in shared clusters request quotas often matter more than bandwidth quotas, "because excessive broker CPU usage for processing requests reduces the effective bandwidth the broker can serve." A client sending many tiny requests can stay far under a byte quota while consuming a large share of the request handler threads.
Monitoring: the MBeans that exist
The Kafka 4.3 monitoring reference lists the broker-side quota metrics as kafka.server:type={Produce|Fetch},user=...,client-id=... with attributes throttle-time (ms, "Ideally = 0") and byte-rate, and kafka.server:type=Request,... with throttle-time and request-time. The name attribute of the MBean follows the quota level, so a client-id quota produces kafka.server:type=Produce,client-id=<id>. JmxTool against the lab broker right after the perf run:
### kafka.server:type=Produce,client-id=*
kafka.server:type=Produce,client-id=perf-throttled-20814:byte-rate=795583.6511195854
kafka.server:type=Produce,client-id=perf-throttled-20814:throttle-time=213.83333333333334
### kafka.server:type=Request,client-id=*
kafka.server:type=Request,client-id=cpu-bound-2e81cb79:request-time=0.0
kafka.server:type=Request,client-id=cpu-bound-2e81cb79:throttle-time=NaN
### kafka.network:type=Quota,*
No matched attributes for the queried objects [kafka.network:type=Quota,*].
Two practical notes from that dump. The sensors are per client group and expire after a client goes idle, so clients from earlier runs showed byte-rate=0.0 and throttle-time=NaN; a dashboard has to treat NaN as "no recent traffic", not as zero. And the client-side view is often the more useful one for application teams: produce-throttle-time-avg/max on kafka.producer:type=producer-metrics and fetch-throttle-time-avg/max on kafka.consumer:type=consumer-fetch-manager-metrics are what the runs above read, and they are visible to the team that owns the client without broker access.
What this does not cover
- Multiple brokers. Quotas are per broker by design; this lab has one, so nothing here shows how a quota behaves as partitions move between brokers.
(user, client-id)quotas were configured and described but not driven under load.controller_mutation_rateand theipsconnection-rate quota.- The exact size of the burst allowance under non-default windows.
- Non-Java clients. Older or third-party clients that ignore the throttle delay are still held back by the muted channel, per the design docs, but that was not measured here.
- Anything about Kafka 2.x or 3.x. Every number above comes from 4.3.1.
Reproduce it
The lab is examples/kafka-quotas in the site repository: a Gradle project with QuotaTests.java (nine tests, each named for the behavior it pins), perf-demo.sh, jmx-check.sh, and a server.properties for a single KRaft node with a PLAINTEXT and a SASL/PLAIN listener.
./start-broker.sh
gradle test --no-daemon --rerun-tasks # 9 tests, about 90 s
./perf-demo.sh
./jmx-check.sh
./stop-broker.sh
Sources
- Apache Kafka 4.3 design: Quotas (quota types, client groups, precedence order, enforcement, windows)
- Apache Kafka 4.3 operations: Setting quotas (
kafka-configs.shexamples) - Apache Kafka 4.3 operations: Multi-tenancy, Isolating tenants
- Apache Kafka 4.3 operations: Monitoring (quota MBean names and client throttle metrics)
- Apache Kafka 4.3 broker configs (
quota.window.num,quota.window.size.seconds,num.io.threads,num.network.threads) - Apache Kafka 4.3 operations: Java version
- KIP-546: Add Client Quota APIs to the Admin Client
- Admin interface, kafka-clients 4.3 Javadoc
