Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
69 changes: 45 additions & 24 deletions docs/modules/kafka/pages/usage-guide/kraft-controller.adoc
Original file line number Diff line number Diff line change
@@ -1,8 +1,6 @@
= KRaft mode (experimental)
= KRaft mode
:description: Apache Kafka KRaft mode with the Stackable Operator for Apache Kafka

WARNING: The Kafka KRaft mode is currently experimental, and subject to change.

Apache Kafka's KRaft mode replaces Apache ZooKeeper with Kafka’s own built-in consensus mechanism based on the Raft protocol.
This simplifies Kafka’s architecture, reducing operational complexity by consolidating cluster metadata management into Kafka itself.

Expand All @@ -18,8 +16,7 @@ WARNING: The Stackable Operator for Apache Kafka currently does not support auto

== Configuration

The Stackable Kafka operator introduces a new xref:concepts:roles-and-role-groups.adoc[role] in the KafkaCluster CRD called KRaft `Controller`.
Configuring the `Controller` will put Kafka into KRaft mode. Apache ZooKeeper will not be required anymore.
Enable KRaft mode by adding `spec.clusterConfig.metadataManager: KRaft` and a `Controller` xref:concepts:roles-and-role-groups.adoc[role] to your cluster manifest as in the example below:

[source,yaml]
----
Expand All @@ -29,7 +26,7 @@ metadata:
name: kafka
spec:
clusterConfig:
metadataManager: kraft
metadataManager: KRaft
image:
productVersion: "3.9.2"
brokers:
Expand All @@ -42,22 +39,39 @@ spec:
replicas: 3
----

NOTE: Using `spec.controllers` is mutually exclusive with `spec.clusterConfig.zookeeperConfigMapName`.
NOTE: Using `spec.controllers` is mutually exclusive with `spec.clusterConfig.zookeeperConfigMapName`. Configuring both would result in an error.

=== Recommendations
=== Resources

A minimal KRaft setup consisting of at least 3 Controllers has the following https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/[resource requirements]:
Each Pod's https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/[resource requirements]
are the sum of its containers. Controller Pods run an additional `quorum-manager` sidecar
(see <<_internal_operator_details, Internal operator details>>) whose resources are fixed and *not* configurable through
`spec.controllers.config.resources`. Only the `kafka` container is eligible.

* `600m` CPU request
* `3000m` CPU limit
* `3000Mi` memory request and limit
* `6Gi` persistent storage
[cols="2,3,2,2,2,2", options="header"]
|===
| Role | Container | CPU request | CPU limit | Memory (request = limit) | Storage (`logDirs`)

NOTE: The Controller replicas should sum up to an odd number for the Raft consensus.
| Controller | `kafka` (configurable) | 250m | 1000m | 1Gi | 2Gi
| Controller | `quorum-manager` (fixed) | 100m | 500m | 512Mi | --
| Controller | *Pod total* | *350m* | *1500m* | *1.5Gi* | *2Gi*
| Broker | `kafka` (configurable) | 250m | 1000m | 2Gi | 2Gi
|===

=== Resources
A minimal KRaft setup consists of 3 Controllers and 3 Brokers:

[cols="2,2,2,2,2", options="header"]
|===
| | CPU request | CPU limit | Memory | Storage

Corresponding to the values above, the operator uses the following resource defaults:
| 3 Controllers | 1050m | 4500m | 4.5Gi | 6Gi
| 3 Brokers | 750m | 3000m | 6Gi | 6Gi
| *Total* | *1800m* | *7500m* | *10.5Gi* | *12Gi*
|===

NOTE: The Controller replicas should sum up to an odd number for the Raft consensus.

Corresponding to the `kafka` container values above, the operator uses the following resource defaults:

[source,yaml]
----
Expand All @@ -72,15 +86,22 @@ controllers:
storage:
logDirs:
capacity: 2Gi
brokers:
config:
resources:
memory:
limit: 2Gi
cpu:
min: 250m
max: 1000m
storage:
logDirs:
capacity: 2Gi
----

=== Overrides

The configuration of overrides, JVM arguments etc. is similar to the Broker and documented on the xref:concepts:overrides.adoc[concepts page].

== Internal operator details

KRaft mode requires major configuration changes compared to ZooKeeper:
KRaft mode requires major configuration changes compared to ZooKeeper mode:

* `cluster-id`: This is set to the `metadata.name` of the KafkaCluster resource during initial formatting
* `node.id`: This is a calculated integer, hashed from the `role` and `rolegroup` and added `replica` id.
Expand All @@ -93,7 +114,7 @@ KRaft mode requires major configuration changes compared to ZooKeeper:
Every other controller formats with `--no-initial-controllers` and joins purely through the sidecar's `add-controller` call.
Brokers always format with `--no-initial-controllers` too; they are never voters.
* `controller.quorum.bootstrap.servers` points at each controller role group's own headless Service DNS name, not
individual pod addresses -- unless Kerberos is enabled. See below for details related to Kerberos.
individual pod addresses, unless Kerberos is enabled. See below for details related to Kerberos.

== Kerberos

Expand Down Expand Up @@ -146,8 +167,8 @@ The example cluster will be kept minimal without any additional configuration.

We'll use Kafka version `3.9.2` for this purpose. This is because this is the last version from the 3.x Kafka series that runs on ZooKeeper mode and is supported by the SDP.

We'll also assign broker ids manually from the beginning to simplify this guide. In a real-workd scenario, you do not have this option at this step because your cluster is already running.
In a real world-scenario you'll have to collect these ids and configure manual assignment at the second step of the migration.
We'll also assign broker ids manually from the beginning to simplify this guide. In a real-world scenario, you do not have this option at cluster creation because your cluster is already running.
Instead, you'll have to collect these ids from the running cluster and configure manual assignment as part of step 1 below.

We start by creating a dedicated namespace to work in and deploy the Kafka cluster including ZooKeeper and credentials.

Expand Down
166 changes: 166 additions & 0 deletions rust/operator-binary/src/controller/build/resource/statefulset.rs
Original file line number Diff line number Diff line change
Expand Up @@ -1241,6 +1241,172 @@ mod tests {
.expect("the kafka container is built")
}

/// `spec.controllers.config.resources` and `spec.brokers.config.resources` must be applied
/// verbatim to each role's `kafka` container: CPU request/limit come from `cpu.min`/`cpu.max`
/// and memory request equals memory limit, both taken from `memory.limit` (there is no
/// separate memory request field in the CRD).
#[test]
fn kafka_container_resources_match_configured_role_resources() {
let kafka = minimal_kafka(
r#"
apiVersion: kafka.stackable.tech/v1alpha1
kind: KafkaCluster
metadata:
name: simple-kafka
namespace: default
uid: 12345678-1234-1234-1234-123456789012
spec:
image:
productVersion: 3.9.2
clusterConfig:
metadataManager: kraft
controllers:
config:
resources:
memory:
limit: 2Gi
cpu:
min: 260m
max: 1024m
storage:
logDirs:
capacity: 2Gi
roleGroups:
default:
replicas: 3
brokers:
config:
resources:
memory:
limit: 2Gi
cpu:
min: 300m
max: 500m
storage:
logDirs:
capacity: 2Gi
roleGroups:
default:
replicas: 3
"#,
);
let cluster = validated_cluster(&kafka);

let controller_resources = controller_kafka_container(&cluster)
.resources
.expect("the controller kafka container has resources");
assert_cpu_and_memory(&controller_resources, "260m", "1024m", "2Gi", "2Gi");

let broker_resources = broker_kafka_container(&cluster)
.resources
.expect("the broker kafka container has resources");
assert_cpu_and_memory(&broker_resources, "300m", "500m", "2Gi", "2Gi");
}

/// Asserts the container's CPU request/limit and memory request/limit against the given
/// quantity strings (memory request and limit are always equal, see the CRD's
/// `Resources`/`MemoryLimits` type).
fn assert_cpu_and_memory(
resources: &stackable_operator::k8s_openapi::api::core::v1::ResourceRequirements,
expected_cpu_request: &str,
expected_cpu_limit: &str,
expected_memory_request: &str,
expected_memory_limit: &str,
) {
use stackable_operator::k8s_openapi::apimachinery::pkg::api::resource::Quantity;

let requests = resources
.requests
.as_ref()
.expect("the container has resource requests");
let limits = resources
.limits
.as_ref()
.expect("the container has resource limits");

assert_eq!(
requests.get("cpu"),
Some(&Quantity(expected_cpu_request.to_string())),
"cpu request"
);
assert_eq!(
limits.get("cpu"),
Some(&Quantity(expected_cpu_limit.to_string())),
"cpu limit"
);
assert_eq!(
requests.get("memory"),
Some(&Quantity(expected_memory_request.to_string())),
"memory request"
);
assert_eq!(
limits.get("memory"),
Some(&Quantity(expected_memory_limit.to_string())),
"memory limit"
);
}

/// Unlike the `kafka` container, the `quorum-manager` sidecar's resources are hardcoded in
/// [`build_quorum_manager_container`] and not exposed through `config.resources` (which only
/// applies to the `kafka` container, see
/// [`kafka_container_resources_match_configured_role_resources`]). `podOverrides` is the only
/// way to change them, applied via a strategic merge on `spec.containers` keyed by container
/// `name` (see the `merge_from` call at the end of `build_controller_rolegroup_statefulset`).
#[test]
fn quorum_manager_resources_can_be_changed_by_pod_overrides() {
let kafka = minimal_kafka(
r#"
apiVersion: kafka.stackable.tech/v1alpha1
kind: KafkaCluster
metadata:
name: simple-kafka
namespace: default
uid: 12345678-1234-1234-1234-123456789012
spec:
image:
productVersion: 3.9.2
clusterConfig:
metadataManager: kraft
controllers:
podOverrides:
spec:
containers:
- name: quorum-manager
resources:
requests:
cpu: 200m
memory: 256Mi
limits:
cpu: 750m
memory: 768Mi
roleGroups:
default:
replicas: 3
brokers:
roleGroups:
default:
replicas: 3
"#,
);
let cluster = validated_cluster(&kafka);

let containers = controller_containers(&cluster);
let sidecar = containers
.iter()
.find(|c| c.name == QUORUM_MANAGER_CONTAINER_NAME.to_string())
.expect("the quorum-manager sidecar is built");
let resources = sidecar
.resources
.as_ref()
.expect("the override gives the sidecar resources");

// Request and limit are asserted independently (unlike
// `kafka_container_resources_match_configured_role_resources`'s memory check) to show
// that `podOverrides` sets the raw Kubernetes fields directly, without the
// request-equals-limit constraint the CRD's `config.resources` memory field imposes.
assert_cpu_and_memory(resources, "200m", "750m", "256Mi", "768Mi");
}

/// The startup and liveness probes share the same check - TCP reachability (a genuinely
/// dead/hung process must still be restarted) plus the broker's JMX `BrokerState` metric
/// reporting `RUNNING` (state `3`); see `probes::broker_running_probe`'s doc comment.
Expand Down
Loading