diff --git a/docs/modules/kafka/pages/usage-guide/kraft-controller.adoc b/docs/modules/kafka/pages/usage-guide/kraft-controller.adoc index 1d176745..0bb755ce 100644 --- a/docs/modules/kafka/pages/usage-guide/kraft-controller.adoc +++ b/docs/modules/kafka/pages/usage-guide/kraft-controller.adoc @@ -1,8 +1,6 @@ -= KRaft mode (experimental) += KRaft mode :description: Apache Kafka KRaft mode with the Stackable Operator for Apache Kafka -WARNING: The Kafka KRaft mode is currently experimental, and subject to change. - Apache Kafka's KRaft mode replaces Apache ZooKeeper with Kafka’s own built-in consensus mechanism based on the Raft protocol. This simplifies Kafka’s architecture, reducing operational complexity by consolidating cluster metadata management into Kafka itself. @@ -18,8 +16,7 @@ WARNING: The Stackable Operator for Apache Kafka currently does not support auto == Configuration -The Stackable Kafka operator introduces a new xref:concepts:roles-and-role-groups.adoc[role] in the KafkaCluster CRD called KRaft `Controller`. -Configuring the `Controller` will put Kafka into KRaft mode. Apache ZooKeeper will not be required anymore. +Enable KRaft mode by adding `spec.clusterConfig.metadataManager: KRaft` and a `Controller` xref:concepts:roles-and-role-groups.adoc[role] to your cluster manifest as in the example below: [source,yaml] ---- @@ -29,7 +26,7 @@ metadata: name: kafka spec: clusterConfig: - metadataManager: kraft + metadataManager: KRaft image: productVersion: "3.9.2" brokers: @@ -42,22 +39,39 @@ spec: replicas: 3 ---- -NOTE: Using `spec.controllers` is mutually exclusive with `spec.clusterConfig.zookeeperConfigMapName`. +NOTE: Using `spec.controllers` is mutually exclusive with `spec.clusterConfig.zookeeperConfigMapName`. Configuring both would result in an error. -=== Recommendations +=== Resources -A minimal KRaft setup consisting of at least 3 Controllers has the following https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/[resource requirements]: +Each Pod's https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/[resource requirements] +are the sum of its containers. Controller Pods run an additional `quorum-manager` sidecar +(see <<_internal_operator_details, Internal operator details>>) whose resources are fixed and *not* configurable through +`spec.controllers.config.resources`. Only the `kafka` container is eligible. -* `600m` CPU request -* `3000m` CPU limit -* `3000Mi` memory request and limit -* `6Gi` persistent storage +[cols="2,3,2,2,2,2", options="header"] +|=== +| Role | Container | CPU request | CPU limit | Memory (request = limit) | Storage (`logDirs`) -NOTE: The Controller replicas should sum up to an odd number for the Raft consensus. +| Controller | `kafka` (configurable) | 250m | 1000m | 1Gi | 2Gi +| Controller | `quorum-manager` (fixed) | 100m | 500m | 512Mi | -- +| Controller | *Pod total* | *350m* | *1500m* | *1.5Gi* | *2Gi* +| Broker | `kafka` (configurable) | 250m | 1000m | 2Gi | 2Gi +|=== -=== Resources +A minimal KRaft setup consists of 3 Controllers and 3 Brokers: + +[cols="2,2,2,2,2", options="header"] +|=== +| | CPU request | CPU limit | Memory | Storage -Corresponding to the values above, the operator uses the following resource defaults: +| 3 Controllers | 1050m | 4500m | 4.5Gi | 6Gi +| 3 Brokers | 750m | 3000m | 6Gi | 6Gi +| *Total* | *1800m* | *7500m* | *10.5Gi* | *12Gi* +|=== + +NOTE: The Controller replicas should sum up to an odd number for the Raft consensus. + +Corresponding to the `kafka` container values above, the operator uses the following resource defaults: [source,yaml] ---- @@ -72,15 +86,22 @@ controllers: storage: logDirs: capacity: 2Gi +brokers: + config: + resources: + memory: + limit: 2Gi + cpu: + min: 250m + max: 1000m + storage: + logDirs: + capacity: 2Gi ---- -=== Overrides - -The configuration of overrides, JVM arguments etc. is similar to the Broker and documented on the xref:concepts:overrides.adoc[concepts page]. - == Internal operator details -KRaft mode requires major configuration changes compared to ZooKeeper: +KRaft mode requires major configuration changes compared to ZooKeeper mode: * `cluster-id`: This is set to the `metadata.name` of the KafkaCluster resource during initial formatting * `node.id`: This is a calculated integer, hashed from the `role` and `rolegroup` and added `replica` id. @@ -93,7 +114,7 @@ KRaft mode requires major configuration changes compared to ZooKeeper: Every other controller formats with `--no-initial-controllers` and joins purely through the sidecar's `add-controller` call. Brokers always format with `--no-initial-controllers` too; they are never voters. * `controller.quorum.bootstrap.servers` points at each controller role group's own headless Service DNS name, not - individual pod addresses -- unless Kerberos is enabled. See below for details related to Kerberos. + individual pod addresses, unless Kerberos is enabled. See below for details related to Kerberos. == Kerberos @@ -146,8 +167,8 @@ The example cluster will be kept minimal without any additional configuration. We'll use Kafka version `3.9.2` for this purpose. This is because this is the last version from the 3.x Kafka series that runs on ZooKeeper mode and is supported by the SDP. -We'll also assign broker ids manually from the beginning to simplify this guide. In a real-workd scenario, you do not have this option at this step because your cluster is already running. -In a real world-scenario you'll have to collect these ids and configure manual assignment at the second step of the migration. +We'll also assign broker ids manually from the beginning to simplify this guide. In a real-world scenario, you do not have this option at cluster creation because your cluster is already running. +Instead, you'll have to collect these ids from the running cluster and configure manual assignment as part of step 1 below. We start by creating a dedicated namespace to work in and deploy the Kafka cluster including ZooKeeper and credentials. diff --git a/rust/operator-binary/src/controller/build/resource/statefulset.rs b/rust/operator-binary/src/controller/build/resource/statefulset.rs index 8dc01278..eb92b9c6 100644 --- a/rust/operator-binary/src/controller/build/resource/statefulset.rs +++ b/rust/operator-binary/src/controller/build/resource/statefulset.rs @@ -1241,6 +1241,172 @@ mod tests { .expect("the kafka container is built") } + /// `spec.controllers.config.resources` and `spec.brokers.config.resources` must be applied + /// verbatim to each role's `kafka` container: CPU request/limit come from `cpu.min`/`cpu.max` + /// and memory request equals memory limit, both taken from `memory.limit` (there is no + /// separate memory request field in the CRD). + #[test] + fn kafka_container_resources_match_configured_role_resources() { + let kafka = minimal_kafka( + r#" + apiVersion: kafka.stackable.tech/v1alpha1 + kind: KafkaCluster + metadata: + name: simple-kafka + namespace: default + uid: 12345678-1234-1234-1234-123456789012 + spec: + image: + productVersion: 3.9.2 + clusterConfig: + metadataManager: kraft + controllers: + config: + resources: + memory: + limit: 2Gi + cpu: + min: 260m + max: 1024m + storage: + logDirs: + capacity: 2Gi + roleGroups: + default: + replicas: 3 + brokers: + config: + resources: + memory: + limit: 2Gi + cpu: + min: 300m + max: 500m + storage: + logDirs: + capacity: 2Gi + roleGroups: + default: + replicas: 3 + "#, + ); + let cluster = validated_cluster(&kafka); + + let controller_resources = controller_kafka_container(&cluster) + .resources + .expect("the controller kafka container has resources"); + assert_cpu_and_memory(&controller_resources, "260m", "1024m", "2Gi", "2Gi"); + + let broker_resources = broker_kafka_container(&cluster) + .resources + .expect("the broker kafka container has resources"); + assert_cpu_and_memory(&broker_resources, "300m", "500m", "2Gi", "2Gi"); + } + + /// Asserts the container's CPU request/limit and memory request/limit against the given + /// quantity strings (memory request and limit are always equal, see the CRD's + /// `Resources`/`MemoryLimits` type). + fn assert_cpu_and_memory( + resources: &stackable_operator::k8s_openapi::api::core::v1::ResourceRequirements, + expected_cpu_request: &str, + expected_cpu_limit: &str, + expected_memory_request: &str, + expected_memory_limit: &str, + ) { + use stackable_operator::k8s_openapi::apimachinery::pkg::api::resource::Quantity; + + let requests = resources + .requests + .as_ref() + .expect("the container has resource requests"); + let limits = resources + .limits + .as_ref() + .expect("the container has resource limits"); + + assert_eq!( + requests.get("cpu"), + Some(&Quantity(expected_cpu_request.to_string())), + "cpu request" + ); + assert_eq!( + limits.get("cpu"), + Some(&Quantity(expected_cpu_limit.to_string())), + "cpu limit" + ); + assert_eq!( + requests.get("memory"), + Some(&Quantity(expected_memory_request.to_string())), + "memory request" + ); + assert_eq!( + limits.get("memory"), + Some(&Quantity(expected_memory_limit.to_string())), + "memory limit" + ); + } + + /// Unlike the `kafka` container, the `quorum-manager` sidecar's resources are hardcoded in + /// [`build_quorum_manager_container`] and not exposed through `config.resources` (which only + /// applies to the `kafka` container, see + /// [`kafka_container_resources_match_configured_role_resources`]). `podOverrides` is the only + /// way to change them, applied via a strategic merge on `spec.containers` keyed by container + /// `name` (see the `merge_from` call at the end of `build_controller_rolegroup_statefulset`). + #[test] + fn quorum_manager_resources_can_be_changed_by_pod_overrides() { + let kafka = minimal_kafka( + r#" + apiVersion: kafka.stackable.tech/v1alpha1 + kind: KafkaCluster + metadata: + name: simple-kafka + namespace: default + uid: 12345678-1234-1234-1234-123456789012 + spec: + image: + productVersion: 3.9.2 + clusterConfig: + metadataManager: kraft + controllers: + podOverrides: + spec: + containers: + - name: quorum-manager + resources: + requests: + cpu: 200m + memory: 256Mi + limits: + cpu: 750m + memory: 768Mi + roleGroups: + default: + replicas: 3 + brokers: + roleGroups: + default: + replicas: 3 + "#, + ); + let cluster = validated_cluster(&kafka); + + let containers = controller_containers(&cluster); + let sidecar = containers + .iter() + .find(|c| c.name == QUORUM_MANAGER_CONTAINER_NAME.to_string()) + .expect("the quorum-manager sidecar is built"); + let resources = sidecar + .resources + .as_ref() + .expect("the override gives the sidecar resources"); + + // Request and limit are asserted independently (unlike + // `kafka_container_resources_match_configured_role_resources`'s memory check) to show + // that `podOverrides` sets the raw Kubernetes fields directly, without the + // request-equals-limit constraint the CRD's `config.resources` memory field imposes. + assert_cpu_and_memory(resources, "200m", "750m", "256Mi", "768Mi"); + } + /// The startup and liveness probes share the same check - TCP reachability (a genuinely /// dead/hung process must still be restarted) plus the broker's JMX `BrokerState` metric /// reporting `RUNNING` (state `3`); see `probes::broker_running_probe`'s doc comment.