From 635467e7e5cfe94a87eca09c26d2540736169d6a Mon Sep 17 00:00:00 2001 From: Shane Lovern Date: Thu, 27 Aug 2026 16:56:37 +0100 Subject: [PATCH] TELCODOCS-2918 - Enable default-preemption for numaresources-operator scheduler --- ...nf-enabling-numa-scheduler-preemption.adoc | 67 +++++++++++++++++++ ...-numa-scheduler-preemption-parameters.adoc | 34 ++++++++++ modules/cnf-numa-scheduler-preemption.adoc | 34 ++++++++++ .../cnf-numa-aware-scheduling.adoc | 11 +++ 4 files changed, 146 insertions(+) create mode 100644 modules/cnf-enabling-numa-scheduler-preemption.adoc create mode 100644 modules/cnf-numa-scheduler-preemption-parameters.adoc create mode 100644 modules/cnf-numa-scheduler-preemption.adoc diff --git a/modules/cnf-enabling-numa-scheduler-preemption.adoc b/modules/cnf-enabling-numa-scheduler-preemption.adoc new file mode 100644 index 00000000000..5ca8974c38f --- /dev/null +++ b/modules/cnf-enabling-numa-scheduler-preemption.adoc @@ -0,0 +1,67 @@ +// Module included in the following assemblies: +// +// *scalability_and_performance/cnf-numa-aware-scheduling.adoc + +:_mod-docs-content-type: PROCEDURE +[id="cnf-enabling-numa-scheduler-preemption_{context}"] += Enabling preemption for the NUMA-aware scheduler + +[role="_abstract"] +To enable priority-based preemption for the NUMA Resources Operator scheduler, configure the `NUMAResourcesScheduler` custom resource. Preemption is disabled by default. When enabled, the scheduler can preempt lower-priority pods to make room for higher-priority pods on constrained cluster nodes. + +.Prerequisites + +* The NUMA Resources Operator is installed. +* You have access to the cluster as a user with the `cluster-admin` role. + +.Procedure + +. Edit the `NUMAResourcesScheduler` custom resource by running the following command: ++ +[source,terminal] +---- +$ oc edit numaresourcesscheduler numaresourcesscheduler +---- + +. Add the `preemptionMode` field to enable preemption: ++ +[source,yaml] +---- +apiVersion: nodetopology.openshift.io/v1 +kind: NUMAResourcesScheduler +metadata: + name: numaresourcesscheduler +spec: + preemptionMode: Enabled +# ... +---- + +. Save and exit the editor. ++ +The NUMA Resources Operator automatically restarts the scheduler pod to apply the new configuration. + +.Verification + +. Verify that the scheduler pod has restarted with the new configuration by running the following command: ++ +[source,terminal] +---- +$ oc get pods -n openshift-numaresources -l app=secondary-scheduler +---- ++ +The output shows a new pod with a recent `AGE` value. + +. Verify that the `NUMAResourcesScheduler` custom resource shows the updated configuration: ++ +[source,terminal] +---- +$ oc get numaresourcesscheduler numaresourcesscheduler -o jsonpath='{.spec.preemptionMode}' +---- ++ +.Example output +[source,terminal] +---- +Enabled +---- + +. Optional: To test that preemption is working, create workloads with different `PriorityClass` values and verify that higher-priority pods preempt lower-priority pods when resources are constrained. diff --git a/modules/cnf-numa-scheduler-preemption-parameters.adoc b/modules/cnf-numa-scheduler-preemption-parameters.adoc new file mode 100644 index 00000000000..20c91edcf44 --- /dev/null +++ b/modules/cnf-numa-scheduler-preemption-parameters.adoc @@ -0,0 +1,34 @@ +// Module included in the following assemblies: +// +// *scalability_and_performance/cnf-numa-aware-scheduling.adoc + +:_mod-docs-content-type: REFERENCE +[id="cnf-numa-scheduler-preemption-parameters_{context}"] += NUMA Resources Operator scheduler preemption configuration + +[role="_abstract"] +You can configure preemption for the NUMA Resources Operator scheduler by using the `preemptionMode` field in the `NUMAResourcesScheduler` custom resource. Preemption is disabled by default. This field controls whether the scheduler uses the `DefaultPreemption` plugin to preempt lower-priority pods. + +The following table describes the `NUMAResourcesScheduler` preemption configuration field: + +[cols="2,1,3", options="header"] +|=== +|Field |Type |Description + +|`spec.preemptionMode` +|`string` +|Enables preemption functionality for the NUMA-aware scheduler. The scheduler includes the `DefaultPreemption` plugin, but preemption is inactive until this field is set. When set to `Enabled`, the scheduler can preempt lower-priority pods to schedule higher-priority pods. When not set or set to `Disabled`, preemption is disabled. Accepted values: `Enabled`, `Disabled`. +|=== + +The following example shows a `NUMAResourcesScheduler` custom resource with preemption enabled: + +[source,yaml] +---- +apiVersion: nodetopology.openshift.io/v1 +kind: NUMAResourcesScheduler +metadata: + name: numaresourcesscheduler +spec: + imageSpec: "registry.redhat.io/openshift5/noderesourcetopology-scheduler-rhel9:v{product-version}" + preemptionMode: Enabled +---- diff --git a/modules/cnf-numa-scheduler-preemption.adoc b/modules/cnf-numa-scheduler-preemption.adoc new file mode 100644 index 00000000000..6b1d9bff09b --- /dev/null +++ b/modules/cnf-numa-scheduler-preemption.adoc @@ -0,0 +1,34 @@ +// Module included in the following assemblies: +// +// *scalability_and_performance/cnf-numa-aware-scheduling.adoc + +:_mod-docs-content-type: CONCEPT +[id="cnf-numa-scheduler-preemption_{context}"] += Scheduling critical workloads with preemption + +[role="_abstract"] +You can configure the NUMA Resources Operator scheduler to use the default Kubernetes scheduler preemption feature. Priority-based preemption enables the scheduler to preempt lower-priority pods so that higher-priority pods can be scheduled when cluster resources are constrained. This helps ensure critical workloads get the resources they need while making better use of available capacity. + +When you enable preemption for the NUMA Resources Operator scheduler, the scheduler uses the standard `DefaultPreemption` plugin from Kubernetes, not a separate NUMA-aware preemption mechanism. + +Preemption works with the standard Kubernetes `PriorityClass` resources. When a high-priority pod cannot be scheduled because of insufficient resources, the scheduler identifies lower-priority pods that can be preempted to make room for the pending pod. The scheduler then preempts those pods and schedules the high-priority pod in their place. + +Benefits of enabling preemption include: + +* Critical workloads are guaranteed resources even on heavily utilized clusters +* Cluster resources are used more efficiently by allowing important pods to displace less important ones +* Service degradation is reduced by ensuring that infrastructure and high-priority application pods are always prioritized + +[NOTE] +==== +NRT cache timing considerations: + +* *Scheduling queue behavior:* When a pod cannot be scheduled, it enters the unschedulable queue and waits for either node resource events or the unschedulable timeout period (5 minutes by default). If the NRT plugin rejects a pod because of stale cache information, the pod remains in this queue. Even after the NRT cache refreshes, the pod does not immediately retry scheduling — it continues to wait for node events or the timeout to trigger a new scheduling attempt. + +* *False eviction during preemption:* When a pod fails the NRT filter plugin, the scheduler initiates default preemption and evaluates sets of pods for eviction. If the NRT cache updates during this evaluation to show sufficient free resources, the preemption plugin might evict pods unnecessarily. To mitigate this, ensure all pods are managed by controllers such as ReplicaSets or DaemonSets. This allows evicted pods to respawn automatically, although scheduling order is not guaranteed. +==== + +[NOTE] +==== +Preemption handles priority-based scheduling decisions. Quality of Service (QoS) based eviction, which evicts pods based on resource pressure on a node, is handled separately by the kubelet's node-pressure eviction feature and is not affected by this configuration. +==== diff --git a/scalability_and_performance/cnf-numa-aware-scheduling.adoc b/scalability_and_performance/cnf-numa-aware-scheduling.adoc index 607e87e37b8..48752c242ea 100644 --- a/scalability_and_performance/cnf-numa-aware-scheduling.adoc +++ b/scalability_and_performance/cnf-numa-aware-scheduling.adoc @@ -77,6 +77,17 @@ include::modules/cnf-deploying-the-numa-aware-scheduler.adoc[leveloffset=+2] * xref:../disconnected/updating/disconnected-update.adoc#images-configuration-registry-mirror_updating-disconnected-cluster[Configuring image registry repository mirroring] +include::modules/cnf-numa-scheduler-preemption.adoc[leveloffset=+2] + +[role="_additional-resources"] +.Additional resources + +* xref:../nodes/pods/nodes-pods-priority.adoc#nodes-pods-priority-preempt-about_nodes-pods-priority[Understanding pod priority] + +include::modules/cnf-enabling-numa-scheduler-preemption.adoc[leveloffset=+3] + +include::modules/cnf-numa-scheduler-preemption-parameters.adoc[leveloffset=+3] + include::modules/cnf-scheduling-numa-aware-workloads.adoc[leveloffset=+2] include::modules/cnf-nrop-support-schedulable-resources.adoc[leveloffset=+1]