craigcondit commented on code in PR #323: URL: https://github.com/apache/yunikorn-site/pull/323#discussion_r1301973455
########## docs/user_guide/preemption.md: ########## @@ -0,0 +1,252 @@ +--- +id: preemption_cases +title: Preemption +--- + +<!-- +Licensed to the Apache Software Foundation (ASF) under one +or more contributor license agreements. See the NOTICE file +distributed with this work for additional information +regarding copyright ownership. The ASF licenses this file +to you under the Apache License, Version 2.0 (the +"License"); you may not use this file except in compliance +with the License. You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, +software distributed under the License is distributed on an +"AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY +KIND, either express or implied. See the License for the +specific language governing permissions and limitations +under the License. +--> + +Preemption is an essential feature found in most schedulers, and it plays a crucial role in enabling key system functionalities like DaemonSets in K8s, as well as SLA and prioritization-based features. + +This document provides a brief introduction to the concepts and configuration methods of preemption in YuniKorn. For a more comprehensive understanding of YuniKorn's design and practical ideas related to preemption, please refer to the [design document](design/preemption.md). + +## Kubernetes Preemption + +Preemption in Kubernetes operates based on priorities. Starting from Kubernetes 1.14, you can configure preemption by adding a `preemptionPolicy` to the `PriorityClass`. However, it is important to note that preemption in Kubernetes is solely based on the priority of the pod during scheduling. The full documentation can be found [here](https://kubernetes.io/docs/concepts/scheduling-eviction/pod-priority-preemption/#preemption). + +While Kubernetes does support preemption, it does have some limitations. Preemption in Kubernetes only occurs during the scheduling cycle and does not change once the scheduling is complete. However, when considering batch or data processing workloads, it becomes necessary to account for the possibility of opting out at runtime. + +## YuniKorn Preemption + +In YuniKorn, we have introduced node-centric preemption for DaemonSet pods. With YuniKorn, we guarantee that a pod will run exclusively on a particular node, and no other pods will be scheduled on that node until the DaemonSet pod is scheduled. Review Comment: Might be worth being a little more explicit here about the fact that there are two different preemption types - generic and DaemonSet. DaemonSet preemption is much more straightforward, as it ensures that pods which must run on a particular node are allowed to do so. The remainder of the documentation here applies only to generic preemption. ########## docs/user_guide/preemption.md: ########## @@ -0,0 +1,252 @@ +--- +id: preemption_cases +title: Preemption +--- + +<!-- +Licensed to the Apache Software Foundation (ASF) under one +or more contributor license agreements. See the NOTICE file +distributed with this work for additional information +regarding copyright ownership. The ASF licenses this file +to you under the Apache License, Version 2.0 (the +"License"); you may not use this file except in compliance +with the License. You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, +software distributed under the License is distributed on an +"AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY +KIND, either express or implied. See the License for the +specific language governing permissions and limitations +under the License. +--> + +Preemption is an essential feature found in most schedulers, and it plays a crucial role in enabling key system functionalities like DaemonSets in K8s, as well as SLA and prioritization-based features. + +This document provides a brief introduction to the concepts and configuration methods of preemption in YuniKorn. For a more comprehensive understanding of YuniKorn's design and practical ideas related to preemption, please refer to the [design document](design/preemption.md). + +## Kubernetes Preemption + +Preemption in Kubernetes operates based on priorities. Starting from Kubernetes 1.14, you can configure preemption by adding a `preemptionPolicy` to the `PriorityClass`. However, it is important to note that preemption in Kubernetes is solely based on the priority of the pod during scheduling. The full documentation can be found [here](https://kubernetes.io/docs/concepts/scheduling-eviction/pod-priority-preemption/#preemption). + +While Kubernetes does support preemption, it does have some limitations. Preemption in Kubernetes only occurs during the scheduling cycle and does not change once the scheduling is complete. However, when considering batch or data processing workloads, it becomes necessary to account for the possibility of opting out at runtime. + +## YuniKorn Preemption + +In YuniKorn, we have introduced node-centric preemption for DaemonSet pods. With YuniKorn, we guarantee that a pod will run exclusively on a particular node, and no other pods will be scheduled on that node until the DaemonSet pod is scheduled. + +Additionally, YuniKorn's generic preemption is based on a hierarchical queue model, enabling pods to opt out of running. Preemption is triggered after a specified delay, ensuring that each queue's resource usage reaches at least the guaranteed amount of resources. To configure the delay time for preemption triggering, you can utilize the `preemption.delay` property in the configuration. + +To prevent the occurrence of preemption storms or loops, where subsequent preemption tasks trigger additional preemption tasks, we have designed seven preemption laws. These laws are as follows: + +1. Preemption policies are strong suggestions, not guarantees +2. Preemption can never leave a queue lower than its guaranteed capacity +3. A task cannot preempt other tasks in the same application +4. A task cannot trigger preemption unless its queue is under its guaranteed capacity +5. A task cannot be preempted unless its queue is over its guaranteed capacity +6. A task can only preempt a task with lower or equal priority +7. A task cannot preempt tasks outside its preemption fence + +For a detailed explanation of these preemption laws, please refer to the preemption [design document](design/preemption.md#the-laws-of-preemption). + +Next, we will provide a few examples to help you understand the functionality and impact of preemption, allowing you to deploy it effectively in your environment. You can find the necessary files for the examples in the yunikorn-k8shim/deployment/example/preemption directory. + +Included in the files is a YuniKorn configuration that defines the queue configuration as follows: + +```bash +queues.yaml: | + partitions: + - name: default + placementrules: + - name: provided + create: true + queues: + - name: root + submitacl: '*' + properties: + preemption.policy: fence + preemption.delay: 10s + queues: + - name: 1-normal ... + - name: 2-no-guaranteed ... + - name: 3-priority-class ... + - name: 4-priority-queue ... + - name: 5-fence ... +``` + +Each queue corresponds to a different example, and the preemption will be triggered 10 seconds after deployment, as indicated in the configuration `preemption.delay: 10s`. + +### General Preemption Case + +In this case, we will demonstrate the outcome of triggering preemption when the queue resources are distributed unevenly in a general scenario. + +We will deploy 10 pods with a resource requirement of 1 to both `queue-1` and `queue-2`. First, we deploy to `queue-1` and then introduce a few seconds delay before deploying to `queue-2`. This ensures that the resource usage in `queue-1` will exceed that of `queue-2`, depleting all resources in the parent queue and triggering preemption. + +| Queue | Max Resource | Guaranteed Resource | +| ---------------- | ------------ | ------------------- | +| `normal` | 12 | - (not configured) | +| `normal.queue-1` | 10 | 5 | +| `normal.queue-2` | 10 | 5 | + +Result: + +When a set of guaranteed resources is defined, preemption aims to ensure that all queues satisfy their guaranteed resources. Preemption stops once the guaranteed resources are met (law 4). A queue may be preempted if it has more resources than its guaranteed amount. For instance, in this case, if queue-1 has fewer resources than its guaranteed amount (<5), it will not be preempted (law 5). + +| Queue | Resource before preemption | Resource after preemption | +| ---------------- | -------------------------- | ------------------------- | +| `normal.queue-1` | 10 (victim) | 7 | +| `normal.queue-2` | 2 | 5 (guaranteed minimum) | + + + +### Preemption Case Without Guaranteed Resources Review Comment: This isn't how preemption works. There's no concept of "equal sharing" mediated by preemption, as queues can only be preempted if they are both 1) over their guaranteed quota and 2) would remain at or above that guarantee if preempted from). Additionally, a pending pod in a queue is not allowed to trigger preemption of anything else unless its queue is both 1) under guaranteed quota and 2) would be at or below guaranteed quota if the pod was launched. Clearly, these cases conflict with what is described here, as you must have guaranteed resources so that preemption can happen. ########## docs/user_guide/preemption.md: ########## @@ -0,0 +1,252 @@ +--- +id: preemption_cases +title: Preemption +--- + +<!-- +Licensed to the Apache Software Foundation (ASF) under one +or more contributor license agreements. See the NOTICE file +distributed with this work for additional information +regarding copyright ownership. The ASF licenses this file +to you under the Apache License, Version 2.0 (the +"License"); you may not use this file except in compliance +with the License. You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, +software distributed under the License is distributed on an +"AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY +KIND, either express or implied. See the License for the +specific language governing permissions and limitations +under the License. +--> + +Preemption is an essential feature found in most schedulers, and it plays a crucial role in enabling key system functionalities like DaemonSets in K8s, as well as SLA and prioritization-based features. + +This document provides a brief introduction to the concepts and configuration methods of preemption in YuniKorn. For a more comprehensive understanding of YuniKorn's design and practical ideas related to preemption, please refer to the [design document](design/preemption.md). + +## Kubernetes Preemption + +Preemption in Kubernetes operates based on priorities. Starting from Kubernetes 1.14, you can configure preemption by adding a `preemptionPolicy` to the `PriorityClass`. However, it is important to note that preemption in Kubernetes is solely based on the priority of the pod during scheduling. The full documentation can be found [here](https://kubernetes.io/docs/concepts/scheduling-eviction/pod-priority-preemption/#preemption). + +While Kubernetes does support preemption, it does have some limitations. Preemption in Kubernetes only occurs during the scheduling cycle and does not change once the scheduling is complete. However, when considering batch or data processing workloads, it becomes necessary to account for the possibility of opting out at runtime. + +## YuniKorn Preemption + +In YuniKorn, we have introduced node-centric preemption for DaemonSet pods. With YuniKorn, we guarantee that a pod will run exclusively on a particular node, and no other pods will be scheduled on that node until the DaemonSet pod is scheduled. + +Additionally, YuniKorn's generic preemption is based on a hierarchical queue model, enabling pods to opt out of running. Preemption is triggered after a specified delay, ensuring that each queue's resource usage reaches at least the guaranteed amount of resources. To configure the delay time for preemption triggering, you can utilize the `preemption.delay` property in the configuration. + +To prevent the occurrence of preemption storms or loops, where subsequent preemption tasks trigger additional preemption tasks, we have designed seven preemption laws. These laws are as follows: + +1. Preemption policies are strong suggestions, not guarantees +2. Preemption can never leave a queue lower than its guaranteed capacity +3. A task cannot preempt other tasks in the same application +4. A task cannot trigger preemption unless its queue is under its guaranteed capacity +5. A task cannot be preempted unless its queue is over its guaranteed capacity +6. A task can only preempt a task with lower or equal priority +7. A task cannot preempt tasks outside its preemption fence + +For a detailed explanation of these preemption laws, please refer to the preemption [design document](design/preemption.md#the-laws-of-preemption). + +Next, we will provide a few examples to help you understand the functionality and impact of preemption, allowing you to deploy it effectively in your environment. You can find the necessary files for the examples in the yunikorn-k8shim/deployment/example/preemption directory. + +Included in the files is a YuniKorn configuration that defines the queue configuration as follows: + +```bash +queues.yaml: | + partitions: + - name: default + placementrules: + - name: provided + create: true + queues: + - name: root + submitacl: '*' + properties: + preemption.policy: fence + preemption.delay: 10s + queues: + - name: 1-normal ... + - name: 2-no-guaranteed ... + - name: 3-priority-class ... + - name: 4-priority-queue ... + - name: 5-fence ... +``` + +Each queue corresponds to a different example, and the preemption will be triggered 10 seconds after deployment, as indicated in the configuration `preemption.delay: 10s`. + +### General Preemption Case + +In this case, we will demonstrate the outcome of triggering preemption when the queue resources are distributed unevenly in a general scenario. + +We will deploy 10 pods with a resource requirement of 1 to both `queue-1` and `queue-2`. First, we deploy to `queue-1` and then introduce a few seconds delay before deploying to `queue-2`. This ensures that the resource usage in `queue-1` will exceed that of `queue-2`, depleting all resources in the parent queue and triggering preemption. + +| Queue | Max Resource | Guaranteed Resource | +| ---------------- | ------------ | ------------------- | +| `normal` | 12 | - (not configured) | +| `normal.queue-1` | 10 | 5 | +| `normal.queue-2` | 10 | 5 | + +Result: + +When a set of guaranteed resources is defined, preemption aims to ensure that all queues satisfy their guaranteed resources. Preemption stops once the guaranteed resources are met (law 4). A queue may be preempted if it has more resources than its guaranteed amount. For instance, in this case, if queue-1 has fewer resources than its guaranteed amount (<5), it will not be preempted (law 5). + +| Queue | Resource before preemption | Resource after preemption | +| ---------------- | -------------------------- | ------------------------- | +| `normal.queue-1` | 10 (victim) | 7 | +| `normal.queue-2` | 2 | 5 (guaranteed minimum) | + + + +### Preemption Case Without Guaranteed Resources + +Similar to the previous example, but this time the queue does not have a "guaranteed resource" defined. + +| Queue | Max Resource | Guaranteed Resource | +| -------------- | ------------ | ------------------- | +| `root` | 12 | - | +| `root.queue-1` | 10 | - | +| `root.queue-2` | 10 | - | + +Result: + +In the absence of a guaranteed resource setting, preemption allows each queue to utilize an equal amount of resources. + +| Queue | Resource before preemption | Resource after preemption | +| -------------- | -------------------------- | ------------------------- | +| `root.queue-1` | 10 (victim) | 6 | +| `root.queue-2` | 2 | 6 | + + + +### Priority + +In general, a pod can preempt a pod with equal or lower priority. You can set the priority by defining a [PriorityClass](https://kubernetes.io/docs/concepts/scheduling-eviction/pod-priority-preemption/) or by utilizing [queue priorities](priorities). + +While preemption allows service-type pods to scale up or down through preemption, it can also lead to the preemption of pods that should not be preempted in certain scenarios: + +1. Spark Jobs, where the driver pod manages a large number of jobs, and if preempted, all jobs will be affected. +2. Interactive pods, such as Python notebooks, have a significant impact when restarted and should be avoided from preemption. + +To address this issue, we have designed a "do not preempt me" flag. You can set the annotation `yunikorn.apache.org/allow-preemption` to `false` in the PriorityClass to prevent pod requests from being preempted. +> **_NOTE:_** The flag `yunikorn.apache.org/allow-preemption` is a request only. It is not guaranteed but Pods annotated with this flag will be preempted last. + + +### PriorityClass + +In this example, we will demonstrate the configuration of `yunikorn.apache.org/allow-preemption` using PriorityClass and observe its effect. The default value for this configuration is set to `true`. + +```bash +apiVersion: scheduling.k8s.io/v1 +kind: PriorityClass +metadata: + name: preemption-not-allow + annotations: + "yunikorn.apache.org/allow-preemption": "false" +value: 0 +``` + +We will deploy 8 pods with a resource requirement of 1 to `queue-1`, `queue-2`, and `queue-3`, respectively. We will deploy to `queue-1` and `queue-2` first, followed by a few seconds delay before deploying to `queue-3`. This ensures that the resource usage in `queue-1` and `queue-2` will be greater than that in `queue-3`, depleting all resources in the parent queue and triggering preemption. + +| Queue | Max Resource | Guaranteed Resource | `allow-preemption` | +| ------------ | ------------ | ------------------- | ------------------ | +| `rt` | 16 | - | | +| `rt.queue-1` | 8 | 3 | `true` | +| `rt.queue-2` | 8 | 3 | `false` | +| `rt.queue-3` | 8 | 3 | `true` | + +Result: + +When preemption is triggered, `queue-3` will start searching for a victim. However, since `queue-2` is set with allow-preemption as false, the resources of `queue-1` will be preempted. + +Please note that setting `yunikorn.apache.org/allow-preemption` is a strong recommendation but does not guarantee the lack of preemption. When this flag is set to `false`, it moves the Pod to the back of the preemption list, giving it a lower priority for preemption compared to other Pods. However, in certain scenarios, such as when no other preemption options are available, Pods with this flag may still be preempted. + +For example, even with `allow-preemption` set to `false`, DaemonSet pods can still be preempted. Additionally, if an application in `queue-1` has a higher priority than one in `queue-3`, the application in `queue-2` will be preempted because an application can never preempt another application with a higher priority. In such cases where no other preemption options exist, the `allow-preemption` flag may not prevent preemption. Review Comment: Instead of "DaemonSet pods can still be preempted" this should probably be "DaemonSet pods can still trigger preemption". The logic as given is inverted from how it works, as DaemonSet pods are never preempted. ########## docs/user_guide/preemption.md: ########## @@ -0,0 +1,252 @@ +--- +id: preemption_cases +title: Preemption +--- + +<!-- +Licensed to the Apache Software Foundation (ASF) under one +or more contributor license agreements. See the NOTICE file +distributed with this work for additional information +regarding copyright ownership. The ASF licenses this file +to you under the Apache License, Version 2.0 (the +"License"); you may not use this file except in compliance +with the License. You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, +software distributed under the License is distributed on an +"AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY +KIND, either express or implied. See the License for the +specific language governing permissions and limitations +under the License. +--> + +Preemption is an essential feature found in most schedulers, and it plays a crucial role in enabling key system functionalities like DaemonSets in K8s, as well as SLA and prioritization-based features. + +This document provides a brief introduction to the concepts and configuration methods of preemption in YuniKorn. For a more comprehensive understanding of YuniKorn's design and practical ideas related to preemption, please refer to the [design document](design/preemption.md). + +## Kubernetes Preemption + +Preemption in Kubernetes operates based on priorities. Starting from Kubernetes 1.14, you can configure preemption by adding a `preemptionPolicy` to the `PriorityClass`. However, it is important to note that preemption in Kubernetes is solely based on the priority of the pod during scheduling. The full documentation can be found [here](https://kubernetes.io/docs/concepts/scheduling-eviction/pod-priority-preemption/#preemption). + +While Kubernetes does support preemption, it does have some limitations. Preemption in Kubernetes only occurs during the scheduling cycle and does not change once the scheduling is complete. However, when considering batch or data processing workloads, it becomes necessary to account for the possibility of opting out at runtime. + +## YuniKorn Preemption + +In YuniKorn, we have introduced node-centric preemption for DaemonSet pods. With YuniKorn, we guarantee that a pod will run exclusively on a particular node, and no other pods will be scheduled on that node until the DaemonSet pod is scheduled. + +Additionally, YuniKorn's generic preemption is based on a hierarchical queue model, enabling pods to opt out of running. Preemption is triggered after a specified delay, ensuring that each queue's resource usage reaches at least the guaranteed amount of resources. To configure the delay time for preemption triggering, you can utilize the `preemption.delay` property in the configuration. + +To prevent the occurrence of preemption storms or loops, where subsequent preemption tasks trigger additional preemption tasks, we have designed seven preemption laws. These laws are as follows: + +1. Preemption policies are strong suggestions, not guarantees +2. Preemption can never leave a queue lower than its guaranteed capacity +3. A task cannot preempt other tasks in the same application +4. A task cannot trigger preemption unless its queue is under its guaranteed capacity +5. A task cannot be preempted unless its queue is over its guaranteed capacity +6. A task can only preempt a task with lower or equal priority +7. A task cannot preempt tasks outside its preemption fence + +For a detailed explanation of these preemption laws, please refer to the preemption [design document](design/preemption.md#the-laws-of-preemption). + +Next, we will provide a few examples to help you understand the functionality and impact of preemption, allowing you to deploy it effectively in your environment. You can find the necessary files for the examples in the yunikorn-k8shim/deployment/example/preemption directory. + +Included in the files is a YuniKorn configuration that defines the queue configuration as follows: + +```bash +queues.yaml: | + partitions: + - name: default + placementrules: + - name: provided + create: true + queues: + - name: root + submitacl: '*' + properties: + preemption.policy: fence + preemption.delay: 10s + queues: + - name: 1-normal ... + - name: 2-no-guaranteed ... + - name: 3-priority-class ... + - name: 4-priority-queue ... + - name: 5-fence ... +``` + +Each queue corresponds to a different example, and the preemption will be triggered 10 seconds after deployment, as indicated in the configuration `preemption.delay: 10s`. + +### General Preemption Case + +In this case, we will demonstrate the outcome of triggering preemption when the queue resources are distributed unevenly in a general scenario. + +We will deploy 10 pods with a resource requirement of 1 to both `queue-1` and `queue-2`. First, we deploy to `queue-1` and then introduce a few seconds delay before deploying to `queue-2`. This ensures that the resource usage in `queue-1` will exceed that of `queue-2`, depleting all resources in the parent queue and triggering preemption. + +| Queue | Max Resource | Guaranteed Resource | +| ---------------- | ------------ | ------------------- | +| `normal` | 12 | - (not configured) | +| `normal.queue-1` | 10 | 5 | +| `normal.queue-2` | 10 | 5 | + +Result: + +When a set of guaranteed resources is defined, preemption aims to ensure that all queues satisfy their guaranteed resources. Preemption stops once the guaranteed resources are met (law 4). A queue may be preempted if it has more resources than its guaranteed amount. For instance, in this case, if queue-1 has fewer resources than its guaranteed amount (<5), it will not be preempted (law 5). + +| Queue | Resource before preemption | Resource after preemption | +| ---------------- | -------------------------- | ------------------------- | +| `normal.queue-1` | 10 (victim) | 7 | +| `normal.queue-2` | 2 | 5 (guaranteed minimum) | + + + +### Preemption Case Without Guaranteed Resources Review Comment: This entire case should probably be removed. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
