tigerquoll commented on code in PR #1060: URL: https://github.com/apache/yunikorn-k8shim/pull/1060#discussion_r3889399190
########## pkg/client/eventsink.go: ########## Review Comment: Agreed that not all events are the same; the bucket approach was type blind - moved to a new prioritisation system: - **Warning events are never shed by the bucket.** `FailedScheduling`, `TaskRejected`, `NodeRejected`, `ApplicationFailed`, the gang scheduling failures — everything that explains why a pod is not running — always goes out and does not consume a token. - **`PodUnschedulable` becomes a Warning.** It was typed Normal, which would have put it in the sheddable class; kube-scheduler's equivalent, `FailedScheduling`, is a Warning too. - **Only Normal events are shed above `eventQPS`.** Counting what the shim emits, the volume is three Normal events per *successfully* scheduled pod (`Scheduling`, `Scheduled`, `PodBindSuccessful`) plus the core's `Informational` records. At the ~1,900 binds/s measured on the KWOK rig that is ~5,700 events/s saying "this worked", and that is what a storm sheds. - **`kubernetes.eventLevel`** (`normal` | `warning` | `none`, default `normal`) is the verbosity knob you suggested. It is applied at the recorder, so a suppressed event is never copied, cached or written, and it is hot-reloadable like the log level. The server-side mute still applies to every type. Once APF is rejecting events, client-go drops them itself after its Retry-After retries (`recordEvent` treats every `StatusError` as "Server rejected event (will not retry!)"), so muting only saves the round trips and the parked goroutines; it does not lose an event that would otherwise have landed. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
