Hi community, Reviving https://issues.apache.org/jira/browse/YUNIKORN-829 . We are running Spark on YuniKorn, and have a requirement to provide more observability of *actual* resource usage for our customers, data engineers/scientists who wrote Spark jobs who may not have deep expertise in Spark job optimization.
- requirement: - have actual resource usage metrics at both job level and queue level (YK already have requested resource usage metrics) - key use case: - as indicators of job optimization for ICs like data engineers/scientists, to show users how much resources they requested v.s. how much resources their jobs actually used - as indicator for managers on their team's resource utilization. In our setup or a typical YK setup, each customer team has their own YuniKorn queue in a shared, multi tenant environment. Managers of the team would want high level (queue) metrics rather than low level (job) ones Currently we haven't found a good product on the market to do this, so would be great if YuniKorn can support it. Would like your input here on feasibility (seems feasible according Weiwei's comment in Jira), priority, and timeline/complexity of the projects. Thanks, Bowen
