zhanqian-zhang commented on issue #49983:
URL: https://github.com/apache/airflow/issues/49983#issuecomment-6008118617

   I'd also like to pick this up and focus on the memory side of task resource 
metrics. For anyone joining the issue, a quick recap:
   
   Airflow 2.10 added periodic task CPU and memory metrics in #39650. Its 
parent task runner periodically observed the task subprocess and reported 
memory as a percentage. Those metrics are no longer emitted in Airflow 3, as 
reported here.
   
   #56690 later proposed measuring the Python Task Runner's RSS in MB, taking a 
snapshot after task execution rather than sampling it periodically. That PR 
wasn't merged. Its review also raised a question that matters here: should a 
memory reading identify a particular execution, or is grouping by 
dag_id/task_id sufficient?
   
   My proposed V1 would borrow the periodic sampling approach from 2.10, but 
measure only the Python Task Runner process's RSS, in bytes, under an 
explicitly scoped name such as task.runner_rss_bytes.
   
   I think that could be a useful diagnostic: when high runner RSS is observed, 
it gives us a clue about which tasks to investigate. But it would not include 
subprocess or container/Pod memory, show the total across concurrent runs, or 
reliably capture an execution's peak. A short spike could be missed entirely.
   
   There is also a question to settle before implementing it. If two runs of 
the same task report at once, what should a dag_id/task_id-level reading mean? 
Their Gauges do not automatically produce a total or maximum, and we need to 
check whether the supported metrics backends keep the observations and task 
identities interpretable. Calling the metric “best effort” does not resolve 
that question.
   
   I've outlined two broader scopes in this Discussion: V2 would aggregate 
active runner readings through a stable local owner; V3 would investigate the 
memory of the local execution workload, including subprocesses. They answer 
different questions rather than being alternative names for V1.
   More details: https://github.com/apache/airflow/discussions/73645
   
   Would this narrowly defined runner diagnostic be useful enough to pursue 
first, if its export behavior can be made clear? Or should a task-memory metric 
aggregate concurrent runs or cover the broader workload from the outset?


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to