awsomesud347 commented on issue #12123:
URL: https://github.com/apache/seatunnel/issues/12123#issuecomment-5554351184

   #12130 covers part of this. The trace and the numbers are in that pull 
request description, so I
   will not repeat them here. This is what it leaves open, for whoever picks it 
up.
   
   The listing endpoints were never the problem. `/running-jobs` and 
`/finished-jobs` do not reach
   `getJobStatusData()` at all, since `JobInfoService` reads the IMaps 
directly. Worth knowing for
   anyone else working #12126. What does reach it is `GET /metrics` and `GET 
/openmetrics`, once per
   Prometheus scrape on the master, plus the CLI which is one shot per 
invocation.
   
   What remains is that scrape path. Counting ten thousand retained jobs 
allocates roughly 200 MB per
   operation, almost entirely from deserialising `JobState` values, and it 
happens on every scrape.
   
   One approach is already ruled out, so nobody spends time on it twice. 
Folding statuses without
   building a `JobStatusData` per retained job changes nothing measurable: 
0.19% latency difference
   with overlapping error bars, and 198.66 against 198.71 MB allocated per 
operation. The cost is the
   IMap fetch itself, not the objects built from it, so avoiding allocation is 
the wrong target.
   
   The untested idea is member side aggregation, so that only counts cross the 
network instead of
   every value. I did not measure it, because my benchmarks ran against a 
single member cluster where
   nothing crosses a network in the first place. It needs a multi member setup 
to evaluate, and I
   would not assume it helps without that.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to