awsomesud347 commented on issue #12123: URL: https://github.com/apache/seatunnel/issues/12123#issuecomment-5554351184
#12130 covers part of this. The trace and the numbers are in that pull request description, so I will not repeat them here. This is what it leaves open, for whoever picks it up. The listing endpoints were never the problem. `/running-jobs` and `/finished-jobs` do not reach `getJobStatusData()` at all, since `JobInfoService` reads the IMaps directly. Worth knowing for anyone else working #12126. What does reach it is `GET /metrics` and `GET /openmetrics`, once per Prometheus scrape on the master, plus the CLI which is one shot per invocation. What remains is that scrape path. Counting ten thousand retained jobs allocates roughly 200 MB per operation, almost entirely from deserialising `JobState` values, and it happens on every scrape. One approach is already ruled out, so nobody spends time on it twice. Folding statuses without building a `JobStatusData` per retained job changes nothing measurable: 0.19% latency difference with overlapping error bars, and 198.66 against 198.71 MB allocated per operation. The cost is the IMap fetch itself, not the objects built from it, so avoiding allocation is the wrong target. The untested idea is member side aggregation, so that only counts cross the network instead of every value. I did not measure it, because my benchmarks ran against a single member cluster where nothing crosses a network in the first place. It needs a multi member setup to evaluate, and I would not assume it helps without that. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
