[
https://issues.apache.org/jira/browse/HDDS-16669?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Chu Cheng Li updated HDDS-16669:
--------------------------------
Attachment: screenshot-1.png
> [Ozone Local] Investigate heap and memory overhead under lightweight S3
> workloads
> ---------------------------------------------------------------------------------
>
> Key: HDDS-16669
> URL: https://issues.apache.org/jira/browse/HDDS-16669
> Project: Apache Ozone
> Issue Type: Sub-task
> Reporter: Chu Cheng Li
> Priority: Major
> Attachments: kuberay-history-server-resource-comparison.png
>
>
> Following the [Slack
> discussion|https://opensource4you.slack.com/archives/C07PLV9QNLF/p1790868752976099?thread_ts=1790686067.612819&cid=C07PLV9QNLF],
> investigate where memory is used in the single-JVM Ozone implementation
> under HDDS-14893. The KubeRay History Server S3/e2e workload provides a
> small-workload example, although this is not Ozone's primary optimization
> scenario.
> h3. Baseline
> The same TestCollector + TestHistoryServer workload at KubeRay commit
> 888e6145ab863cc4a23a5a0d9d1ceb9fa27c7f9b was run three times per backend,
> sequentially in rotated order, with an 8-CPU/8-GiB limit per store. All 12
> formal runs passed the main cases; the same upstream nested PID check was
> skipped in every run.
> ||Backend||Mean working set (MiB)||Mean RSS (MiB)||Mean CPU (cores)||Mean e2e
> (min)||
> |Single-JVM Ozone prototype|626|601|0.280|29.64|
> |Ozone 2.1.2 all-in-one (four JVMs)|1272|1238|0.302|29.35|
> |RustFS 1.0.0|243|221|0.022|29.32|
> |SeaweedFS 4.48|285|269|0.020|29.24|
> These are container working-set/RSS measurements, *not Java heap
> measurements*. Summaries cover the timed e2e window, excluding backend
> setup/readiness and the idle baseline. This is a functional workload with
> three repetitions, not a throughput or scalability benchmark.
> The prototype is the historical 2.2.0-SNAPSHOT implementation from
> [peterxcli/ozone PR #13|https://github.com/peterxcli/ozone/pull/13] at
> 8f12228177907df2a6d979b1d0e90911b4814397: SCM, OM, DataNode and S3G in one
> JVM, 512-MiB maximum heap, standalone/one replica. The all-in-one comparator
> uses the original KubeRay configuration with four JVMs at 128 MiB each.
> Version and replication/configuration differences prevent attributing the
> observed difference solely to JVM consolidation.
> h3. Investigation
> * Profile startup, idle and e2e phases, recording heap used/committed/max
> separately from RSS and container working set.
> * Identify the largest retained heap consumers and quantify native/off-heap
> contributors where possible, including RocksDB, direct buffers and caches.
> * Document the reproduction configuration, profiling procedure and findings;
> identify actionable follow-up optimizations.
> * Validate proposed changes against the same functional suite and report
> memory/runtime tradeoffs.
> The resource comparison chart accompanies this baseline. Its traces show one
> observed median-duration repetition per backend; the summary above averages
> all three repetitions. CPU and memory cover store processes; pod network
> includes the common client and excludes loopback. Per-store disk I/O was
> unavailable in the rootless environment.
> This task tracks investigation and follow-up opportunities; it does not set a
> numerical reduction target or release deadline.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]