[ 
https://issues.apache.org/jira/browse/HDDS-16669?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121789#comment-18121789
 ] 

Chu Cheng Li commented on HDDS-16669:
-------------------------------------

cc [~weichiu]

> [Ozone Local] Investigate heap and memory overhead under lightweight S3 
> workloads
> ---------------------------------------------------------------------------------
>
>                 Key: HDDS-16669
>                 URL: https://issues.apache.org/jira/browse/HDDS-16669
>             Project: Apache Ozone
>          Issue Type: Sub-task
>            Reporter: Chu Cheng Li
>            Priority: Major
>         Attachments: kuberay-history-server-resource-comparison.png
>
>
> Investigate where memory is used in the single-JVM Ozone implementation under 
> HDDS-14893. The KubeRay History Server S3/e2e workload provides a 
> small-workload example, although this is not Ozone's primary optimization 
> scenario.
> h3. Baseline
> The same TestCollector + TestHistoryServer workload at KubeRay commit 
> [888e6145ab863cc4a23a5a0d9d1ceb9fa27c7f9b|https://github.com/win5923/kuberay/commit/888e6145ab863cc4a23a5a0d9d1ceb9fa27c7f9b]
>  was run three times per backend, sequentially in rotated order, with an 
> 8-CPU/8-GiB limit per store. All 12 formal runs passed the main cases; the 
> same upstream nested PID check was skipped in every run.
> ||Backend||Mean working set (MiB)||Mean RSS (MiB)||Mean CPU (cores)||Mean e2e 
> (min)||
> |Single-JVM Ozone prototype|626|601|0.280|29.64|
> |Ozone 2.1.2 all-in-one (four JVMs)|1272|1238|0.302|29.35|
> |RustFS 1.0.0|243|221|0.022|29.32|
> |SeaweedFS 4.48|285|269|0.020|29.24|
> These are container working-set/RSS measurements, *not Java heap 
> measurements*. Summaries cover the timed e2e window, excluding backend 
> setup/readiness and the idle baseline. This is a functional workload with 
> three repetitions, not a throughput or scalability benchmark.
> The prototype is the historical 2.2.0-SNAPSHOT implementation from 
> [peterxcli/ozone PR #13|https://github.com/peterxcli/ozone/pull/13] at 
> 8f12228177907df2a6d979b1d0e90911b4814397: SCM, OM, DataNode and S3G in one 
> JVM, 512-MiB maximum heap, standalone/one replica. The all-in-one comparator 
> uses the original KubeRay configuration with four JVMs at 128 MiB each. 
> Version and replication/configuration differences prevent attributing the 
> observed difference solely to JVM consolidation.
> h3. Investigation
> * Profile startup, idle and e2e phases, recording heap used/committed/max 
> separately from RSS and container working set.
> * Identify the largest retained heap consumers and quantify native/off-heap 
> contributors where possible, including RocksDB, direct buffers and caches.
> * Document the reproduction configuration, profiling procedure and findings; 
> identify actionable follow-up optimizations.
> * Validate proposed changes against the same functional suite and report 
> memory/runtime tradeoffs.
> The resource comparison chart accompanies this baseline. Its traces show one 
> observed median-duration repetition per backend; the summary above averages 
> all three repetitions. CPU and memory cover store processes; pod network 
> includes the common client and excludes loopback. Per-store disk I/O was 
> unavailable in the rootless environment.
> This task tracks investigation and follow-up opportunities; it does not set a 
> numerical reduction target or release deadline.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to