Hi Gregory,

Thanks for sharing the working branch.

Some of us at Samsung are testing the series (both v4 & v5) on a 
compression capable CXL expander and are observing encouraging results.

Good things first! We feel that the isolation using private nodes and 
the capability selection using NODE_PRIVATE_CAP_* work functionally well 
for the compressed memory use cases.

On improvement points, we observed much higher 'page_faults', 
'allocation_stalls' etc. which contributed to the higher tail latencies 
in some tests. I hope these are already part of the optimization plans. 
The higher latencies might have resulted from the write_protection 
applied to the private node, I believe.

On the tools side, we used MLC and Taobench for the tests.

Note that our intention for this phase of experiments was to test only 
the 'private node' layer based isolation, and not the cram and below 
layers for the compressed memory. Therefore we did not enable/utilize 
the compression capability in the hardware for these set of experiments 
(enabling compression would require the cram level ballooning/memory 
shrinking and cxl level interfacing driver to the compression device). 
That would be different set of experiments where we would be testing the 
cram balloon shrinkers and our alternate algorithms (upstream targeted). 
So cram and below layers were used only for enumeration of private node 
regions for these tests and not for the run-time memory shrinkers.

======= Test Methodology ==============

We explored these 3 cases:

    1) (Baseline – existing CXL infra):  DRAM 16GB + CXL 32GB. Here we 
used the existing CXL driver infra to enable the device. No private node 
is involved. And compression is disabled on device.

    2) (private node in default write protected path): DRAM 16GB + CXL 
Private Node 32GB. Here private node infra is used to enable the device. 
And compression is disabled on device.

    3) (private node with no write protection): DRAM 16GB + CXL Private 
Node 32GB. Here private node infra is used to enable the device without 
the write protection/fencing enabled. Compression is disabled on device. 
We could not complete these runs as they resulted in kernel panics. 
Guess the code path is not stable yet for this (We had hoped that case 1 
and case 3 results would be similar). We also observed some unmovable 
page warnings logs in 'dmesg' before crash (might be related to the 
panic). Adding the log snippets at the end.

========= Results summary ===================

- TaoBench: ‘private node’ showed much higher latencies. We believe this 
could be due to the migration back into the DRAM for overwrites (due to 
write-protection) based on the kernel VM stats.

- MLC:  For lower inject delay, the private node shows better latency; 
but when inject delay goes high, the baseline is better.

Overall kernel VM stats suggested much higher ‘page faults, ‘allocation 
stalls’ etc. in ‘private node’ case which is corroborated with the results.

========== Setup ======================

CPU: 144 Cores (single socket), DRAM: 16GB, CXL: 32GB,
Kernel: 7.2.0-rc2+ (Private Node), 6.18.32 (Baseline)

Private node opt CAPs:
#define CRAM_NP_CAPS    (NODE_PRIVATE_CAP_RECLAIM |
                         NODE_PRIVATE_CAP_HOTUNPLUG | \
                         NODE_PRIVATE_CAP_DEMOTION |
                         NODE_PRIVATE_CAP_USER_NUMA | \
                         NODE_PRIVATE_CAP_NUMA_BALANCING |
                         NODE_PRIVATE_POLICY_WRITE_FENCE)

Enabled 'numa balancing' and 'demotion'.
   $ echo 1 > /proc/sys/kernel/numa_balancing
   $ echo true > /sys/kernel/mm/numa/demotion_enabled

========== TaoBench Experiments ====================

CMD:
$ python3 benchpress_cli.py run tao_bench_standalone -i '{"bind_mem": 0, 
"bind_cpu": 0, "memsize": 32, "test_time": 300, "warmup_time": 0, 
"clients_per_thread": 100, "set_get_ratio": "1:9"}'

1. (baseline) DRAM 16GB + CXL 32GB   ==>

$ grep -A8 "^ALL STATS" benchmark_metrics_27cfff0e/client_0.log

===================================================
Type  Avg. Latency    p50        p95        p99
----------------------------------------------------
Sets   2.75215     2.73500      4.57500     5.79100
Gets   2.84236     2.81500      4.76700     5.95100

  2. (private node in default write protected path)
      DRAM 16GB + CXL Private Node(uncompressed) 32GB   ===>

$ grep -A8 "^ALL STATS" benchmark_metrics_5181fda5/client_0.log

===================================================
Type  Avg. Latency    p50        p95        p99
----------------------------------------------------
Sets   5.58395     4.73500      12.35100     21.75900
Gets   4.83344     4.19100      10.43100     17.91900

Vmstat metrics comparison of both cases for Taobench
($cat /proc/vmstat):

-------------------------------------------------------------------
Metric           |  Baseline    |   CXL Private Node |   Ratio
-------------------------------------------------------------------
pgfault            37,121,808       171,512,616          4.6x
pgactivate          5,360,928       161,947,594         30.2x
pgreuse             4,346,396           407,397         10.7x less
pgrefill           88,775,854       436,486,518          4.9x
pgdemote_kswapd     5,933,439       129,586,750         21.8x
pgdemote_direct        24,837        35,505,114       1429x
allocstall_normal         469           153,600        327x
kswapd_low_wmark_hit_quickly 671         27,821         41.5x
pageoutrun              1,333            29,217         21.9x
numa_pte_updates   24,508,077                 0          na
numa_hint_faults   21,152,264                 0          na
numa_pages_migrated 3,330,897                 0          na
numa_hit           46,667,202       353,812,703          7.6x

---------- Takeaways -----

Private (CRAM) node performance is lower compared to a normal CXL memory 
allocation path. VM stats shows more page_faults and allocation_stalls 
on the private node case. Could it be the allocator waits for migration 
path to demote pages to private node? Or the actual hot pages in dram 
got demoted to private node to make space during overwrites? We will try 
further analysis on this.

========== MLC Experiments ====================

CMD:
   $./mlc --loaded_latency -j0 -c0 -b1g -k1-15 -W5 -r

Experiment Results:

1. (Baseline) DRAM 16GB + CXL 32GB

Inject  Latency Bandwidth
Delay   (ns)    MB/sec
================
  00000  686.18   40137.5
  00002  686.69   40124.0
  00008  709.74   40362.7
  00015  711.75   40350.0
  00050  710.72   40357.4
  00100  712.22   40316.0
  00200  695.29   40026.5
  00300  602.91   38845.3
  00400  482.28   36240.4
  00500  410.03   33606.6
  00700  275.13   26352.5
  01000  241.82   18851.7
  01300  230.91   14695.8
  01700  215.27   11403.2
  02500  217.76    7900.9
  03500  194.42    5787.0
  05000  192.17    4166.4
  09000  190.36    2475.0
  20000  188.09    1312.1

2. (private node in default write protected path)
     DRAM 16GB + CXL Private Node(Uncompressed) 32GB

Inject  Latency Bandwidth
Delay   (ns)    MB/sec
================
  00000  320.09   12186.2
  00002  543.42   22726.1
  00008  482.11   28385.2
  00015  488.26   26304.6
  00050  485.19   26898.4
  00100  477.65   24668.9
  00200  483.68   21273.8
  00300  486.44   17655.9
  00400  480.41   16470.0
  00500  481.30   13589.9
  00700  493.84    9960.8
  01000  488.66    8267.7
  01300  491.31    6433.6
  01700  476.97    6211.2
  02500  472.97    5037.0
  03500  466.05    3533.8
  05000  446.44    2783.1
  09000  417.03    1877.5
  20000  387.22    1035.3

Vmstat metrics comparison of both cases for MLC
($cat /proc/vmstat):

-------------------------------------------------------------------
Metric           |  Baseline    |   CXL Private Node |   Ratio
-------------------------------------------------------------------
pgfault           10,287,329         46540,782           4.6x
pgactivate           1141468        65,098,272          57x
pgrefill             260,699        47,592,751         182.5x
pgdemote_kswapd    1,085,783        26,063,754          24x
pgdemote_direct    4,562,539        15,809,485           3.5x
pgmigrate_fail        41,929         2,763,725          65.9x
allocstall_normal        161               425           2.6x
kswapd_low_wmark_hit_quickly  47           763          16.2x
pageoutrun                55               936          17x

Takeaway: Our MLC tests show a trade-off between the two configurations. 
During low inject delay, the private Node is faster. But when inject 
delay goes high, the baseline shows better latency. Could it be the TLB 
cache effects?

--------------------------------------------------------------------
Case 3 dmesg log snippet:

[   82.525741] page: refcount:1 mapcount:0 mapping:0000000000000000 
index:0x0 pfn:0x5f5800
[   82.525753] flags: 
0x17ffffc0002000(reserved|node=0|zone=2|lastcpupid=0x1fffff)
[   82.525762] raw: 0017ffffc0002000 ffefc85717d60008 ffefc85717d60008 
0000000000000000
[   82.525764] raw: 0000000000000000 0000000000000000 00000001ffffffff 
0000000000000000
[   82.525766] page dumped because: unmovable page
..............

We will update on the compression enabled experiments further.

~ Arun

On 23-07-2026 10:23 pm, Gregory Price wrote:
> On Thu, Jul 23, 2026 at 02:08:31PM +0530, Arun George/Arun George wrote:
>> On 21-07-2026 01:03 am, Gregory Price wrote:
>> We intend to test this series on a compression capable CXL expander
>> hardware. Since compressed ram example (cram) is not part of this
>> series, how do you suggest to do that? Do you have a version of cram
>> module compatible with this series?
>>
>> ~Arun
> 
> Hi Arun,
> > I plan to RFC the new setup for compressed ram this a bit later, but I
> will share my working branch with you for testing.
> 
> ~Gregory



Reply via email to