On Thu, 13 Aug 2026 16:36:27 GMT, Andrew Haley <[email protected]> wrote:
>> Please use [this >> link](https://github.com/openjdk/jdk/pull/28541/changes?w=1) to view the >> files changed. >> >> Profile counters scale very badly. >> >> The overhead for profiled code isn't too bad with one thread, but as the >> thread count increases, things go wrong very quickly. >> >> For example, here's a benchmark from the OpenJDK test suite, run at >> TieredLevel 3 with one thread, then three threads: >> >> >> Benchmark (randomized) Mode Cnt Score Error Units >> InterfaceCalls.test2ndInt5Types false avgt 4 27.468 ± 2.631 ns/op >> InterfaceCalls.test2ndInt5Types false avgt 4 240.010 ± 6.329 ns/op >> >> >> This slowdown is caused by high memory contention on the profile counters. >> Not only is this slow, but it can also lose profile counts. >> >> This patch is for C1 only. It'd be easy to randomize C1 counters as well in >> another PR, if anyone thinks it's worth doing. >> >> One other thing to note is that randomized profile counters degrade very >> badly with small decimation ratios. For example, using a ratio of 2 with >> `-XX:ProfileCaptureRatio=2` with a single thread results in >> >> >> Benchmark (randomized) Mode Cnt Score Error >> Units >> InterfaceCalls.test2ndInt5Types false avgt 4 80.147 ± 9.991 >> ns/op >> >> >> The problem is that the branch prediction rate drops away very badly, >> leading to many mispredictions. It only really makes sense to use higher >> decimation ratios, e.g. 64. >> >> --------- >> - [x] I confirm that I make this contribution in accordance with the >> [OpenJDK Interim AI Policy](https://openjdk.org/legal/ai). > > Andrew Haley has updated the pull request with a new target base due to a > merge or a rebase. The pull request now contains 201 commits: > > - Fix RISC V > - Fix RISCV > - Fix merge > - Fix RISCV > - Fix Arm-32 after merge. > - Merge branch 'JDK-8134940' of https://github.com/theRealAph/jdk into > JDK-8134940 > - Merge branch 'master' into JDK-8134940 > - _c1_needs_stack_repair is useless and should be removed > - Tidy remove_frame() > - Fix conflict > - ... and 191 more: https://git.openjdk.org/jdk/compare/b1467bd6...5119f86c Finally had some cycles to do some performance investigations. Were you able to find a larger workload where this improves performance? I assume it should have a positive impact on warmup on multi-threaded tests that linger in profiled code for a while, accruing coherence stalls. I have been looking through Dacapo and Renaissance ...and all I see so far are regressions the moment we do `PCR > 1`. I'll keep looking. For example, here is Renaissance/Dotty on my 16-core 7950X3D, using [this simple script](https://github.com/user-attachments/files/31304499/run-renaissance.sh): $ ./run-renaissance.sh dotty -r 20 dotty -r 20 --- Patched, PCR=1 Run 1: 3792 1332 819 698 605 564 515 512 493 479 450 450 441 451 444 426 413 420 398 406 Run 2: 3927 1183 809 642 618 600 597 552 577 524 479 479 451 465 471 436 411 430 409 405 Run 3: 3851 1172 816 616 573 562 523 495 483 528 503 523 519 518 507 449 442 453 432 428 Tier1: 0.068 s, 9631 bytes, 2110 methods; Tier2: 0.016 s, 12180 bytes, 56 methods; Tier3: 5.013 s, 2747152 bytes, 17149 methods; Tier4: 53.617 s, 6588104 bytes, 6917 methods; Uncommon traps hit: 673 --- Patched, PCR=4 Run 1: 4875 1660 1081 929 792 729 665 613 591 571 554 554 506 485 491 474 501 506 470 465 Run 2: 4934 1605 1029 816 755 779 663 630 563 550 535 543 512 498 479 454 463 465 437 496 Run 3: 5196 1720 1072 847 749 733 686 668 590 602 568 554 525 502 508 464 448 467 451 431 Tier1: 0.067 s, 9647 bytes, 2115 methods; Tier2: 0.015 s, 11627 bytes, 55 methods; Tier3: 5.205 s, 2972539 bytes, 17387 methods; Tier4: 57.010 s, 7038828 bytes, 7314 methods; Uncommon traps hit: 963 --- Patched, PCR=16 Run 1: 4757 1631 1107 838 752 672 617 602 556 611 576 593 544 516 495 483 465 475 454 429 Run 2: 4814 1693 1177 930 747 708 633 584 534 531 523 510 502 479 474 474 465 457 442 446 Run 3: 4657 1721 1100 903 742 701 645 582 561 528 513 511 491 479 460 448 441 427 449 439 Tier1: 0.066 s, 9708 bytes, 2126 methods; Tier2: 0.008 s, 5857 bytes, 41 methods; Tier3: 5.724 s, 3304344 bytes, 17844 methods; Tier4: 62.403 s, 7877781 bytes, 7714 methods; Uncommon traps hit: 1357 --- Patched, PCR=64 Run 1: 4594 1601 1184 914 824 789 693 645 595 569 541 517 513 486 472 466 457 471 469 436 Run 2: 4433 1663 1152 931 750 724 726 661 612 597 541 539 532 521 492 473 457 474 446 441 Run 3: 4845 1657 1129 943 761 696 627 681 625 617 585 515 512 503 500 469 468 454 453 443 Tier1: 0.075 s, 9804 bytes, 2146 methods; Tier2: 0.011 s, 6797 bytes, 57 methods; Tier3: 6.618 s, 3650119 bytes, 18345 methods; Tier4: 71.218 s, 8939902 bytes, 8315 methods; Uncommon traps hit: 1895 --- Patched, PCR=256 Run 1: 4515 1712 1271 986 837 720 655 608 587 563 541 554 551 546 553 547 479 470 491 478 Run 2: 4447 1879 1225 1006 843 749 697 634 607 567 536 526 582 672 530 528 505 463 485 478 Run 3: 4467 1705 1315 1051 891 765 689 644 598 559 529 511 501 495 477 473 479 459 448 446 Tier1: 0.070 s, 9887 bytes, 2163 methods; Tier2: 0.014 s, 9360 bytes, 52 methods; Tier3: 6.707 s, 3850573 bytes, 18907 methods; Tier4: 76.892 s, 9560744 bytes, 8857 methods; Uncommon traps hit: 2424 --- Patched, PCR=1024 Run 1: 4267 1704 1202 966 959 794 740 661 584 550 533 547 525 522 497 487 461 508 474 454 Run 2: 4358 1607 1199 961 818 819 748 676 627 585 548 545 532 521 479 470 475 467 460 439 Run 3: 4307 1625 1124 909 809 714 677 685 690 627 583 552 523 515 499 472 499 469 464 445 Tier1: 0.086 s, 10013 bytes, 2190 methods; Tier2: 0.018 s, 12772 bytes, 75 methods; Tier3: 7.179 s, 3953801 bytes, 19313 methods; Tier4: 72.033 s, 9002379 bytes, 9647 methods; Uncommon traps hit: 2894 --- Patched, PCR=4096 Run 1: 4351 1754 1279 1073 889 804 723 687 646 610 697 661 672 612 567 566 548 539 523 512 Run 2: 4134 1761 1365 1064 907 811 774 713 667 620 613 586 577 565 562 544 528 522 519 524 Run 3: 4124 1782 1344 1054 920 819 772 726 678 652 615 599 575 600 576 549 541 529 511 511 Tier1: 0.064 s, 9824 bytes, 2149 methods; Tier2: 0.000 s, 0 bytes, 0 methods; Tier3: 6.098 s, 3793398 bytes, 19034 methods; Tier4: 46.714 s, 5889248 bytes, 9216 methods; Uncommon traps hit: 2571 --- Patched, PCR=16384 Run 1: 4176 1844 1684 1382 1116 1063 973 964 872 833 808 784 756 745 715 685 690 671 662 648 Run 2: 4234 2083 1560 1271 1126 1016 959 894 879 817 807 778 752 719 722 678 671 651 644 650 Run 3: 4404 1970 1474 1296 1139 1005 975 907 851 811 792 762 736 703 697 703 671 665 659 641 Tier1: 0.072 s, 9979 bytes, 2178 methods; Tier2: 0.000 s, 74 bytes, 1 methods; Tier3: 5.741 s, 3518511 bytes, 18537 methods; Tier4: 26.872 s, 3289376 bytes, 8434 methods; Uncommon traps hit: 2046 --- Patched, PCR=65536 Run 1: 4619 2346 1982 1993 1791 1621 1427 1374 1285 1230 1190 1355 1264 1150 1082 1045 1033 966 983 952 Run 2: 4826 2336 1924 1757 1577 1676 1522 1349 1286 1249 1185 1157 1102 1065 1035 1009 986 973 950 925 Run 3: 4730 2669 2036 1781 1593 1474 1417 1554 1433 1284 1220 1184 1165 1113 1069 1009 985 960 937 921 Tier1: 0.055 s, 10090 bytes, 2204 methods; Tier2: 0.000 s, 116 bytes, 3 methods; Tier3: 4.707 s, 3231221 bytes, 18031 methods; Tier4: 14.596 s, 1842916 bytes, 5499 methods; Uncommon traps hit: 1421 I think this speaks to the _hypothesis_ I had in my first comment here: https://github.com/openjdk/jdk/pull/28541#issuecomment-3586901495: a randomized profile, even at lower PCRs, is distorted enough to cause significant recompilation churn due to subsampling 1->0 updates that would yield about-to-be-hit uncommon traps in C2 generated code. We are hitting 3x (!) more uncommon traps on mid-level PCRs; I suspect that is why warmup is that much worse. With high-level PCRs, from the look at compilation times, we seem to just linger in profiling code much longer, never getting to tier4, so warmup is worse again. That seems to speak against the idea that we have enough profiling traffic updates to level out the "stepped" increases in counters. So the more I look into performance model of this approach, the weirder and weirder feelings I feel. I get that ultimately it might come to selecting the good PCR for the concrete workload. But I struggle to find a workload yet where "good PCR" even exists... Something is off somewhere. ------------- PR Comment: https://git.openjdk.org/jdk/pull/28541#issuecomment-5369545884
