[
https://issues.apache.org/jira/browse/GROOVY-12191?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18099181#comment-18099181
]
ASF GitHub Bot commented on GROOVY-12191:
-----------------------------------------
daniellansun commented on PR #2736:
URL: https://github.com/apache/groovy/pull/2736#issuecomment-5080089470
# GROOVY-12191 Performance Verification Report
**Scoped indy SwitchPoint invalidation**
| Item | Value |
|---|---|
| Issue | [GROOVY-12191](https://issues.apache.org/jira/browse/GROOVY-12191)
|
| Commit under test | `ae8e49741e74067371823148c8993e8c0b73e317` |
| Baseline (parent) | `0ca665dad0d4070bfa218d9b5577ec423c8b2181` |
| Date | 2026-07-26 |
| Verdict | **PASS** — cross-type MetaClass churn no longer globally
deoptimizes hot monomorphic call sites; steady-state hot path is parity;
correctness semantics retained |
---
## 1. Executive summary
This commit changes Groovy’s invokedynamic MOP `SwitchPoint` invalidation
domain from a **single process-wide SwitchPoint** to a **per-class domain**
attached to `ClassInfo` (with hierarchy fan-out to loaded
subtypes/implementors). Category enter/leave and unscoped MetaClass events
still bulk-retire loaded class domains. Linked call sites still carry **one**
SwitchPoint guard on the hot path—the same monomorphic shape as before 6.0.
On the same host, JDK, and JMH annotation settings, A/B measurement shows:
1. **Near-zero regression on steady-state monomorphic hot paths** (baseline
ratios ≈ 0.99–1.05×, within noise).
2. **Large gains under cross-type MetaClass churn** — the primary win:
- `CallSiteInvalidationBench.crossTypeInvalidationEvery1000`: **405.6 →
2.77 ms/op** (≈ **146×**).
- `ScopedInvalidationBench.hotLoop_unrelatedMetaClassChurn`: throughput ≈
**156×**.
3. **Framework-style “unrelated burst then steady request loop”**
(`burstThenSteadyState`): ≈ **76×**.
4. **Same-type invalidation and Category semantics are not “falsely
optimized”**: same-type churn remains far slower than baseline (forced
re-link); category enter/leave stays bulk and matches parent throughput.
5. **Correctness: 56/56 tests passed** (full scoped-invalidation suite
introduced by this commit).
> In one line: without adding a second hot-path guard, the change shrinks
the blast radius of unrelated MetaClass mutations from process-global to
type-scoped, with measured gains on the order of **10¹–10²×** (depending on
churn frequency and site count).
---
## 2. Change mechanism and performance hypotheses
### 2.1 Previous model (parent)
- `IndyInterface.switchPoint` was a **process-wide shared** `SwitchPoint`.
- Any MetaClass registry change → `invalidateSwitchPoints()` → **every**
linked indy site’s guard failed → mass re-link.
- Typical symptom: a framework/plugin extends MetaClass on *ColdType* at
startup or hot-reload; application hot paths on *HotTarget.compute()* were
deoptimized together with it.
### 2.2 New model (`ae8e497`)
| Component | Role |
|---|---|
| `SwitchPointInvalidator` | Single-domain lifecycle: CAS `get` /
`detachLive` / `invalidate` |
| `IndyInvalidation` | Policy: `invalidateClass` (+ hierarchy),
`invalidateCategory` (bulk), `guardWithMopSwitchPoints` |
| `ClassInfo` | Holds one `indySwitchPoint` per class; coordinates version
bumps with domain retirement |
| `Selector` / `IndyCompoundAssign` / `ColdReflectiveMethodHandleWrapper` |
Install a **per-receiver-class** SwitchPoint guard at link time |
**Hot-path shape (intentionally monomorphic):**
```text
handle = classSp.guardWithTest(fast, fallback)
```
Still a single guard—matching the pre-6.0 monomorphic shape. There is **no**
second category SwitchPoint on the hot path.
**Invalidation policy:**
| Event | Behavior | Expected hot-path cost |
|---|---|---|
| Unrelated-type MetaClass change | Retire that type’s domain (and loaded
subtypes) only | Hot sites **do not** re-link |
| Same-type / parent MetaClass change | Retire class + hierarchy fan-out |
Must re-link (correctness) |
| Category enter/leave, `invalidateCallSites()` | Bulk-retire all loaded
class domains | Same order of cost as old global invalidation |
| Unattributed MetaClass event | `invalidateUnscoped()` bulk | Conservative,
correctness-first |
### 2.3 Hypotheses under test
| ID | Hypothesis | How verified |
|---|---|---|
| H1 | Baseline hot loop has no material regression | A/B ratio ≈ 1 |
| H2 | Cross-type churn throughput/latency improves substantially |
Unrelated ≫ parent; on HEAD, unrelated ≫ same-type |
| H3 | Pure hot loop after unrelated burst approaches baseline |
`afterUnrelatedBurst` ≈ `baseline` |
| H4 | Same-type / hierarchy / category still pay re-link tax | Clearly
slower than baseline |
| H5 | Correctness suite stays green | Unit / integration tests |
---
## 3. Methodology
### 3.1 Environment
| Item | Value |
|---|---|
| OS | Linux 6.15.5 x86_64 (hostname: hera) |
| CPU | AMD EPYC 7763 (6 vCPUs visible in container) |
| Memory | 23 GiB |
| JDK | Amazon Corretto **25.0.2** (`25.0.2+10-LTS`) |
| Build | Gradle wrapper, `./gradlew :perf:jmh`, indy enabled by default |
| Trees | HEAD: project worktree; parent: worktree at `0ca665dad0` |
| Bench source parity | `ScopedInvalidationBench` workload identical on
parent (Javadoc-only diff) |
### 3.2 Benchmark suites
1. **`org.apache.groovy.bench.ScopedInvalidationBench`** (added by this
commit; Throughput, ops/ms)
- 100 000 monomorphic calls per op; MetaClass write every 1000 iterations
in churn cases.
- JMH: `@Warmup(3×2s) @Measurement(5×2s) @Fork(2)` → 10 samples per
benchmark.
2. **`org.apache.groovy.perf.grails.CallSiteInvalidationBench`** (existing
Grails-oriented suite; AverageTime, ms/op)
- Same 100 000-iteration loops; cross-type / same-type / burst patterns.
- Same JMH annotation settings.
### 3.3 Correctness
```text
./gradlew :test --tests Groovy12191 \
--tests org.apache.groovy.runtime.indy.IndyInvalidationTest \
--tests org.apache.groovy.runtime.indy.SwitchPointInvalidatorTest \
--tests org.codehaus.groovy.vmplugin.v8.IndyScopedSwitchPointTest
--rerun-tasks
```
**Result: 56 passed / 0 failed / 0 error.**
| Test class | Cases |
|---|---:|
| `bugs.Groovy12191` | 8 |
| `IndyInvalidationTest` | 27 |
| `SwitchPointInvalidatorTest` | 12 |
| `IndyScopedSwitchPointTest` | 9 |
### 3.4 Statistics and artifacts
- Scores are JMH means; `±` is JMH’s default **99.9% CI**.
- **Throughput** ratio = HEAD / parent (>1 is better).
- **AverageTime** speedup = parent / HEAD (>1 is better).
- Builds ran **sequentially** to avoid dual-JVM core contention; no CPU
pinning. Conclusions rest on **order-of-magnitude and structural separation**,
not ±1% micro-deltas.
Raw JSON (local run artifacts, not committed):
- `/tmp/groovy-12191-perf-results/head/scoped-results.json`
- `/tmp/groovy-12191-perf-results/parent/scoped-results.json`
- `/tmp/groovy-12191-perf-results/head/callsite-results.json`
- `/tmp/groovy-12191-perf-results/parent/callsite-results.json`
---
## 4. Results
### 4.1 ScopedInvalidationBench (Throughput, ops/ms — higher is better)
| Benchmark | Parent | HEAD | HEAD/Parent | Interpretation |
|---|---:|---:|---:|---|
| `hotLoop_baseline` | 3.891 ± 0.224 | 3.869 ± 0.219 | **0.99×** | H1: no
hot-path regression |
| `hotLoop_afterUnrelatedBurst` | 3.774 ± 0.201 | 4.522 ± 0.168 | **1.20×**
| H3: full speed after unrelated burst |
| `hotLoop_unrelatedMetaClassChurn` | 0.00247 ± 0.00017 | 0.386 ± 0.078 |
**≈156×** | **H2 primary win** |
| `hotLoop_sameTypeMetaClassChurn` | 0.00249 ± 0.00017 | 0.0188 ± 0.0012 |
**≈7.6×** | Still forced re-link; narrower blast radius than parent |
| `hotLoop_categoryEnterLeave` | 0.00225 ± 0.00045 | 0.00231 ± 0.00028 |
**≈1.03×** | H4: bulk semantics retained |
| `hierarchy_baseline` | 3.325 ± 0.118 | 2.896 ± 0.113 | 0.87× | Noise /
load variance; both full-speed order |
| `hierarchy_parentMetaClassChurn` | 0.00233 ± 0.00021 | 0.00957 ± 0.00080 |
**≈4.1×** | H4: parent change still invalidates child sites |
**Structural fingerprint on HEAD only:**
```text
baseline ≈ 3.87 ops/ms
unrelated churn ≈ 0.39 ops/ms (~20× slower than baseline: mostly MetaClass
write cost)
same-type churn ≈ 0.019 ops/ms (~200× slower than baseline: write +
same-class re-link)
category ≈ 0.002 ops/ms (bulk re-link tax)
```
On parent, `unrelated` and `same-type` **both collapse to ~0.002**, showing
that the old global SwitchPoint made “unrelated churn” equivalent to “global
deopt.” HEAD separates them by ~**20×** (0.386 / 0.019)—a direct fingerprint
that scoping is live.
### 4.2 CallSiteInvalidationBench (AverageTime, ms/op — lower is better)
| Benchmark | Parent | HEAD | Speedup (P/H) | Interpretation |
|---|---:|---:|---:|---|
| `baselineHotLoop` | 0.257 ± 0.017 | 0.248 ± 0.014 | **1.04×** | H1 |
| `baselineListSize` | 0.242 ± 0.010 | 0.236 ± 0.006 | **1.03×** | H1 |
| `baselineSteadyStateNoBurst` | 0.612 ± 0.029 | 0.583 ± 0.033 | **1.05×** |
H1 |
| `baselineMultipleCallSites` | 20.42 ± 1.03 | 18.86 ± 0.54 | **1.08×** | H1
|
| `crossTypeInvalidationEvery10000` | 288.5 ± 24.3 | 2.73 ± 1.20 | **≈106×**
| H2 |
| `crossTypeInvalidationEvery1000` | 405.6 ± 27.3 | 2.77 ± 0.88 | **≈146×**
| **H2 primary win** |
| `crossTypeInvalidationEvery100` | 137.4 ± 10.2 | 12.5 ± 1.0 | **≈11×** |
Still large under denser churn |
| `listSizeWithCrossTypeInvalidation` | 330.7 ± 25.0 | 2.21 ± 0.53 |
**≈150×** | JDK-type hot sites benefit too |
| `multipleCallSitesWithInvalidation` | 1098 ± 89 | 25.9 ± 1.7 | **≈42×** |
Multi-site still much improved |
| `burstThenSteadyState` | 145.8 ± 30.6 | 1.93 ± 0.16 | **≈76×** |
“Framework MC extend + request steady state” |
| `sameTypeInvalidationEvery1000` | 407.5 ± 29.4 | 53.3 ± 3.2 | **≈7.6×** |
Same-type still slow; local invalidation cheaper than global rotate |
**Penalty vs baseline on HEAD:**
| Scenario | Relative cost | Meaning |
|---|---:|---|
| crossType @1000 | 2.77 / 0.25 ≈ **11×** baselineHotLoop | Dominated by
~100 ColdType MetaClass writes, not hot-site deopt |
| sameType @1000 | 53.3 / 0.25 ≈ **215×** baselineHotLoop | Clear re-link
tax |
| burstThenSteady | 1.93 / 0.58 ≈ **3.3×** steady baseline | Burst write
cost dominates; steady calls recovered |
On parent, crossType@1000 and sameType@1000 are **both ~406 ms/op**, again
confirming global invalidation collapsed cross-type into full deopt.
---
## 5. Why it is faster
```text
Old path New path
ColdType.metaClass.foo = ... ColdType.metaClass.foo = ...
│ │
▼ ▼
invalidate global SwitchPoint invalidate ClassInfo(ColdType) SP
│ │
▼ ▼
ALL sites: guard fails ColdType sites only
HotTarget.compute → full re-select HotTarget.compute → stays linked
(MH rebuild / re-guard / cold path) (single SwitchPoint check passes)
```
Cost breakdown:
1. **Avoided re-links** — After each global invalidation, hot sites re-run
`Selector`, reinstall guards, and rewrite `CallSite`s. At 100 churn events ×
many sites this dominates (parent’s 400+ ms/op).
2. **Avoided JIT churn** — Frequent `SwitchPoint.invalidateAll` breaks
monomorphic inlining assumptions and amplifies deopt/recompile cost (especially
under dense cross-type churn).
3. **Unchanged hot-path guard count** — Still one `guardWithTest`, so
baselines stay flat (H1).
4. **Same-type / hierarchy still correctly charged** — Prevents “fast but
wrong” semantics (H4 + 56 tests).
5. **Category remains bulk** — No second hot-path guard; category visibility
stays simple and correct; cost matches the old global path (measured ~1.0×).
Same-type still improves ~**7–8×** on HEAD vs parent, not because re-link is
skipped, but because the invalidation surface shrinks from “all process sites”
to “this class hierarchy” and batch `invalidateAll` is cheaper—a secondary
benefit, not the primary goal.
---
## 6. Risks and boundaries
| Topic | Assessment |
|---|---|
| Hot-path regression | **Not observed** (baselines 0.99–1.08×) |
| Semantic regression | 56/56 tests for this commit; Category / hierarchy /
per-instance MC covered |
| Dense Category enter/leave | **Still expensive** (by design);
Category-heavy workloads do not get faster |
| Non-final parent MetaClass change | Hierarchy scan over
`ClassInfo.getAllClassInfo()` is O(loaded classes); cold path,
correctness-first |
| Legacy `IndyInterface.switchPoint` | Rotated for binary compatibility
only; **no longer** the real call-site MOP guard; external readers should
migrate to `IndyInvalidation` |
| Environment noise | Shared/container CPUs; conclusions use
order-of-magnitude gaps (10¹–10²×), robust to ±10% noise |
| Not covered here | Concurrent multi-thread churn throughput, full Grails
end-to-end, JDK 17/21 matrix (recommended for CI) |
---
## 7. Hypothesis scorecard
| Hypothesis | Result | Evidence |
|---|---|---|
| H1 baseline parity | **Holds** | Scoped baseline 0.99×; CallSite baselines
1.03–1.08× |
| H2 cross-type speedup | **Holds** | ≈156× thrpt; ≈146× avgt (@1000) |
| H3 post-unrelated-burst full speed | **Holds** | `afterUnrelatedBurst` ≥
baseline on HEAD |
| H4 necessary invalidation still paid | **Holds** | same-type / hierarchy /
category ≪ baseline |
| H5 correctness | **Holds** | 56/56 tests |
---
## 8. Conclusions and recommendations
### Verdict
`ae8e497` (GROOVY-12191) keeps a **single-guard monomorphic hot path** while
scoping indy MOP invalidation to per-class domains (+ hierarchy). For
**unrelated-type MetaClass churn with hot monomorphic call sites**—a
Grails/framework-shaped load—measured improvement is on the order of
**10¹–10²×**. Steady-state undisturbed paths show **no material regression**.
Correctness costs for Category and same-type changes **remain visible and
intentional**.
**Performance verification: PASS.**
### Recommendations
1. Keep `ScopedInvalidationBench` and the cross-type / burst rows of
`CallSiteInvalidationBench` in regular JMH regression (README already documents
the suite) so a return to a global SwitchPoint cannot land silently.
2. In release notes, state clearly that Category and unscoped MetaClass
events may still bulk-invalidate; the win is **type-local MetaClass change**.
3. Optionally re-run CallSite `crossType@1000` and baselines on JDK 17/21 to
confirm cross-LTS consistency.
4. Document migration away from reading `IndyInterface.switchPoint` (field
is `@Deprecated`).
---
## Appendix A — Reproduction commands
```bash
# Correctness
./gradlew :test --tests Groovy12191 \
--tests org.apache.groovy.runtime.indy.IndyInvalidationTest \
--tests org.apache.groovy.runtime.indy.SwitchPointInvalidatorTest \
--tests org.codehaus.groovy.vmplugin.v8.IndyScopedSwitchPointTest
# Performance (HEAD)
./gradlew :perf:jmh -PbenchInclude=ScopedInvalidation -PjmhResultFormat=JSON
./gradlew :perf:jmh -PbenchInclude=CallSiteInvalidation
-PjmhResultFormat=JSON
# Performance (parent @ 0ca665dad0; mount the same ScopedInvalidationBench
sources for a fair A/B)
cd /path/to/parent && ./gradlew :perf:jmh -PbenchInclude=ScopedInvalidation
-PjmhResultFormat=JSON
cd /path/to/parent && ./gradlew :perf:jmh
-PbenchInclude=CallSiteInvalidation -PjmhResultFormat=JSON
```
## Appendix B — Files touched by the commit
- New: `org.apache.groovy.runtime.indy.IndyInvalidation`,
`SwitchPointInvalidator`
- Updated: `ClassInfo`, `IndyInterface`, `Selector`,
`ColdReflectiveMethodHandleWrapper`, `IndyCompoundAssign`
- Tests: `Groovy12191`, `IndyInvalidationTest`,
`SwitchPointInvalidatorTest`, `IndyScopedSwitchPointTest`
- Benchmarks: `ScopedInvalidationBench`; docs:
`subprojects/performance/README.adoc`
---
*Report generated from local JMH A/B measurements on 2026-07-26. Raw JSON
and run logs: `/tmp/groovy-12191-perf-results/`.*
> Scope indy SwitchPoint invalidation
> -----------------------------------
>
> Key: GROOVY-12191
> URL: https://issues.apache.org/jira/browse/GROOVY-12191
> Project: Groovy
> Issue Type: Improvement
> Reporter: Daniel Sun
> Priority: Major
>
> h3. Problem
> With invokedynamic enabled (default since Groovy 4), linked MOP call sites
> were guarded by a *single process-wide* {{SwitchPoint}}
> ({{{}IndyInterface.switchPoint{}}}).
> Any MetaClass registry change or category enter/leave invalidated that switch
> point, so *every* linked site fell back and re-linked — including sites whose
> receiver type was unrelated.
> That global invalidation is expensive when MetaClass churn is common (e.g.
> ExpandoMetaClass / mixins on startup or per-request paths, Grails-like
> patterns). Unrelated hot monomorphic sites pay re-link and JIT deopt cost
> they should not.
> h3. Goal
> Keep linked call sites optimized unless the *relevant* MetaClass state for
> that site actually changed.
> h3. Approach
> One SwitchPoint domain {*}per class{*}, stored on {{{}ClassInfo{}}}:
> * MetaClass change for type {{T}} retires {{{}T{}}}'s SwitchPoint *and*
> those of loaded subtypes / implementors (hierarchy fan-out).
> * Unrelated types keep their SwitchPoints; their call sites stay optimized.
> * Category enter/leave (and {{{}VMPlugin.invalidateCallSites(){}}})
> bulk-retire *all* loaded class SwitchPoints so sites re-link under the new
> category state. There is *no* second category SwitchPoint on the hot path.
> * Linked handles always install a *single* class-domain guard via
> {{IndyInvalidation.guardWithMopSwitchPoints(...)}} — same monomorphic guard
> shape as before, without global deopt on unrelated MetaClass churn.
> Final classes short-circuit hierarchy fan-out (no full {{ClassInfo}} scan).
> Non-final types batch retirements with {{{}SwitchPoint.invalidateAll{}}}.
> h3. Invalidation map
> ||Event||What is retired||
> |MetaClass change for type {{T}} (registry /
> {{{}ClassInfo.incVersion{}}})|{{T}} + loaded subtypes / implementors|
> |Category enter/leave, {{invalidateCallSites()}}|All loaded class
> SwitchPoints (bulk)|
> |Unattributed MetaClass registry event|All loaded class SwitchPoints (bulk)|
> |First MetaClass *install* on a class|Version bump only (no linked sites
> yet); replacement / clear retires that class's SwitchPoint|
>
--
This message was sent by Atlassian Jira
(v8.20.10#820010)