zhengruifeng opened a new pull request, #57932:
URL: https://github.com/apache/spark/pull/57932

   ### What changes were proposed in this pull request?
   
   This PR updates the ML `KMeansModel`, `BisectingKMeansModel`, and 
`GaussianMixtureModel` size estimates to pass only their persisted MLlib data 
fields to `SizeEstimator`.
   
   For Gaussian mixture models, it estimates each Gaussian mean and covariance 
rather than the `MultivariateGaussian` wrapper. For bisecting K-means, it makes 
the MLlib model root available to Spark-internal code so the ML wrapper can 
list the fields directly.
   
   ### Why are the changes needed?
   
   `SizeEstimator` recursively follows object references. Estimating an entire 
parent MLlib model can include transient or lazily initialized derived state, 
such as distance structures, statistics, or logging fields, rather than only 
persisted model data. Explicit fields make the accounting predictable and avoid 
retaining those auxiliary object graphs during estimation.
   
   ### Does this PR introduce _any_ user-facing change?
   
   No.
   
   ### How was this patch tested?
   
   ```
   build/sbt -java-home /usr/lib/jvm/java-17-openjdk-amd64 'mllib / Compile / 
compile'
   ```
   
   ### Was this patch authored or co-authored using generative AI tooling?
   
   Generated-by: Codex (GPT-5)


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to