[ 
https://issues.apache.org/jira/browse/SPARK-59376?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated SPARK-59376:
-----------------------------------
    Labels: pull-request-available  (was: )

> Add frequency encoding to ml.feature
> ------------------------------------
>
>                 Key: SPARK-59376
>                 URL: https://issues.apache.org/jira/browse/SPARK-59376
>             Project: Spark
>          Issue Type: New Feature
>          Components: ML
>    Affects Versions: 5.0.0
>            Reporter: Varun Bhandary
>            Priority: Major
>              Labels: pull-request-available
>
> MLlib has no unsupervised encoder for high cardinality categorical features.
> StringIndexer produces ordinals whose magnitude is meaningless to any model 
> that
> treats its features as numbers. OneHotEncoder is correct but adds a column per
> category, which stops being practical after a few hundred. FeatureHasher 
> avoids
> the width at the cost of collisions and interpretability. TargetEncoder, 
> added in
> 4.0.0 by SPARK-37178, needs a label column, so it is not available for
> unsupervised work at all.
> That leaves clustering, anomaly detection and dimensionality reduction with no
> good option for a column of, say, 50,000 merchant ids.
> Frequency encoding fills that gap. Each category is replaced by how often it
> occurs in the training data, as a proportion or as a raw count. It is one of 
> the
> standard treatments for high cardinality categoricals and it needs no label.
> Proposed API, mirroring TargetEncoder so the two stay consistent:
>   FrequencyEncoder, an Estimator, producing FrequencyEncoderModel
>   inputCol / outputCol and inputCols / outputCols, pairs treated independently
>   handleInvalid, error or keep, where keep maps unseen categories to 0
>   normalize, default true, giving a proportion, or false for the raw count
>   numeric already indexed input, as TargetEncoder requires, so StringIndexer
>     feeds it
> Fit is a single pass. The input columns are packed into an array and 
> posexploded
> so that one groupBy aggregates every column at once, which is the shape
> SPARK-50267 gave TargetEncoder.fit. The fitted map is small, one entry per
> category, and transform applies it with try_element_at against a literal, so
> there is no join and no shuffle at transform time.
> It is fair to point out that a user can write groupBy.count and a join for
> themselves. The same is true of TargetEncoder, and in both cases the value is 
> not
> the arithmetic. It is the fitted model semantics: the mapping is learned once 
> and
> applied identically at train and at serving time, it survives save and load, 
> it
> composes inside a Pipeline, and unseen categories get defined behaviour 
> instead
> of silently becoming null.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to