bito-code-review[bot] commented on PR #44652:
URL: https://github.com/apache/superset/pull/44652#issuecomment-5832020772
<!-- Bito Reply -->
The flagged issue is correct. When grouping by columns that contain null
values, pandas drops those rows by default, which causes the denominator in the
normalization step (`histogram_df.values.sum()`) to exclude those data points.
As a result, the normalized cumulative values do not correctly represent the
share of the total dataset.
To resolve this, you should ensure that null values in the groupby columns
are preserved, for example by using `observed=False` (if using `groupby`) or by
explicitly filling nulls before grouping.
Would you like me to fetch all other comments on this PR to validate and
implement fixes for them as well?
**superset/utils/pandas_postprocessing/histogram.py**
```
# Ensure nulls in groupby columns are not dropped
histogram_df = df.groupby(groupby,
observed=False)[column].apply(hist_values).unstack()
```
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]