SulimanAbdulrazzaq commented on code in PR #44652:
URL: https://github.com/apache/superset/pull/44652#discussion_r4177236711
##########
superset/utils/pandas_postprocessing/histogram.py:
##########
@@ -107,9 +106,14 @@ def hist_values(series: Series) -> np.ndarray:
)
histogram_df.columns = bin_edges_str
+ # normalize by the total count of data points before accumulating, so that
a
+ # cumulative histogram shows the share of data points up to each bin
if normalize:
histogram_df = histogram_df / histogram_df.values.sum()
Review Comment:
Rows whose group value is null are not in `histogram_df` at all, because
`groupby` drops them, so they are not shown and the denominator is the sum of
the rows that are. That is how `normalize` worked before this PR; the change
only normalizes before accumulating instead of after.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]