SulimanAbdulrazzaq commented on code in PR #44652:
URL: https://github.com/apache/superset/pull/44652#discussion_r4177236711


##########
superset/utils/pandas_postprocessing/histogram.py:
##########
@@ -107,9 +106,14 @@ def hist_values(series: Series) -> np.ndarray:
         )
         histogram_df.columns = bin_edges_str
 
+    # normalize by the total count of data points before accumulating, so that 
a
+    # cumulative histogram shows the share of data points up to each bin
     if normalize:
         histogram_df = histogram_df / histogram_df.values.sum()

Review Comment:
   Rows whose group value is null are not in `histogram_df` at all, because 
`groupby` drops them, so they are not shown and the denominator is the sum of 
the rows that are. That is how `normalize` worked before this PR; the change 
only normalizes before accumulating instead of after.
   



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to