[GitHub] [spark] viirya opened a new pull request #25352: [SPARK-28422][SQL][Python] GROUPED_AGG pandas_udf should work without group by clause

GitBox Sun, 04 Aug 2019 11:03:55 -0700

viirya opened a new pull request #25352: [SPARK-28422][SQL][Python] GROUPED_AGG 
pandas_udf should work without group by clause
URL: https://github.com/apache/spark/pull/25352
 
 
   ## What changes were proposed in this pull request?
   
   A GROUPED_AGG pandas python udf can't work, if without group by clause, like 
`select udf(id) from table`.
   
   This doesn't match with aggregate function like sum, count..., and also 
dataset API like `df.agg(udf(df['id']))`.
   
   When we parse a udf (or an aggregate function) like that from SQL syntax, it 
is known as a function in a project. `GlobalAggregates` rule in analysis makes 
such project as aggregate, by looking for aggregate expressions. At the moment, 
we should also look for GROUPED_AGG pandas python udf.
   
   ## How was this patch tested?
   
   Added tests.


----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
 
For queries about this service, please contact Infrastructure at:
[email protected]


With regards,
Apache Git Services

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

[GitHub] [spark] viirya opened a new pull request #25352: [SPARK-28422][SQL][Python] GROUPED_AGG pandas_udf should work without group by clause

Reply via email to