[GitHub] [arrow] ianmcook commented on a change in pull request #11425: ARROW-14304: [R] Update news for 6.0.0

GitBox Tue, 19 Oct 2021 06:47:44 -0700


ianmcook commented on a change in pull request #11425:
URL: https://github.com/apache/arrow/pull/11425#discussion_r731885443




##########
File path: r/NEWS.md
##########
@@ -21,32 +21,64 @@
 
 There are now two ways to query Arrow data:
 
-## 1. Grouped aggregation in Arrow
+## 1. Expanded Arrow-native queries: aggregation and joins
 
 `dplyr::summarize()`, both grouped and ungrouped, is now implemented for Arrow 
Datasets, Tables, and RecordBatches. Because data is scanned in chunks, you can 
aggregate over larger-than-memory datasets backed by many files. Supported 
aggregation functions include `n()`, `n_distinct()`, `min(),` `max()`, `sum()`, 
`mean()`, `var()`, `sd()`, `any()`, and `all()`. `median()` and `quantile()` 
with one probability are also supported and currently return approximate 
results using the t-digest algorithm.
 
+Along with `summarize()`, you can also call `count()`, `tally()`, and 
`distinct()`, which effectively wrap `summarize()`.
+
 This enhancement does change the behavior of `summarize()` and `collect()` in 
some cases: see "Breaking changes" below for details.
 
-New compute functions include `str_to_title()` and `strftime()`.
+In addition to `summarize()`, equality joins (`left_join()`, `inner_join()`, 
`semi_join()`, et al.) are also supported natively in Arrow.
+
+Grouped aggregation and (especially) joins should be considered somewhat 
experimental in this release. We expect them to work, but they may not be well 
optimized for all workloads. To help us focus our efforts on improving them in 
the next release, please let us know if you encounter unexpected behavior or 
poor performance.
+
+New non-aggregating compute functions include string functions like 
`str_to_title()` and `strftime()` as well as compute functions for extracting 
date parts (e.g. `year()`, `month()`) from dates. This is not a complete list 
of additional compute functions; for an exhaustive list of available compute 
functions see `list_compute_functions()`.
+
+We've also worked to fill in support for all data types, such as `Decimal`, 
for functions added in previous releases. All type limitations mentioned in 
previous release notes should be no longer valid, and if you find a function 
that is not implemented for a certain data type, please [report an 
issue](https://issues.apache.org/jira/projects/ARROW/issues).
 
 ## 2. duckdb integration
 
-If you have the [duckdb](https://duckdb.org/) package installed, you can hand 
off an Arrow Dataset or query object to duckdb for further querying using the 
`to_duckdb()` function. This allows you to use duckdb's `dbplyr` methods, as 
well as its SQL interface, to aggregate data. Filtering and column projection 
done before `to_duckdb()` is evaluated in Arrow.
+If you have the [duckdb](https://duckdb.org/) package installed, you can hand 
off an Arrow Dataset or query object to duckdb for further querying using the 
`to_duckdb()` function. This allows you to use duckdb's `dbplyr` methods, as 
well as its SQL interface, to aggregate data. Filtering and column projection 
done before `to_duckdb()` is evaluated in Arrow, and duckdb can push down some 
predicates to Arrow as well. This handoff *does not* copy the data, instead it 
uses Arrow's C-interface (just like passing arrow data between R and Python). 
This means there is no serialization or data copying costs are incurred.

Review comment:
       Distinguish between the database (DuckDB) and the R package (duckdb):
   ```suggestion
   If you have the [duckdb package](https://CRAN.R-project.org/package=duckdb) 
installed, you can hand off an Arrow Dataset or query object to 
[DuckDB](https://duckdb.org/) for further querying using the `to_duckdb()` 
function. This allows you to use duckdb's `dbplyr` methods, as well as its SQL 
interface, to aggregate data. Filtering and column projection done before 
`to_duckdb()` is evaluated in Arrow, and duckdb can push down some predicates 
to Arrow as well. This handoff *does not* copy the data, instead it uses 
Arrow's C-interface (just like passing arrow data between R and Python). This 
means there is no serialization or data copying costs are incurred.
   ```




-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

[GitHub] [arrow] ianmcook commented on a change in pull request #11425: ARROW-14304: [R] Update news for 6.0.0

Reply via email to