[
https://issues.apache.org/jira/browse/ARROW-16598?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17538464#comment-17538464
]
Will Jones edited comment on ARROW-16598 at 5/17/22 8:50 PM:
-------------------------------------------------------------
Yeah. I say column encodings are likely the biggest influence, but columns are
also run through compression algorithms like Snappy, LZ4, and GZIP, and so I
bet there are surprising interactions there. That's why it's hard to give
generic advice.
I don't think we're setup to do this now, but it makes me wonder if we couldn't
make something that does for parquet what
[https://www.squoosh.app|https://www.squoosh.app/] does for images: provide an
easy interface to try out different compression, sort, and partitioning options
and see how they impact file size / access patterns. Could be prototyped as a
Shiny app, but would be cool eventually to have as a WASM in-browser app like
Squoosh.
was (Author: willjones127):
Yeah. I say column encodings are likely the biggest influence, but columns are
also run through compression algorithms like Snappy, LZ4, and GZIP, and so I
bet there are surprising interactions there. That's why it's hard to give
generic advice.
I don't think we're setup to do this now, but it makes me wonder if we couldn't
make something that does for parquet what
[https://www.squoosh.app|https://www.squoosh.app/] does for images: provide an
easy interface to try out different compression, sort, and partitioning options
and see how they impact access patterns. Could be prototyped as a Shiny app,
but would be cool eventually to have as a WASM in-browser app like Squoosh.
> [R] Sorting data.frame prior to writing Parquet affects file size
> -----------------------------------------------------------------
>
> Key: ARROW-16598
> URL: https://issues.apache.org/jira/browse/ARROW-16598
> Project: Apache Arrow
> Issue Type: Bug
> Components: R
> Environment: MacBook Pro (non-M1), other info in R file
> Reporter: Michael Culshaw-Maurer
> Priority: Minor
> Attachments: arrow_parquet_bug.R
>
>
> When using the arrow R package, sorting a data.frame prior to using
> write_parquet() results in different file sizes, depending on how the
> data.frame is sorted. I have attached a reproducible example showing how a
> few different sorting methods can lead to 2-3 fold changes in .parquet file
> size.
> It may be that I don't know enough about Parquet internals, but at the very
> least, I think this behavior should be documented on the arrow R package
> site. Most R users tend to approach sorting as a convenience and don't expect
> it to lead to performance changes when writing to a file.
> {code:java}
> library(tidyverse)
> d <- expand_grid(group = letters[1:4],
> id = 1:100) %>%
> mutate(id_f = paste(group, id, sep = "_")) %>%
> mutate(time = rep(list(1:100)), 400) %>%
> unnest(time) %>%
> group_by(group) %>%
> mutate(id_n = list(sample(id_f, size = 5, replace = F))) %>%
> unnest(id_n) %>%
> ungroup()
> f1 <- tempfile(fileext = ".parquet")
> f2 <- tempfile(fileext = ".parquet")
> f3 <- tempfile(fileext = ".parquet")
> f4 <- tempfile(fileext = ".parquet")
> f5 <- tempfile(fileext = ".parquet")
> d %>%
> arrow::write_parquet(f1)
> d %>%
> arrange(id_n) %>%
> arrow::write_parquet(f2)
> d %>%
> arrange(id_n, time) %>%
> arrow::write_parquet(f3)
> d %>%
> arrange(time, id_f) %>%
> arrow::write_parquet(f4)
> d %>%
> arrange(group, time, id_n, id_f) %>%
> arrow::write_parquet(f5)
> fs::file_info(c(f1, f2, f3, f4, f5))[, "size"]
> #> # A tibble: 5 × 1
> #> size
> #> <fs::bytes>
> #> 1 25.4K
> #> 2 17.3K
> #> 3 28.4K
> #> 4 45.6K
> #> 5 30.1K
> sessioninfo::session_info()
> #> ─ Session info
> ───────────────────────────────────────────────────────────────
> #> setting value
> #> version R version 4.1.3 (2022-03-10)
> #> os macOS Big Sur/Monterey 10.16
> #> system x86_64, darwin17.0
> #> ui X11
> #> language (EN)
> #> collate en_US.UTF-8
> #> ctype en_US.UTF-8
> #> tz America/Chicago
> #> date 2022-05-17
> #>
> #> ─ Packages
> ───────────────────────────────────────────────────────────────────
> #> package * version date lib source
> #> arrow 8.0.0 2022-05-09 [1] CRAN (R 4.1.2)
> #> assertthat 0.2.1 2019-03-21 [1] CRAN (R 4.1.0)
> #> backports 1.4.1 2021-12-13 [1] CRAN (R 4.1.0)
> #> bit 4.0.4 2020-08-04 [1] CRAN (R 4.1.0)
> #> bit64 4.0.5 2020-08-30 [1] CRAN (R 4.1.0)
> #> broom 0.7.9 2021-07-27 [1] CRAN (R 4.1.0)
> #> cellranger 1.1.0 2016-07-27 [1] CRAN (R 4.1.0)
> #> cli 3.3.0 2022-04-25 [1] CRAN (R 4.1.2)
> #> colorspace 2.0-3 2022-02-21 [1] CRAN (R 4.1.2)
> #> crayon 1.5.1 2022-03-26 [1] CRAN (R 4.1.2)
> #> DBI 1.1.1 2021-01-15 [1] CRAN (R 4.1.0)
> #> dbplyr 2.1.1 2021-04-06 [1] CRAN (R 4.1.0)
> #> digest 0.6.29 2021-12-01 [1] CRAN (R 4.1.0)
> #> dplyr * 1.0.9 2022-04-28 [1] CRAN (R 4.1.2)
> #> ellipsis 0.3.2 2021-04-29 [1] CRAN (R 4.1.0)
> #> evaluate 0.14 2019-05-28 [1] CRAN (R 4.1.0)
> #> fansi 1.0.3 2022-03-24 [1] CRAN (R 4.1.2)
> #> fastmap 1.1.0 2021-01-25 [1] CRAN (R 4.1.0)
> #> forcats * 0.5.1 2021-01-27 [1] CRAN (R 4.1.0)
> #> fs 1.5.2 2021-12-08 [1] CRAN (R 4.1.0)
> #> generics 0.1.2 2022-01-31 [1] CRAN (R 4.1.2)
> #> ggplot2 * 3.3.6 2022-05-03 [1] CRAN (R 4.1.3)
> #> glue 1.6.2 2022-02-24 [1] CRAN (R 4.1.2)
> #> gtable 0.3.0 2019-03-25 [1] CRAN (R 4.1.0)
> #> haven 2.4.3 2021-08-04 [1] CRAN (R 4.1.0)
> #> highr 0.9 2021-04-16 [1] CRAN (R 4.1.0)
> #> hms 1.1.0 2021-05-17 [1] CRAN (R 4.1.0)
> #> htmltools 0.5.2 2021-08-25 [1] CRAN (R 4.1.0)
> #> httr 1.4.2 2020-07-20 [1] CRAN (R 4.1.0)
> #> jsonlite 1.8.0 2022-02-22 [1] CRAN (R 4.1.2)
> #> knitr 1.37 2021-12-16 [1] CRAN (R 4.1.0)
> #> lifecycle 1.0.1 2021-09-24 [1] CRAN (R 4.1.0)
> #> lubridate 1.7.10 2021-02-26 [1] CRAN (R 4.1.0)
> #> magrittr 2.0.3 2022-03-30 [1] CRAN (R 4.1.2)
> #> modelr 0.1.8 2020-05-19 [1] CRAN (R 4.1.0)
> #> munsell 0.5.0 2018-06-12 [1] CRAN (R 4.1.0)
> #> pillar 1.7.0 2022-02-01 [1] CRAN (R 4.1.2)
> #> pkgconfig 2.0.3 2019-09-22 [1] CRAN (R 4.1.0)
> #> purrr * 0.3.4 2020-04-17 [1] CRAN (R 4.1.0)
> #> R6 2.5.1 2021-08-19 [1] CRAN (R 4.1.0)
> #> Rcpp 1.0.8.3 2022-03-17 [1] CRAN (R 4.1.2)
> #> readr * 2.0.1 2021-08-10 [1] CRAN (R 4.1.0)
> #> readxl 1.3.1 2019-03-13 [1] CRAN (R 4.1.0)
> #> reprex 2.0.1 2021-08-05 [1] CRAN (R 4.1.0)
> #> rlang 1.0.2 2022-03-04 [1] CRAN (R 4.1.2)
> #> rmarkdown 2.11 2021-09-14 [1] CRAN (R 4.1.0)
> #> rstudioapi 0.13 2020-11-12 [1] CRAN (R 4.1.0)
> #> rvest 1.0.1 2021-07-26 [1] CRAN (R 4.1.0)
> #> scales 1.2.0 2022-04-13 [1] CRAN (R 4.1.2)
> #> sessioninfo 1.1.1 2018-11-05 [1] CRAN (R 4.1.0)
> #> stringi 1.7.6 2021-11-29 [1] CRAN (R 4.1.0)
> #> stringr * 1.4.0 2019-02-10 [1] CRAN (R 4.1.0)
> #> styler 1.4.1 2021-03-30 [1] CRAN (R 4.1.0)
> #> tibble * 3.1.7 2022-05-03 [1] CRAN (R 4.1.3)
> #> tidyr * 1.1.3 2021-03-03 [1] CRAN (R 4.1.0)
> #> tidyselect 1.1.2 2022-02-21 [1] CRAN (R 4.1.2)
> #> tidyverse * 1.3.1 2021-04-15 [1] CRAN (R 4.1.0)
> #> tzdb 0.1.2 2021-07-20 [1] CRAN (R 4.1.0)
> #> utf8 1.2.2 2021-07-24 [1] CRAN (R 4.1.0)
> #> vctrs 0.4.1 2022-04-13 [1] CRAN (R 4.1.2)
> #> withr 2.5.0 2022-03-03 [1] CRAN (R 4.1.2)
> #> xfun 0.30 2022-03-02 [1] CRAN (R 4.1.2)
> #> xml2 1.3.2 2020-04-23 [1] CRAN (R 4.1.0)
> #> yaml 2.3.5 2022-02-21 [1] CRAN (R 4.1.2)
> #>
> #> [1] /Users/MJ/R_Packages_4.1
> #> [2] /Library/Frameworks/R.framework/Versions/4.1/Resources/library {code}
--
This message was sent by Atlassian Jira
(v8.20.7#820007)