ghaarsma opened a new issue, #50667:
URL: https://github.com/apache/arrow/issues/50667

   ### Describe the bug, including details regarding any error messages, 
version, and platform.
   
   In our application we implement a local TimeSeries caching implementation 
that stores TimeSeries data into Parquet files. We heavily read/write parquet 
files via the pandas to_parquet(engine="pyarrow") and read_parquet. We also use 
the pyarrow.read_metadata directly from the pyarrows package. We use default 
(snappy) pandas compression when we write out our files, which consist of a 
single column DataFrame with the index being the timestamps (DateTimeIndex) and 
the single column representing the TimeSeries values (mostly Float32/64).
   
   After installing pyarrow 25.0.0 we saw a significant reduction in 
performance (see trend below). When we downgraded back to 24.0.0 the problem 
went away. No other changes were made and we kept the rest of our dependencies 
the same.
   
   <img width="1923" height="889" alt="Image" 
src="https://github.com/user-attachments/assets/34101978-32b4-44b1-8dc0-924643466158";
 />
   
   The Dev/Prod stack Azure App Service on `Linux 779de6246efc 
6.6.139.1-1.azl3`:
   ```
   root@779de6246efc:/home/site/wwwroot# python
   Python 3.13.13 (main, May  4 2026, 08:38:58) [GCC 12.2.0] on linux
   Type "help", "copyright", "credits" or "license" for more information.
   >>> import pyarrow as pa
   >>> pa.show_info()
   pyarrow version info
   --------------------
   Package kind              : python-wheel-manylinux228
   Arrow C++ library version : 24.0.0
   Arrow C++ compiler        : GNU 14.2.1
   Arrow C++ compiler flags  :  -Wno-noexcept-type -Wno-self-move 
-Wno-subobject-linkage  -fdiagnostics-color=always  -Wall 
-fno-semantic-interposition -msse4.2
   Arrow C++ git revision    :
   Arrow C++ git description :
   Arrow C++ build type      : release
   PyArrow build type        : release
   
   Platform:
     OS / Arch           : Linux x86_64
     SIMD Level          : avx2
     Detected SIMD Level : avx2
   
   Memory:
     Default backend     : mimalloc
     Bytes allocated     : 0 bytes
     Max memory          : 0 bytes
     Supported Backends  : mimalloc, jemalloc, system
   
   Optional modules:
     csv                 : Enabled
     cuda                : -
     dataset             : Enabled
     feather             : Enabled
     flight              : Enabled
     fs                  : Enabled
     gandiva             : -
     json                : Enabled
     orc                 : Enabled
     parquet             : Enabled
     opentelemetry       : Enabled
   
   Filesystems:
     AzureFileSystem     : Enabled
     GcsFileSystem       : Enabled
     HadoopFileSystem    : Enabled
     S3FileSystem        : Enabled
   
   Compression Codecs:
     brotli              : Enabled
     bz2                 : Enabled
     gzip                : Enabled
     lz4_frame           : Enabled
     lz4                 : Enabled
     snappy              : Enabled
     zstd                : Enabled
   ```
   
   
   ### Component(s)
   
   C++, Python


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to