adriangb opened a new issue, #10929:
URL: https://github.com/apache/arrow-rs/issues/10929

   ## Describe the bug
   
   With content-defined chunking (CDC) enabled, writing a `BOOLEAN` column 
panics with `RLE value encoder is not initialized`.
   
   CDC forces a data page break at the end of every chunk except the last:
   
   ```rust
   // parquet/src/arrow/arrow_writer/mod.rs, 
ArrowColumnWriter::write_with_chunker
   // Add a page break after each chunk except the last
   if i + 1 < num_chunks {
       match &mut self.writer {
           ArrowColumnWriterImpl::Column(c) => c.add_data_page()?,
           ArrowColumnWriterImpl::ByteArray(c) => c.add_data_page()?,
       }
   }
   ```
   
   That break is unconditional. Writing the chunk can itself have flushed the 
page already, when the chunk's own values reach `data_page_size_limit` or 
`data_page_row_count_limit` exactly at the chunk boundary. The forced break 
then flushes a page with no buffered values.
   
   `RleValueEncoder` builds its inner encoder lazily on the first `put`, so 
flushing before any value has been written panics:
   
   ```rust
   // parquet/src/encodings/encoding/mod.rs
   let rle_encoder = self
       .encoder
       .take()
       .expect("RLE value encoder is not initialized");
   ```
   
   A `BOOLEAN` column uses `RleValueEncoder` under `WriterVersion::PARQUET_2_0` 
(`fallback_encoding`), or when `Encoding::RLE` is set explicitly.
   
   For every other encoding the same forced break does not panic, but writes a 
data page holding zero values.
   
   Note that `data_page_size_limit` being smaller than `max_chunk_size` is a 
documented, supported configuration. From the `CdcOptions::max_chunk_size` docs:
   
   > Note that the parquet writer has a related `data_page_size_limit` property 
that controls the maximum size of a parquet data page after encoding. While 
setting `data_page_size_limit` to a smaller value than `max_chunk_size` doesn't 
affect the chunking effectiveness, it results in more small parquet data pages.
   
   **Affects:** `parquet` 59.2.0, and any release with content-defined chunking.
   
   ## To Reproduce
   
   `Cargo.toml`:
   
   ```toml
   [dependencies]
   arrow = "59.2.0"
   parquet = "59.2.0"
   ```
   
   `src/main.rs`:
   
   ```rust
   use arrow::array::{ArrayRef, BooleanArray, RecordBatch};
   use parquet::arrow::ArrowWriter;
   use parquet::file::properties::{CdcOptions, WriterProperties, WriterVersion};
   use std::sync::Arc;
   
   fn main() {
       let values: Vec<bool> = (0..500_000).map(|i| i % 7 == 0).collect();
       let col = Arc::new(BooleanArray::from(values)) as ArrayRef;
       let batch = RecordBatch::try_from_iter([("flag", col)]).unwrap();
   
       let props = WriterProperties::builder()
           .set_writer_version(WriterVersion::PARQUET_2_0)
           .set_data_page_size_limit(1024)
           .set_content_defined_chunking(Some(CdcOptions {
               min_chunk_size: 8 * 1024,
               max_chunk_size: 16 * 1024,
               norm_level: 0,
           }))
           .build();
   
       let mut out = Vec::new();
       let mut writer = ArrowWriter::try_new(&mut out, batch.schema(), 
Some(props)).unwrap();
       writer.write(&batch).unwrap();
       writer.close().unwrap();
       println!("wrote {} bytes", out.len());
   }
   ```
   
   ```console
   $ cargo run --release
   thread 'main' panicked at 
parquet-59.2.0/src/encodings/encoding/mod.rs:250:14:
   RLE value encoder is not initialized
   ```
   
   The zero-value pages written for other types are visible by swapping in an 
`Int32Array` and counting pages:
   
   ```rust
   let props = WriterProperties::builder()
       .set_writer_version(WriterVersion::PARQUET_2_0)
       .set_dictionary_enabled(false)
       .set_data_page_row_count_limit(128)
       .set_content_defined_chunking(Some(CdcOptions {
           min_chunk_size: 8 * 1024,
           max_chunk_size: 16 * 1024,
           norm_level: 0,
       }))
       .build();
   // ... write 500_000 values of `(0..500_000).map(|i| i % 97)` into column 
"a", then:
   // iterate the column's pages and count those with `page.num_values() == 0`
   ```
   
   ```console
   total data pages = 611, pages with zero values = 121
   ```
   
   ## Expected behavior
   
   Writing the column succeeds, and no data page holds zero values. A forced 
page break with nothing buffered should be a no-op.
   
   ## Additional context
   
   The two neighbouring call sites in `parquet/src/column/writer/mod.rs` 
already guard this condition, so `add_data_page` is the odd one out:
   
   - `should_add_data_page` returns `false` when 
`page_metrics.num_buffered_values == 0`
   - `dict_fallback` and `flush_data_pages` check 
`page_metrics.num_buffered_values > 0` before calling `add_data_page`
   
   ---
   
   *This issue was written by Claude (Anthropic's AI assistant) working with 
@adriangb. The reproduction above was executed and its output is verbatim.*
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to