github-actions[bot] commented on code in PR #68819:
URL: https://github.com/apache/doris/pull/68819#discussion_r4227605341
##########
regression-test/suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy:
##########
@@ -105,7 +105,10 @@ suite("test_hdfs_parquet_group0", "p0,external") {
uri = "${defaultFS}" +
"/user/doris/tvf_data/test_hdfs_parquet/group0/large_string_map.brotli.parquet"
- order_qt_test_11 """ select count(arr) from HDFS(
+ // Read both 1 GiB keys one row per batch to avoid a 4 GiB output
buffer allocation.
+ // Disable aggregate pushdown to retain full decoding of the >2
GiB column chunk.
+ order_qt_test_11 """ select /*+ SET_VAR(batch_size=1,
enable_push_down_no_group_agg=false) */
Review Comment:
[P2] Honor the one-row cap when V2 adaptive batching is off. With the
default `enable_file_scanner_v2=true` and mutable BE
`enable_adaptive_batch_size=false`, V2 never calls
`TableReader::set_batch_size`: both ordinary-scan calls require an adaptive
predictor. `TableReader` stays at zero, so native Parquet keeps its 4096-row
default and can read both 1 GiB Map keys into one output batch, recreating the
4 GiB allocation this test is meant to avoid. Initialize V2's reader from the
query `batch_size` even when adaptive sizing is disabled, or pin this test to
V1.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]