github-actions[bot] commented on code in PR #68819:
URL: https://github.com/apache/doris/pull/68819#discussion_r4227605341


##########
regression-test/suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy:
##########
@@ -105,7 +105,10 @@ suite("test_hdfs_parquet_group0", "p0,external") {
 
 
             uri = "${defaultFS}" + 
"/user/doris/tvf_data/test_hdfs_parquet/group0/large_string_map.brotli.parquet"
-            order_qt_test_11 """ select count(arr) from HDFS(
+            // Read both 1 GiB keys one row per batch to avoid a 4 GiB output 
buffer allocation.
+            // Disable aggregate pushdown to retain full decoding of the >2 
GiB column chunk.
+            order_qt_test_11 """ select /*+ SET_VAR(batch_size=1, 
enable_push_down_no_group_agg=false) */

Review Comment:
   [P2] Honor the one-row cap when V2 adaptive batching is off. With the 
default `enable_file_scanner_v2=true` and mutable BE 
`enable_adaptive_batch_size=false`, V2 never calls 
`TableReader::set_batch_size`: both ordinary-scan calls require an adaptive 
predictor. `TableReader` stays at zero, so native Parquet keeps its 4096-row 
default and can read both 1 GiB Map keys into one output batch, recreating the 
4 GiB allocation this test is meant to avoid. Initialize V2's reader from the 
query `batch_size` even when adaptive sizing is disabled, or pin this test to 
V1.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to