andygrove commented on code in PR #6498:
URL: https://github.com/apache/datafusion-comet/pull/6498#discussion_r4160584542


##########
native/spark-expr/src/array_funcs/list_extract.rs:
##########
@@ -359,6 +360,113 @@ fn list_extract<O: OffsetSizeTrait>(
     )))
 }
 
+// Inspect owned output capacity directly: nested builders can propagate a 
reservation
+// into deeper children even when the immediate child count was estimated 
correctly.
+// Keep ordinary growth headroom and small aligned buffers to avoid 
unnecessary copies.
+fn has_excess_nested_capacity(array: &dyn Array) -> bool {

Review Comment:
   `map_extract.rs` has the same problem, and this PR doesn't cover it. 
`spark_map_extract` ends with `take(map_array.values(), &indices, None)`, so a 
map whose values are arrays gets the same reservation from `take_list`. I 
probed 8192 rows of `map<int, array<int>>` holding `{1 -> [], 2 -> 128 ints}` 
and looked up key 1. The output retained 2,129,924 bytes, with 2 MiB in the 
empty child, which is the same number as #6225.
   
   Could we move these helpers into a shared module and use them in 
`map_extract` here too? If you'd rather keep this PR to `list_extract`, could 
you file a follow-up issue for the map path?



##########
spark/src/test/resources/sql-tests/expressions/array/list_extract_nested.sql:
##########
@@ -0,0 +1,73 @@
+-- Licensed to the Apache Software Foundation (ASF) under one
+-- or more contributor license agreements.  See the NOTICE file
+-- distributed with this work for additional information
+-- regarding copyright ownership.  The ASF licenses this file
+-- to you under the Apache License, Version 2.0 (the
+-- "License"); you may not use this file except in compliance
+-- with the License.  You may obtain a copy of the License at
+--
+--   http://www.apache.org/licenses/LICENSE-2.0
+--
+-- Unless required by applicable law or agreed to in writing,
+-- software distributed under the License is distributed on an
+-- "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+-- KIND, either express or implied.  See the License for the
+-- specific language governing permissions and limitations
+-- under the License.
+
+
+-- Config: spark.sql.ansi.enabled=false
+-- Config: spark.comet.batchSize=3
+
+statement
+CREATE TABLE test_list_extract_nested(a array<array<int>>, m 
array<map<int,int>>, s array<struct<n:array<int>>>, idx int) USING parquet
+
+statement
+INSERT INTO test_list_extract_nested VALUES

Review Comment:
   The fixture never reaches `compact_nested_data`. Every batch here holds at 
most three rows of short arrays, so the largest reservation `take` makes still 
fits in one 64-byte block and `has_excess_capacity` never fires. I checked with 
a temporary `eprintln!` on the compaction branch and it printed nothing. So 
these queries cover nested `GetArrayItem` and `ElementAt` in general, but no 
compacted array ever reaches the JVM.
   
   Could we add a row whose unselected element is long? These two rows were 
enough locally. With them the branch fired for all five value types and the 
file still matched Spark.
   
   ```sql
   -- test_list_extract_nested
   (array(array(), sequence(1, 64)), array(map(), map_from_arrays(sequence(1, 
64), sequence(1, 64))), array(named_struct('n', array()), named_struct('n', 
sequence(1, 64))), 0)
   -- test_list_extract_deep
   (array(array_repeat(array(), 32), array_repeat(sequence(1, 4), 32)), 
array(map_from_arrays(sequence(1, 32), array_repeat(array(), 32)), 
map_from_arrays(sequence(1, 32), array_repeat(sequence(1, 4), 32))), 0)
   ```



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to