Nandor Kollar has posted comments on this change. ( http://gerrit.cloudera.org:8080/24733 )
Change subject: IMPALA-14996: Optimize count(*) for Iceberg V3 DV-only tables ...................................................................... Patch Set 1: (2 comments) http://gerrit.cloudera.org:8080/#/c/24733/1/fe/src/main/java/org/apache/impala/catalog/FeIcebergTable.java File fe/src/main/java/org/apache/impala/catalog/FeIcebergTable.java: http://gerrit.cloudera.org:8080/#/c/24733/1/fe/src/main/java/org/apache/impala/catalog/FeIcebergTable.java@1237 PS1, Line 1237: long dataRecords = fileStore.getDataFilesWithoutDeletes().stream() You can combine the two iterations (fileStore.getDataFilesWithoutDeletes and fileStore.getDataFilesWithDeletes) into a single one, if you use fileStore.getAllDataFiles() Actually, can't we just skip iterating over the data files? Can't we say that: v1 record count (read from the snapshot, no iteration at all) - deletedRecords (we need to iterate over the delete vectors here)? http://gerrit.cloudera.org:8080/#/c/24733/1/fe/src/main/java/org/apache/impala/catalog/FeIcebergTable.java@1243 PS1, Line 1243: if (fileStore.getDataFileToDV().values().stream() nit: Here we iterate through the getDataFileToDV twice (first over the values, then over the entrySet), I think we could combine this to a single loop. -- To view, visit http://gerrit.cloudera.org:8080/24733 To unsubscribe, visit http://gerrit.cloudera.org:8080/settings Gerrit-Project: Impala-ASF Gerrit-Branch: master Gerrit-MessageType: comment Gerrit-Change-Id: I0e6239629ee779a64808a0a4bbe041ef756c5bb2 Gerrit-Change-Number: 24733 Gerrit-PatchSet: 1 Gerrit-Owner: Arnab Karmakar <[email protected]> Gerrit-Reviewer: Impala Public Jenkins <[email protected]> Gerrit-Reviewer: Nandor Kollar <[email protected]> Gerrit-Reviewer: Peter Rozsa <[email protected]> Gerrit-Comment-Date: Thu, 27 Aug 2026 14:14:19 +0000 Gerrit-HasComments: Yes
