Nandor Kollar has posted comments on this change. ( 
http://gerrit.cloudera.org:8080/24733 )

Change subject: IMPALA-14996: Optimize count(*) for Iceberg V3 DV-only tables
......................................................................


Patch Set 1:

(2 comments)

http://gerrit.cloudera.org:8080/#/c/24733/1/fe/src/main/java/org/apache/impala/catalog/FeIcebergTable.java
File fe/src/main/java/org/apache/impala/catalog/FeIcebergTable.java:

http://gerrit.cloudera.org:8080/#/c/24733/1/fe/src/main/java/org/apache/impala/catalog/FeIcebergTable.java@1237
PS1, Line 1237:         long dataRecords = 
fileStore.getDataFilesWithoutDeletes().stream()
You can combine the two iterations (fileStore.getDataFilesWithoutDeletes and 
fileStore.getDataFilesWithDeletes) into a single one, if you use 
fileStore.getAllDataFiles()

Actually, can't we just skip iterating over the data files? Can't we say that: 
v1 record count (read from the snapshot, no iteration at all) - deletedRecords 
(we need to iterate over the delete vectors here)?


http://gerrit.cloudera.org:8080/#/c/24733/1/fe/src/main/java/org/apache/impala/catalog/FeIcebergTable.java@1243
PS1, Line 1243:         if (fileStore.getDataFileToDV().values().stream()
nit: Here we iterate through the getDataFileToDV twice (first over the values, 
then over the entrySet), I think we could combine this to a single loop.



--
To view, visit http://gerrit.cloudera.org:8080/24733
To unsubscribe, visit http://gerrit.cloudera.org:8080/settings

Gerrit-Project: Impala-ASF
Gerrit-Branch: master
Gerrit-MessageType: comment
Gerrit-Change-Id: I0e6239629ee779a64808a0a4bbe041ef756c5bb2
Gerrit-Change-Number: 24733
Gerrit-PatchSet: 1
Gerrit-Owner: Arnab Karmakar <[email protected]>
Gerrit-Reviewer: Impala Public Jenkins <[email protected]>
Gerrit-Reviewer: Nandor Kollar <[email protected]>
Gerrit-Reviewer: Peter Rozsa <[email protected]>
Gerrit-Comment-Date: Thu, 27 Aug 2026 14:14:19 +0000
Gerrit-HasComments: Yes

Reply via email to