[
https://issues.apache.org/jira/browse/HIVE-30064?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Ryu Kobayashi updated HIVE-30064:
---------------------------------
Description:
h2. Summary
Copy-on-write UPDATE and MERGE operations can report an incorrect number of
affected rows.
The FileSink currently counts all rows written by the rewritten query. However,
copy-on-write rewrites entire data files, so the output also contains unchanged
survivor rows. MERGE additionally combines UPDATE, INSERT, DELETE, and survivor
branches into a shared output.
h2. Root cause
The affected-row count is derived from the number of rows written by the
FileSink rather than from the logical operation that produced each row.
As a result:
- CoW UPDATE can include unchanged rows from rewritten files.
- CoW MERGE can include INSERT rows and unchanged survivor rows in the same
output.
- The reported `numModifiedRows` value does not necessarily represent the rows
actually updated by the statement.
h2. Suggested change
Add a `cow_update_matched` marker to the projections generated by the CoW
UPDATE and MERGE rewriters.
- Set the marker to `true` for rows produced by the matched UPDATE branch.
- Set the marker to `false` for unchanged survivor rows and other branches
that should not contribute to the UPDATE count.
- Make `FileSinkOperator` count only rows whose marker is `true` when the
marker mode is enabled.
- Keep the marker as an internal column so it is not written to the target
table.
This avoids relying on the physical Tez plan shape or operator layout to infer
the affected-row count.
h2. Confirmed behavior
The change should be covered by regression tests for:
- CoW UPDATE with matching rows
- CoW UPDATE with no matching rows
- CoW MERGE with UPDATE, DELETE, and INSERT branches
- CoW MERGE with only INSERT or DELETE branches
- Partitioned and unpartitioned tables
The tests should verify both the reported affected-row count and the final
table contents.
was:
h2. Summary
Copy-on-write UPDATE and MERGE operations can report an incorrect number of
affected rows.
The FileSink currently counts all rows written by the rewritten query. However,
copy-on-write rewrites entire data files, so the output also contains unchanged
survivor rows. MERGE additionally combines UPDATE, INSERT, DELETE, and survivor
branches into a shared output.
h2. Root cause
The affected-row count is derived from the number of rows written by the
FileSink rather than from the logical operation that produced each row.
As a result:
- CoW UPDATE can include unchanged rows from rewritten files.
- CoW MERGE can include INSERT rows and unchanged survivor rows in the same
output.
- The reported `numModifiedRows` value does not necessarily represent the rows
actually updated by the statement.
h2. Suggested change
Add a `cow_update_matched` marker to the projections generated by the CoW
UPDATE and MERGE rewriters.
- Set the marker to `true` for rows produced by the matched UPDATE branch.
- Set the marker to `false` for unchanged survivor rows and other branches that
should not contribute to the UPDATE count.
- Make `FileSinkOperator` count only rows whose marker is `true` when the
marker mode is enabled.
- Keep the marker as an internal column so it is not written to the target
table.
This avoids relying on the physical Tez plan shape or operator layout to infer
the affected-row count.
h2. Confirmed behavior
The change should be covered by regression tests for:
- CoW UPDATE with matching rows
- CoW UPDATE with no matching rows
- CoW MERGE with UPDATE, DELETE, and INSERT branches
- CoW MERGE with only INSERT or DELETE branches
- Partitioned and unpartitioned tables
The tests should verify both the reported affected-row count and the final
table contents.
> Report accurate affected-row counts for copy-on-write UPDATE and MERGE
> ----------------------------------------------------------------------
>
> Key: HIVE-30064
> URL: https://issues.apache.org/jira/browse/HIVE-30064
> Project: Hive
> Issue Type: Bug
> Reporter: Ryu Kobayashi
> Priority: Major
>
> h2. Summary
> Copy-on-write UPDATE and MERGE operations can report an incorrect number of
> affected rows.
> The FileSink currently counts all rows written by the rewritten query.
> However, copy-on-write rewrites entire data files, so the output also
> contains unchanged survivor rows. MERGE additionally combines UPDATE, INSERT,
> DELETE, and survivor branches into a shared output.
> h2. Root cause
> The affected-row count is derived from the number of rows written by the
> FileSink rather than from the logical operation that produced each row.
> As a result:
> - CoW UPDATE can include unchanged rows from rewritten files.
> - CoW MERGE can include INSERT rows and unchanged survivor rows in the same
> output.
> - The reported `numModifiedRows` value does not necessarily represent the
> rows actually updated by the statement.
> h2. Suggested change
> Add a `cow_update_matched` marker to the projections generated by the CoW
> UPDATE and MERGE rewriters.
> - Set the marker to `true` for rows produced by the matched UPDATE branch.
> - Set the marker to `false` for unchanged survivor rows and other branches
> that should not contribute to the UPDATE count.
> - Make `FileSinkOperator` count only rows whose marker is `true` when the
> marker mode is enabled.
> - Keep the marker as an internal column so it is not written to the target
> table.
> This avoids relying on the physical Tez plan shape or operator layout to
> infer the affected-row count.
> h2. Confirmed behavior
> The change should be covered by regression tests for:
> - CoW UPDATE with matching rows
> - CoW UPDATE with no matching rows
> - CoW MERGE with UPDATE, DELETE, and INSERT branches
> - CoW MERGE with only INSERT or DELETE branches
> - Partitioned and unpartitioned tables
> The tests should verify both the reported affected-row count and the final
> table contents.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)