Hari Krishna Dara created PHOENIX-8012:
------------------------------------------
Summary: Report new-row count for UPSERT … ON DUPLICATE KEY IGNORE
batches via getUpdateCount()
Key: PHOENIX-8012
URL: https://issues.apache.org/jira/browse/PHOENIX-8012
Project: Phoenix
Issue Type: Improvement
Reporter: Hari Krishna Dara
When rows are inserted with {{UPSERT}} … {{ON DUPLICATE KEY IGNORE}} in batch
mode (autocommit off), there is currently no way to learn how many rows were
genuinely new versus already-existing (ignored). Both {{executeUpdate()}} and
{{getUpdateCount()}} return the client-side buffered row count (1 per
statement), captured before the server decides ignore-vs-insert, so their sum
equals the number of statements rather than the number of new rows.
Today, an application that needs the ignored/new count has only two
workarounds, both undesirable:
* Switch to point puts (autocommit on, one row per upsert), where each atomic
op reports its status individually. This forfeits the throughput benefit of
batching and issues one RPC per row.
* Run {{SELECT}} queries in advance to check which keys already exist. This is
not atomic — another writer can insert a key between the existence check and
the upsert — so the count can be wrong, and it doubles the round trips.
This tracking is possible without either workaround because the server already
emits a per-row "ignored" status (UPSERT_STATUS_CQ = 0) for every ignored row
in a multi-op batch; the client just never captures it in batch mode. The
new-row count can be derived as: rows buffered in the commit window minus rows
the server reported as ignored.
Make this count available through {{Statement.getUpdateCount()}} after
{{{}commit(){}}}, scoped to where it is unambiguous: within a single commit
window, exactly one Statement object executed mutations and all of them were
{{ON DUPLICATE KEY IGNORE}} upserts. When these preconditions are not met,
getUpdateCount() returns its existing value unchanged.
This behavior must be strictly opt-in via a new connection property (default
off) so that the standard getUpdateCount() behavior is preserved and there is
no backward-compatibility breakage for existing clients. Only when the property
is enabled does commit() populate the corrected new-row count on the statement.
Acceptance criteria:
* With the property enabled and the preconditions met, getUpdateCount() after
commit() returns the exact number of newly inserted rows.
* With the property disabled, or when preconditions are not met,
getUpdateCount() behavior is unchanged.
* No regression in existing {{ON DUPLICATE KEY}} / atomic-upsert tests.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)