zhang-arvin opened a new pull request, #18202:
URL: https://github.com/apache/iceberg/pull/18202

   Fixes #18162
   
   ## What
   
   Adds a new Spark session configuration `spark.sql.iceberg.row-level-mode` 
that lets a Spark session pick the row-level operation mode (copy-on-write vs 
merge-on-read) for `DELETE`, `UPDATE` and `MERGE` without changing any table 
property.
   
   Accepted values are `copy-on-write` and `merge-on-read`; the value is parsed 
with the existing `RowLevelOperationMode.fromName`, so an invalid value fails 
fast with `Unknown row-level operation mode: <value>`.
   
   ## Priority
   
   ```
   spark.sql.iceberg.row-level-mode (session)  >  
write.{delete,update,merge}.mode (table)  >  copy-on-write (default)
   ```
   
   When the session config is not set, behaviour is byte-for-byte identical to 
before — the table property is consulted, and the hard-coded `copy-on-write` 
default applies when the table property is absent. Only when the session config 
is present does the per-command table property get skipped.
   
   ## How
   
   `SparkRowLevelOperationBuilder` already resolves the mode per command from 
the table properties. It now reads 
`spark.conf().get(SparkSQLProperties.ROW_LEVEL_OPERATION_MODE, null)` and, when 
non-null, uses it in preference to the table property. This mirrors how a 
single `spark.sql.iceberg.distribution-mode` session key overrides the 
per-command `write.(delete|update|merge).distribution-mode` table properties.
   
   The property is added to `SparkSQLProperties` next to the existing 
`DISTRIBUTION_MODE` key.
   
   ## Tests
   
   New `TestSessionRowLevelOperationMode` (spark-extensions), covering:
   
   - `testSessionModeOverridesTableProperties` — the table is explicitly 
configured for copy-on-write on all three commands, the session is set to 
`merge-on-read`; `DELETE`/`UPDATE`/`MERGE` must all produce delete files 
(merge-on-read) rather than rewritten data files.
   - `testTablePropertiesUsedWhenSessionModeIsNotSet` — with the session key 
unset, the copy-on-write table properties still win for all three commands, and 
no delete files are added.
   
   The new test runs against the existing parameter matrix (Hive/REST catalog, 
ORC/PARQUET/AVRO, v2/v3 format, local/distributed planning).
   
   ## Notes for reviewers
   
   - Applied to all Spark version modules (3.5, 4.0, 4.1, 4.2).
   - `spark.conf().get(key, null)` is used rather than `SparkConfParser` 
because the Spark option parser's `sessionConf` applies to write options keyed 
by `spark.sql.iceberg.*`; the row-level operation builder has no write-option 
map to feed it. Happy to switch to `SparkConfParser` if you'd prefer a single 
parsing path — it would mean threading the session conf into the builder's 
`mode` resolution.
   - I did not touch the docs (`docs/spark-configuration.md`); let me know if 
you want it documented there as part of this PR.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to