anew opened a new pull request, #57527:
URL: https://github.com/apache/spark/pull/57527

   ### What changes were proposed in this pull request?
   What changes were proposed in this pull request?
   
   This PR replaces https://github.com/apache/spark/pull/57488, which had to be 
abandoned due to severe semantic merge conflicts. It is an exact cherry-pick of 
the previous commits, plus a fix for the test cases that failed due to the 
merge conflicts.
     
   What were the merge conflicts? Two PRs disagreeing on the semantics when a 
source contains a reserved column but does not include that column in the 
column selection. This PR rejects that because it validates pre-column 
selection. The other PR (SPARK-58313) allowed that because it was doing the 
validation port column selection. I decided to fallbacj to the original 
behavior that already existed for SCD type 1 validate pre column selection. And 
we have SPARK-58325 to reconsider whether we want to do validation post column 
selection. 
   
   AutoCdcMergeFlow validates at construction time that a flow's source 
change-data feed does not carry columns that collide with AutoCDC's internal 
columns. Until now, that validation 
(requireReservedPrefixAbsentInSourceColumns) only rejected column names 
starting with the reserved prefix __spark_autocdc_.
     
     SCD2, however, persists two framework columns to the target that do not 
carry that prefix — __START_AT and __END_AT (see 
Scd2BatchProcessor.reservedFrameworkColNames). A source column named __START_AT 
or __END_AT therefore slipped past the guard and would be silently overwritten 
during microbatch preprocessing.
     
     This PR closes that gap (the TODO(SPARK-57251) in Scd2BatchProcessor):
     
     - Adds requireReservedFrameworkColumnsAbsentInSourceColumns() to the 
AutoCdcMergeFlow constructor. For SCD2 flows it rejects any source  column 
whose name collides (by exact name, resolver-aware so it respects 
spark.sql.caseSensitive) with a non-prefixed reserved framework column. SCD1 
targets carry no such columns, so the check is a no-op for SCD1.
     - Runs the check before the flow's schema val is forced, so the actionable 
reserved-name error surfaces ahead of the temporary AUTOCDC_SCD2_NOT_SUPPORTED 
gate, and the check remains correct once SCD2 support lands.
     - Adds a new error condition AUTOCDC_RESERVED_COLUMN_NAME_CONFLICT 
(SQLSTATE 42710), distinct from the existing prefix-based 
AUTOCDC_RESERVED_COLUMN_NAME_PREFIX_CONFLICT, since this collision is by exact 
name rather than by prefix.
     - Widens the visibility of Scd2BatchProcessor.reservedFrameworkColNames to 
private[pipelines] so the flow layer can validate against the single source of 
truth.
   
   ### Why are the changes needed?
    Without this check, a user whose CDC source happens to contain a __START_AT 
or __END_AT column would have that data silently overwritten by AutoCDC's SCD2 
framework columns, with no error and no diagnostic — a data-correctness 
footgun. Failing fast at flow construction with a user-actionable error 
("rename or remove the column") is the intended UX, consistent with the 
existing reserved-prefix guard.    
    
   ### Does this PR introduce _any_ user-facing change?
   Yes. An AutoCDC SCD2 flow whose source change-data feed contains a column 
named __START_AT or __END_AT (subject to case-sensitivity settings) now fails 
at flow construction with AUTOCDC_RESERVED_COLUMN_NAME_CONFLICT instead of 
silently overwriting the column. There is no change for SCD1 flows, and no 
change for SCD2 flows whose sources do not use these names. (Note: SCD2 AutoCDC 
flows are not yet generally supported on master — they are still gated by 
AUTOCDC_SCD2_NOT_SUPPORTED — so no released behavior changes.)
    
   ### How was this patch tested?
   New unit tests in AutoCdcFlowSuite covering:                                 
                                                           
     - an SCD2 flow with a __START_AT/__END_AT source column is rejected with 
AUTOCDC_RESERVED_COLUMN_NAME_CONFLICT;                         
     - the reserved-name check fires before the AUTOCDC_SCD2_NOT_SUPPORTED 
gate;                                                             
     - an SCD1 flow with the same column name is allowed and the column 
survives into the flow schema;                                       
     - case-sensitivity behavior (spark.sql.caseSensitive true/false);          
                                                             
     - a guard test asserting the set of non-prefixed reserved names is exactly 
{__START_AT, __END_AT}, so a future rename can't silently un-cover the 
validation.                                                                     
                                           
   
   ### Was this patch authored or co-authored using generative AI tooling?
    Generated-by: Claude Code (Opus 4.8) 
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to