hutiefang76 opened a new issue, #12497:
URL: https://github.com/apache/seatunnel/issues/12497

   ### Problem
   
   SeaTunnel's JDBC sink can write to a DuckLake table through DuckDB JDBC, but 
a JDBC `PreparedStatement.executeBatch()` of 100 INSERT rows with the currently 
pinned DuckDB JDBC 1.3.1 driver produced 100 Parquet files in a local DuckLake 
catalog. The same rows staged in a DuckDB temporary table and transferred with 
one `INSERT INTO lake.main.events SELECT ... FROM stage` produced one Parquet 
file. The row count remained correct after closing and reopening the connection.
   
   This matters for repeated SeaTunnel flushes into DuckLake: `batch_size` 
limits the Java batch but does not by itself bound DuckLake's small-file count.
   
   ### Proposed scope
   
   Add an opt-in insert-only bulk path to the existing JDBC/DuckDB sink. Keep 
ordinary DuckDB writes unchanged. Require an existing DuckLake table, avoid 
unsupported XA/upsert semantics, and document that job replay after an 
uncertain commit can still duplicate rows.
   
   ### Validation
   
   The behavior was reproduced with DuckDB JDBC 1.3.1 and matching DuckLake 
extension. A separate PostgreSQL metadata catalog using a non-public schema and 
MinIO S3 data path also accepted a 100-row staged insert and read back all rows 
after reconnect; it produced one Parquet object.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to