hutiefang76 opened a new pull request, #12498: URL: https://github.com/apache/seatunnel/pull/12498
### Purpose of this pull request Closes #12497. With DuckDB JDBC 1.3.1, a 100-row `executeBatch` into DuckLake produced 100 Parquet files in a local reproduction. The JDBC sink's `batch_size` did not prevent this. This PR adds an opt-in `ducklake_bulk_write` path that stages each batch in a connection-local DuckDB temporary table, then performs one `INSERT ... SELECT` into the attached DuckLake table. The ordinary DuckDB/JDBC path remains the default. ### Does this PR introduce any user-facing change? Yes. For an existing DuckLake table, users can set `generate_sink_sql = true`, `database = "lake"`, `table = "main.events"`, and `ducklake_bulk_write = true`. An attached lake can be initialized on each JDBC connection through DuckDB's `session_init_sql_file` URL parameter; the PostgreSQL database, metadata schema, and data path are configured separately in that SQL file. The sink's dry-run connection check validates the attached target and input columns without DDL/DML. The mode is append-only and requires `schema_save_mode = IGNORE`, `data_save_mode = APPEND_DATA`, `auto_commit = true`, `max_retries = 0`, and a positive `batch_size`. It rejects custom write SQL, primary keys/upserts, COPY, and XA. A task replay after an uncertain commit can still duplicate rows; this change does not claim exactly-once delivery. English and Chinese JDBC/DuckDB docs describe the setup and limits. The DuckDB docs also remove an XA example that referred to a class absent from the pinned 1.3.1 driver. ### How was this patch tested? - JDK 17: `mvn -o -pl seatunnel-connectors-v2/connector-jdbc -am -DskipTests -q verify` passed; module Spotless applied. - Targeted JDK 17 tests: `DuckLakeBulkWriteTest` (3), `JdbcSinkFactoryTest` (18), and existing `DuckDBSourceAndSinkTest` (1) all passed. The opt-in DuckLake test uses the matching DuckLake and SQLite scanner extension files, compares 100 files from ordinary JDBC batching with one additional file from the SeaTunnel bulk sink, runs the dry-run check, and reads back 100 distinct rows after reconnect. The other tests cover configuration rejection, an unsuccessful transfer followed by job-style replay, and decimal/timestamptz/null values through staging. - Manual deployment smoke: DuckDB JDBC 1.3.1 + matching extensions attached DuckLake to an isolated PostgreSQL `lake_metadata` database with a non-public `dl_meta` metadata schema and a MinIO S3 data path. A 100-row staged insert containing nullable text, decimal, and timestamptz values yielded one Parquet object and all 100 rows were readable after reconnect. The temporary containers were removed afterward. ### Check list * [x] No new Jar binary package or connector module is added. * [x] English and Chinese connector documentation is updated. * [x] The option defaults to false, so existing DuckDB writes are unchanged. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
