DanielLeens commented on PR #11892: URL: https://github.com/apache/seatunnel/pull/11892#issuecomment-5363681434
Thanks @davidzollo — agreed on the direction, but I want to flag that CI is currently failing for a reason that traces directly back to this PR's diff, not an unrelated flake, so we shouldn't merge on green-CI faith yet. **The `Build` check is currently failing** (fork run: https://github.com/JacobZheng0927/seatunnel/actions/runs/32328300864), specifically in `connector-jdbc-e2e-part-7` → `JdbcOpenGaussIT` under the **Spark** engine job: ``` Exception in thread "main" org.apache.seatunnel.core.starter.exception.CommandExecuteException: Run SeaTunnel on spark failed Caused by: java.lang.NoClassDefFoundError: org/postgresql/util/PGobject Caused by: java.lang.ClassNotFoundException: org.postgresql.util.PGobject ``` (same failure reproduced in both matrix shards 8 and 11, so it's deterministic, not a one-off flake) Root cause, traced through the diff: - `JdbcOpenGaussIT`'s `CREATE_SQL` (`seatunnel-e2e/.../JdbcOpenGaussIT.java`) only has one PG-object-relevant column: `uuid_col UUID`. There's no `geometry`/`inet`/`interval` column in this table, so this test never previously exercised any `PGobject`-typed write path. - Before this PR, `uuid_col` fell through to the generic branch and was written with `statement.setString(...)` — no dependency on `org.postgresql.util.PGobject` at write time. - This PR's new branch in `PostgresJdbcRowConverter.setValueToStatementByDataType` (`.../dialect/psql/PostgresJdbcRowConverter.java`) now routes `PG_UUID` (and `PG_JSON`/`PG_JSONB`) through `new PGobject(); statement.setObject(...)`, so the write path now hard-requires the `org.postgresql.util.PGobject` class to be resolvable at runtime. - Under the Zeta engine that class is resolvable (plugin classloader has the pgjdbc driver jar), but under the **Spark** engine job in this E2E module it is not — hence `NoClassDefFoundError` only in the "Run SeaTunnel on spark failed" leg, and only now that `uuid_col` actually goes through `PGobject`. So this isn't a pre-existing/unrelated CI issue we can wait out — it's this PR's own new code path hitting a real classloader gap between the Zeta and Spark execution paths for `org.postgresql.util.PGobject`. Given the existing `geometry`/`inet`/`macaddr`/`interval` `PGobject` branches directly above the new code were apparently never covered by a Spark-engine E2E test with those column types, this may be a latent issue in that pre-existing code too — but this PR is what makes it observable and blocking, since `uuid_col` is a real column in this suite. @JacobZheng0927 could you take a look at why `org.postgresql.util.PGobject` isn't resolvable under the Spark job's classloader in this module (e.g. is the postgres driver dependency scope/shading different between the Zeta and Spark JDBC connector artifacts, or is it a Spark `--jars`/executor classpath assembly gap)? Once you have a fix, please push it and I'll re-review the updated head — happy to help dig into the Spark classpath assembly if useful. Given this, I'm holding off on re-affirming "ready to merge" until the Spark-engine failure above is actually resolved (not just re-run) — the `mergeStateStatus: BLOCKED` here reflects a real bug, not just a required-check gate. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
