moneycat957 commented on issue #10443:
URL: https://github.com/apache/seatunnel/issues/10443#issuecomment-5566019567
I'd like to pick up the **CockroachDB** row.
@sug-ghosh claimed this on 2026-03-23 and I don't see a CockroachDB PR in
`apache/seatunnel`
(open or merged) since. Per @davidzollo's four-week rule at the top of this
thread I'd like to
take it, but @sug-ghosh — if you're still working on this, say so and I'll
drop it immediately
and pick another row instead.
**Proposed shape: a `connector-jdbc` dialect, not a standalone connector**
CockroachDB speaks the PostgreSQL wire protocol, so the obvious first
question is whether it
needs anything of its own at all. The repo already answers that:
`connector-jdbc` carries 31
dialects, and five of them are PostgreSQL-wire-compatible databases that
each got a dedicated
dialect rather than reusing `psql` — `greenplum`, `highgo`, `opengauss`,
`kingbase` and
`redshift`. There is no CockroachDB dialect today.
Concrete reasons a dedicated dialect earns its place, rather than
"compatible but different"
in the abstract:
1. **Upsert.** `JdbcDialect.getUpsertStatement()` is a first-class dialect
method.
`PostgresDialect` emits `INSERT ... ON CONFLICT (...) DO UPDATE SET ...`.
CockroachDB has a
native `UPSERT` statement with different semantics (it does not evaluate
a conflict target and
is cheaper for the common case), so the sink path should emit the native
form rather than
inherit the Postgres one.
2. **Type surface.** CockroachDB does not implement several PostgreSQL types
the `psql` type
converter maps (`money`, `oid` family, some array/range behaviours), so
`PostgresTypeConverter`'s mapping is not a safe inheritance.
3. **Identity/serial semantics.** `SERIAL` resolves through `unique_rowid()`
rather than a
sequence, which changes what a generated-key path can assume.
**V1 scope**
Following the shape you asked for on the DocumentDB row (#12046):
*In scope* — a new `cockroachdb` dialect package under
`connector-jdbc/.../internal/dialect/`, mirroring the `psql` package's five
main classes
(`Dialect`, `DialectFactory`, `JdbcRowConverter`, `TypeConverter`,
`TypeMapper`), with the
native `UPSERT` statement, a CockroachDB-specific type mapping, dialect
registration, unit
tests matching the four `psql` test classes, and bilingual docs.
*Out of scope for V1* — CDC/changefeeds, catalog auto-DDL beyond what the
JDBC connector
already does generically, multi-region/locality-aware routing, and
transaction-retry
(`40001` serialization failure) retry policy, which I'd rather propose
separately once the
basic dialect is in.
**Two questions before I start**
1. Is a `connector-jdbc` dialect the direction you want here, or does the
tracker row
("Source/Sink CockroachDB") imply you'd rather have a standalone
`connector-cockroachdb`
module? The dialect route is far smaller and matches how every other
Postgres-compatible
database in this repo is handled, so that's my default unless you say
otherwise.
2. For E2E — `cockroachdb/cockroach` publishes an official Docker image that
starts a
single-node instance with `start-single-node --insecure`, so unlike the
DocumentDB row a
real E2E is genuinely achievable here. Should I include a
`connector-jdbc-cockroachdb-e2e`
module in the first PR, or land the dialect with unit tests first and add
E2E as a
follow-up?
I'll open the PR against `dev` within the four-week window.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]