SEZ9 commented on issue #12267:
URL: https://github.com/apache/seatunnel/issues/12267#issuecomment-5738788087

   Thanks @1362227089 for the config, the DDL, and the plain-text catalog 
output. One row each in `pg_class` and `pg_namespace` for `test_a` rules out a 
duplicate relation inside the catalog itself, so the `Duplicate key` must be 
introduced further along the discovery path.
   
   To recap where things stand on `dev` at 
`1325a44b2f4013198fdb2c12f9df345de1cde4bc`: table discovery reads `pg_database` 
first, then queries each database's `INFORMATION_SCHEMA.TABLES`, and the 
resulting `TableId` list is collected into a map without a merge function. 
Since your config only lists `highgo.testdb.test_a`, the same `TableId` is 
apparently being produced twice before schema serialization. We don't want to 
simply de-duplicate or add a merge function, because that could mask a wrong 
cross-catalog discovery result on HighGo.
   
   Remaining asks from the failing run (sanitized as before):
   
   1. The connector log lines that list the available databases, plus every 
`including '...' for further processing` line.
   2. Output of `SELECT datname FROM pg_database ORDER BY datname;`.
   3. For each catalog the connector lists, the rows returned for 
`testdb.test_a` from that catalog's `INFORMATION_SCHEMA.TABLES`, with 
`table_catalog`, `table_schema`, `table_name`, and `table_type`.
   
   That should tell us whether HighGo exposes the relation through more than 
one catalog query or whether the duplicate comes from the connector's own 
filtering. Once the source is confirmed, the fix should include a regression 
covering both an identical duplicate (handled deterministically) and 
conflicting metadata (fails with a clear message rather than silently picking 
one).
   
   One more note: if the credentials in the earlier configuration were real, 
please rotate them, and keep using redacted values in future updates.
   
   <!-- streview-comment:1163 -->


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to