SEZ9 commented on issue #12267:
URL: https://github.com/apache/seatunnel/issues/12267#issuecomment-5747035074

   Thanks @1362227089, this settles where the duplicate comes from.
   
   Your results show:
   
   - `pg_database` lists `cflag`, `highgo`, `template0`, `template1`, matching 
the connector log (`list of available databases is: [highgo, template1, 
template0, cflag]`).
   - `pg_class` / `pg_namespace` contain exactly one `testdb.test_a` (oid 
`334287`, schema oid `334286`).
   - `information_schema.tables` in `highgo` returns six identical rows for 
`testdb.test_a`, and the connector logs six `including 'highgo.testdb.test_a' 
for further processing` lines. The same query in `cflag` returns nothing.
   
   So the duplicate is not introduced by the connector's filtering path or by 
cross-catalog discovery. HighGo's `information_schema.tables` returns the same 
relation multiple times, `TableDiscoveryUtils.listTables()` passes each row 
through, and `PostgresIncrementalSource.tableChanges()` then fails in 
`Collectors.toMap` because there is no merge function.
   
   Since the six rows carry identical `table_catalog`, `table_schema`, 
`table_name` and `table_type`, the direction from earlier applies: identical 
discovered `TableId` entries should be collapsed deterministically, while 
conflicting metadata for the same identifier should still produce a descriptive 
failure rather than a silent pick. The regression test should cover both cases.
   
   One small thing: the last SQL block in your comment appears to be cut off 
after `WHERE c.relname = 't`. If there was more output there, please re-paste 
it.
   
   <!-- streview-comment:1177 -->


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to