Hi, I think we're missing a recheck in ProcessSyncingTablesForApply(). We take the subscription lock before marking a table READY, but continue to use the SYNCDONE state we cached earlier. A concurrent REFRESH PUBLICATION can remove the table while we're waiting for the lock. Once we get it, there's no pg_subscription_rel row left to update, and apply errors out.
I asked an LLM to review this and it could reproduce the missing-row error. With disable_on_error = true, that disables the whole subscription. The attached fix reads the state and LSN again under the lock. Just checking for a row wouldn't be enough: two refreshes could remove and re-add the table, and using the old SYNCDONE state would skip the new initial copy. The TAP test just checks the A/B behaviour: the subscription gets disabled without the fix and stays enabled with it. I tested remove/re-add separately, but left it out of the TAP test because it needs quite a bit more setup and there has been no field report of this issue yet (that I know of). I added the test as a separate patch for the same reason, in case we want to leave it out given how niche the case is. If we want to add specific tests, I'm happy to send revised patches with those. Thoughts? Regards, Ayush
0002-Test-table-sync-after-a-concurrent-refresh.patch
Description: Binary data
0001-Recheck-table-sync-state-after-refresh.patch
Description: Binary data
