jjj-n commented on PR #12265:
URL: https://github.com/apache/seatunnel/pull/12265#issuecomment-5746840891

   Thanks for the second review. I checked the cleanup branch against the 
merged dev parent and the remaining CI failures.
   
   The early-fire cleanup rescheduling at `processPendingJobCleanup` was 
introduced upstream by #11503, commit ec1b1b8b51a7aa39c208cf39473f8f44104b3939 
(September 13). It is already present in this merge's dev parent 
c7304ace6e18d350314e92480df1fd3c0962f1f2 and is unchanged by this PR relative 
to that parent. It appeared between reviewed heads because 4a5fbc94 merged dev. 
I agree dedicated coverage would be useful, but it is not a cleanup behavior 
change introduced by the executor split. Could you reassess that blocking item 
against the base-to-head diff?
   
   The BackpressureSlowSinkIT change has a local differential result: the 
original 90-second loop observed only two additional completed checkpoints both 
with this PR and with only CoordinatorService replaced by the unmodified dev 
version. Observed checkpoint duration reached 34 seconds despite a 15-second 
trigger interval, and the old loop did not sample after its final sleep. The 
revised loop preserves at least 90 seconds of stress, at least three additional 
checkpoints, zero failed checkpoints and the backpressure metric assertions, 
with a 180-second upper bound. The final three embedded-engine integration 
tests passed together (154.707 seconds). This reproduces the completion-count 
failure; it does not claim to reproduce the earlier CI first-checkpoint timeout.
   
   The completed CI run on 4a5fbc94 is not fully green: 
https://github.com/jjj-n/seatunnel/actions/runs/35302309118 . All four Java 
8/11 Windows/Ubuntu unit-test jobs and both Java 8/11 engine-v2-it jobs passed. 
Remaining failed jobs:
   
   - all-connectors-it-2 / Java 8: Maven Central timed out downloading 
hadoop-common:3.3.4 before connector-cdc-mongodb-e2e dependency resolution 
completed (job 105468814056).
   - all-connectors-it-1 / Java 8 and 11: NebulaGraphIT.startUp:109 fails at 
adminPool.init after the graph client cannot ping the server, before connector 
job execution (jobs 105468814101 and 105468814191).
   - rocketmq-connector-it / Java 11: 87 tests, 14 failures and 29 errors; logs 
show missing topic routes, including failures in producer data generation and 
across Flink/Spark/Zeta executions (job 105468814771). This points to shared 
broker/topic setup rather than establishing an executor regression, but I am 
not treating it as a confirmed harmless flake or claiming the full matrix 
passes.
   
   I am rerunning the failed/cancelled checks on the same head to distinguish 
transient failures from repeatable setup problems. The PR remains draft while 
those results are pending.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to