DanielLeens commented on issue #12058: URL: https://github.com/apache/seatunnel/issues/12058#issuecomment-5645798997
Thanks for making the candidate genuinely single-variable. I rechecked the current flush path: it can call an HDFS-aware sync, then a wrapped DFS sync or plain stream sync, followed by hflush. The proposed experiment replaces that path with exactly one hsync-family call per branch and no trailing hflush. That is acceptable as an isolated #12058 experiment, not as proof that #12081 resolves this issue and not as approval to merge its unrelated correctness changes on the strength of a benchmark result. Please keep the contract precise: - apply only the HdfsWriter.flush() control-flow change on the exact current-dev baseline; exclude RequestFuture, WAL failure handling, batch-deadline, and MapStore changes; - preserve write-through, persistence, and the caller's completion wait, with one hsync-family call for every successful append; - describe the local file:/// result as filesystem-specific. A null or weak movement is useful evidence; same-process visibility and mocked call counts are not a crash-durability or device-sync-count proof; - retain branch-level call-path coverage, then attach the same-environment wall, CPU, GC, raw per-fork, Score/Error/CV artifacts. The conclusion must compare the original InvocationFuture.get/park hotspot and the storage sync path, not only aggregate scores. With those boundaries, please run the A/B and report either result before proposing a production conclusion. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
