JingsongLi commented on PR #9390: URL: https://github.com/apache/paimon/pull/9390#issuecomment-5748014991
Requirement fit: NEEDS-EVIDENCE. Implementation: FINDINGS. The corrected `usedInputs` handling addresses the previously reproduced unreferenced-column codegen failure, so the implementation direction is plausible. However, this PR adds roughly 700 lines duplicated across Spark 3.2/3.3/3.4 execution classes and a new user-facing switch, while the only performance evidence is an upstream Spark range (1.7x–32.3x), not a Paimon MERGE workload using this implementation. Please provide a small reproducible benchmark comparing the same Paimon MERGE actions with the option on/off, including generated-code size/compile time and rows/s; that evidence is needed to justify the cross-version maintenance surface. [P2] Regenerate the connector option documentation The PR adds `spark.paimon.write.merge.codegen.enabled` to `SparkConnectorOptions` but does not update `docs/generated/spark_connector_configuration.html`. The Java build reaches `ConfigOptionsDocsCompletenessITCase` and fails with “Documentation is outdated”; this is a PR-caused CI failure, not infrastructure. Regenerate the Spark connector configuration and document the option's default/supported Spark versions. Other red jobs include dependency-download 429s, but this docs failure is deterministic. I am not closing this PR because accelerating real Paimon MERGE execution can have end-to-end value; the value evidence and the failing generated-doc contract need to be completed first. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
