zhangshenghang opened a new pull request, #12525: URL: https://github.com/apache/seatunnel/pull/12525
### Purpose of this pull request Close #12521. `AbstractSchema#indexOf`, `#getColumn` and `#contains` scan the `columnNames` list linearly on every call. These lookups sit on hot paths (for example sink writers resolving field indexes per table), and for wide tables with hundreds of columns the repeated O(n) scans add significant overhead. This PR builds a lazily-initialized name-to-index map (double-checked, thread safe) and uses it for the lookups: - `indexOf`/`contains` become hash lookups; `getColumn` reuses `indexOf` and keeps its previous semantics (including the exception behavior for missing columns). - Duplicate column names keep the first occurrence, matching the former linear-scan semantics. - The constructor now takes a defensive copy of the column list (unmodifiable), so external mutation of the caller's list can no longer change the schema or invalidate the cache. ### Does this PR introduce _any_ user-facing change? No behavior change is intended. The only observable difference is that mutating the list passed to the schema constructor after construction no longer affects the schema (previously an undocumented aliasing hazard), and `getColumns()` now returns an unmodifiable view of an internal copy. ### How was this patch tested? - Added `AbstractSchemaTest` covering lookups, duplicate-name first-match semantics, and the defensive copy. - `mvn -pl seatunnel-api test`: 404 tests, all passing. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
