zhangshenghang opened a new issue, #12521: URL: https://github.com/apache/seatunnel/issues/12521
### Description `AbstractSchema#indexOf`, `#getColumn` and `#contains` scan the `columnNames` list linearly on every call. These lookups sit on hot paths (for example sink writers resolving fields per table per flush), and for wide tables with hundreds of columns the per-row/per-lookup cost becomes O(n) each time, which adds up to significant overhead in wide-table synchronization jobs. Since the column list is immutable after construction, the name-to-index mapping can be built once (lazily) and reused, keeping the same first-match semantics for duplicate names as the linear scan. ### Impact Wide-table jobs pay repeated O(n) scans per lookup; the same lookup repeated for every row degrades throughput as the column count grows. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
