Hi all, I created a PR for an issue I encountered while working on a data lake for a financial org.
Upsert was making in-python comparison, pushing this comparison to PyArrow Compute of course drastically reduces the memory and cpu usage. There was also an inefficient cast and what I would call "daring" solution to the nested comparison issue, as PC cannot compare struct columns. This work was partially done with Claude, originated from my mind, and I went through it a few times. In mho this relatively simple PRs does bring a lot for people doing large upserts. If you have any comments or suggestions I will be more than happy to answer/implement :) Best, Gaspard Merten
