Hi all,

I created a PR for an issue I encountered while working on a data lake for a 
financial org.

Upsert was making in-python comparison, pushing this comparison to PyArrow 
Compute of course drastically reduces the memory and cpu usage.  There was also 
an inefficient cast and what I would call "daring" solution to the nested 
comparison issue, as PC cannot compare struct columns.

This work was partially done with Claude, originated from my mind, and I went 
through it a few times.

In mho this relatively simple PRs does bring a lot for people doing large 
upserts. If you have any comments or suggestions I will be more than happy to 
answer/implement :)

Best,

Gaspard Merten

Reply via email to