David-Banquet commented on issue #2599:
URL: 
https://github.com/apache/iceberg-python/issues/2599#issuecomment-5983252242

   I tried to reproduce this on current `main` (068aae5) with a table that has 
`id` and `some_column`:
   
   ```python
   with table.update_schema() as update:
       update.rename_column("some_column", "renamed_column")
       update.move_first("renamed_column")
   ```
   
   | Catalog | `move_first("renamed_column")` | `move_first("some_column")` |
   |---|---|---|
   | SQL catalog (local SQLite) | `ValueError: Cannot move missing column: 
renamed_column` | OK, schema becomes `renamed_column, id` |
   | Glue + S3 (eu-west-1), 5 runs | 5/5 same `ValueError` | 5/5 OK |
   
   It fails every time for me, on both catalogs, so I don't think Glue 
consistency is involved. Inside one `update_schema()`, the `move_*` methods 
look names up in the schema as it was when the update started, plus columns 
added in the same update. Pending renames are not considered 
([`_find_for_move`](https://github.com/apache/iceberg-python/blob/068aae50402285657b41f5357acb67672b0a053d/pyiceberg/table/update/schema.py#L505-L511),
 same code in 0.10.0). If you pass the original name, the move works and the 
rename is still applied.
   
   The Java implementation behaves the same way. `SchemaUpdate.findForMove` 
checks columns added in the update, then calls `schema.findField(name)` on the 
base schema 
([SchemaUpdate.java#L418-L430](https://github.com/apache/iceberg/blob/94250b2527f9e24082134be7da247ff5fc55df8c/core/src/main/java/org/apache/iceberg/SchemaUpdate.java#L418-L430)).
 Names in the Java API refer to the base schema in general: a Java test renames 
`preferences.feature2` in the same update that renames `preferences`, still 
using the old parent name 
([TestSchemaUpdate.java#L440-L441](https://github.com/apache/iceberg/blob/94250b2527f9e24082134be7da247ff5fc55df8c/core/src/test/java/org/apache/iceberg/TestSchemaUpdate.java#L440-L441)).
   
   #3687 proposed resolving pending renames in `_find_for_move`. That would 
make PyIceberg accept a call that Java rejects, so I'd keep the current 
behavior and make it easier to understand instead: when the name only matches a 
pending rename, the error would say to use the original name, and the "Rename 
column" section of the docs would mention it. I can open a small PR for that if 
maintainers agree.
   
   @din14970, the part I can't explain is the "sometimes". Do you remember what 
changed between the runs that passed and the ones that failed? For example the 
PyIceberg version, the code path, or whether the rename and the move always ran 
in the same `update_schema()` block.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to