David-Banquet commented on issue #2599:
URL:
https://github.com/apache/iceberg-python/issues/2599#issuecomment-5983252242
I tried to reproduce this on current `main` (068aae5) with a table that has
`id` and `some_column`:
```python
with table.update_schema() as update:
update.rename_column("some_column", "renamed_column")
update.move_first("renamed_column")
```
| Catalog | `move_first("renamed_column")` | `move_first("some_column")` |
|---|---|---|
| SQL catalog (local SQLite) | `ValueError: Cannot move missing column:
renamed_column` | OK, schema becomes `renamed_column, id` |
| Glue + S3 (eu-west-1), 5 runs | 5/5 same `ValueError` | 5/5 OK |
It fails every time for me, on both catalogs, so I don't think Glue
consistency is involved. Inside one `update_schema()`, the `move_*` methods
look names up in the schema as it was when the update started, plus columns
added in the same update. Pending renames are not considered
([`_find_for_move`](https://github.com/apache/iceberg-python/blob/068aae50402285657b41f5357acb67672b0a053d/pyiceberg/table/update/schema.py#L505-L511),
same code in 0.10.0). If you pass the original name, the move works and the
rename is still applied.
The Java implementation behaves the same way. `SchemaUpdate.findForMove`
checks columns added in the update, then calls `schema.findField(name)` on the
base schema
([SchemaUpdate.java#L418-L430](https://github.com/apache/iceberg/blob/94250b2527f9e24082134be7da247ff5fc55df8c/core/src/main/java/org/apache/iceberg/SchemaUpdate.java#L418-L430)).
Names in the Java API refer to the base schema in general: a Java test renames
`preferences.feature2` in the same update that renames `preferences`, still
using the old parent name
([TestSchemaUpdate.java#L440-L441](https://github.com/apache/iceberg/blob/94250b2527f9e24082134be7da247ff5fc55df8c/core/src/test/java/org/apache/iceberg/TestSchemaUpdate.java#L440-L441)).
#3687 proposed resolving pending renames in `_find_for_move`. That would
make PyIceberg accept a call that Java rejects, so I'd keep the current
behavior and make it easier to understand instead: when the name only matches a
pending rename, the error would say to use the original name, and the "Rename
column" section of the docs would mention it. I can open a small PR for that if
maintainers agree.
@din14970, the part I can't explain is the "sometimes". Do you remember what
changed between the runs that passed and the ones that failed? For example the
PyIceberg version, the code path, or whether the rename and the move always ran
in the same `update_schema()` block.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]