David-Banquet commented on issue #2512:
URL: 
https://github.com/apache/iceberg-python/issues/2512#issuecomment-5994641680

   I tried to reproduce this with the schema and steps from the description: 
two rows, then `delete`, `overwrite` and `upsert` filtered on the key. Each 
operation ran on a fresh table, with a watchdog that dumps the stack if a call 
takes more than 120 s.
   
   | Catalog | PyIceberg | Runs | Key `'123'` | Key `'123.123'` |
   |---|---|---|---|---|
   | SQL catalog (local SQLite) | `main` (068aae5) and 0.10.0 | 1 each | under 
0.4 s | under 0.1 s |
   | S3 Tables REST endpoint (`s3tables.eu-west-1.amazonaws.com/iceberg`, 
SigV4), Python 3.12 | `main` | 3 | 2.0 to 5.3 s | 2.0 to 4.6 s |
   | Same | 0.10.0 | 3 | 1.9 to 6.6 s | 1.8 to 3.6 s |
   
   Every operation returned the expected rows. I couldn't make it hang, and the 
dot in the key made no difference. The seconds on S3 Tables are network round 
trips.
   
   What I found while looking for a cause:
   
   - Commit retries (`commit.retry.*`, 30 minutes in total by default) only 
exist since 0.12.0 (#3320), so they can't explain a hang on 0.10.
   - @francocalvo's filter is `'123'`, without a dot, and the new data file was 
already in S3 when the call stopped returning. The write went through and 
something after it waited.
   - In 0.10.0 the REST catalog sets no timeout on its HTTP requests, so a 
request whose response never comes back blocks forever. That fits the symptom, 
but I have nothing showing it is what happened here.
   - `main` now has `rest.client.connection-timeout-ms` and 
`rest.client.socket-timeout-ms` (#3418, not released yet). They are off by 
default, and they are ignored when `rest.sigv4-enabled` is set: the SigV4 
adapter is mounted on the catalog URI, a longer prefix than `https://`, so it 
handles every catalog request without the timeout ([comment at 
L577-L581](https://github.com/apache/iceberg-python/blob/068aae50402285657b41f5357acb67672b0a053d/pyiceberg/catalog/rest/__init__.py#L577-L581),
 [mount at 
L1148](https://github.com/apache/iceberg-python/blob/068aae50402285657b41f5357acb67672b0a053d/pyiceberg/catalog/rest/__init__.py#L1148)).
 Against a local server that answers after 8 s, with a 2 s timeout configured, 
the call fails after 2 s without SigV4 and waits the full 8 s with it. That one 
is a bug on its own, so I can open a separate issue for it.
   
   @chidachu77 @francocalvo, a few things would help narrow this down:
   
   1. Which REST endpoint were you using: the S3 Tables one 
(`s3tables.<region>.amazonaws.com/iceberg`) or the Glue one 
(`glue.<region>.amazonaws.com/iceberg` with a 
`<account-id>:s3tablescatalog/<bucket>` warehouse)?
   2. Was the real table as small as the snippet? On large tables, #3129 (slow 
upsert and delete) may be the same problem.
   3. If it happens again, `py-spy dump --pid <pid>` while it hangs shows where 
the process is waiting.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to