TheR1sing3un opened a new pull request, #10044:
URL: https://github.com/apache/paimon/pull/10044

   ### Purpose
   
   `BlobStore.list_objects(prefix=..., limit=...)` currently materializes every 
object's metadata before filtering and applying the limit. A request for one 
matching object therefore allocates metadata for the entire table.
   
   Read metadata in Arrow batches and stop once enough objects match. Push the 
limit into the reader when no prefix is supplied; otherwise retain the existing 
`str(key).startswith(prefix)` behavior, including numeric and null keys. Use 
the managed batch reader to close the underlying iterator on early return and 
exceptions, including when an exception traceback remains alive.
   
   The result remains a list of `ObjectInfo` values. Object bodies are not 
fetched, and an unlimited call still retains its full result list.
   
   ### Tests
   
   - Multimodal table suite: 86 passed; final BlobStore regression suite: 6 
passed.
   - Focused BlobStore tests on PyArrow 16: 6 passed.
   - Regression coverage checks bounded batch consumption, limit placement, 
selected metadata columns, numeric/null key prefixes, empty results, and 
immediate iterator cleanup while an exception traceback is retained.
   - Flake8, changed-file license checks, Python 3.6 grammar checks, and `git 
diff --check` passed.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to