zhoulii opened a new issue, #9696:
URL: https://github.com/apache/paimon/issues/9696

   ### Search before asking
   
   - [x] I searched in the [issues](https://github.com/apache/paimon/issues) 
and found nothing similar.
   
   
   ### Paimon version
   
   master
   
   ### Compute Engine
   
   java & python api
   
   ### Minimal reproduce step
   
   **Java**
   
   1. Create a table with `id` and `embedding ARRAY<FLOAT> COMMENT 
'__VECTOR_FIELD;3'`, with row tracking and data evolution enabled.
   2. Insert a row and create a tag named `before_add`.
   3. Add `embedding_v2` using the same vector directive.
   4. Reload the table and read the `before_add` tag, projecting only `id`.
   
   **Python**
   
   1. Create a data-evolution table with a `__BLOB_DESCRIPTOR_FIELD` or 
`__BLOB_VIEW_FIELD` column. Write a non-null reference and retain the snapshot.
   2. Drop that column, reload the table, and read the retained snapshot.
   
   ### What doesn't meet your expectations?
   
   Expected: historical reads should use the field declarations from the 
selected snapshot. Adding or dropping columns should not break earlier 
snapshots or change their returned values.
   
   Actual:
   
   - Java fails before reading data with `Some of the columns specified as 
vector-field are unknown.` Similar validation failures occur after adding BLOB 
columns.
   - Python returns serialized descriptor/view bytes instead of resolving the 
original payload.
   
   Both paths combine historical fields with current schema options, leaving 
the field declarations inconsistent.
   
   
   ### Anything else?
   
   _No response_
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to