Copilot commented on code in PR #10045:
URL: https://github.com/apache/paimon/pull/10045#discussion_r4059657155
##########
docs/docs/pypaimon/writing.md:
##########
@@ -144,3 +144,23 @@ table_commit.close()
| `snapshot` | `Snapshot` | The committed snapshot (id,
commit_kind, time_millis, next_row_id, …) |
| `commit_entries` | `List[ManifestEntry]` | Delta manifest entries in this
commit (each carries `file.first_row_id` when row-tracking is enabled) |
| `identifier` | `int` | Commit identifier
|
+
+## DATE Partition Naming
+
+For tables with a `DATE` partition key, Python reads and writes follow the
+existing `partition.legacy-name` setting, matching Java. The default `true`
+uses days since `1970-01-01`; `false` uses ISO dates. For example,
+`1970-01-02` maps to `day=1` or `day=1970-01-02`, respectively. Record values
+remain dates. Date-shaped STRING values and separate integer year, month,
+and day columns are outside this DATE fix.
+
+Paths are derived from the table configuration without probing alternative
+partition directories. Explicit external file paths remain authoritative.
+
+Older Python writers used ISO dates even when `partition.legacy-name=true`.
+Such files require a separate migration before using the corrected reader;
+there is no automatic fallback to their historical directories. Upgrading the
+writer does not relocate old files, and older Python readers may not read new
Review Comment:
This statement is too broad: `read/native_plan.py:264-289` still probes the
legacy raw partition directory with `list_status` and restores files found
there. Native planning therefore retains an automatic fallback even though the
regular Python planner does not; qualify this documentation and the following
migration warning by planner type.
##########
paimon-python/pypaimon/read/scanner/split_generator.py:
##########
@@ -38,20 +40,19 @@ class AbstractSplitGenerator(ABC):
def __init__(
self,
- table,
+ table: "FileStoreTable",
target_split_size: int,
open_file_cost: int,
deletion_files_map: Optional[Dict] = None,
snapshot_id: Optional[int] = None,
- ):
+ ) -> None:
self.table = table
self.target_split_size = target_split_size
self.open_file_cost = open_file_cost
self.deletion_files_map = deletion_files_map or {}
self.snapshot_id = snapshot_id
- self.default_part_value = table.options.options.get(
- CoreOptions.PARTITION_DEFAULT_NAME, "__DEFAULT_PARTITION__")
-
+ self.path_factory = table.path_factory()
Review Comment:
This constructor now unconditionally calls `table.path_factory()`, so
existing lightweight table fixtures that only provide `table_path` and
`options` fail before split generation (for example,
`interval_partition_test.py` and `data_evolution_group_stats_test.py`). Update
every such fixture to provide the new path factory, or defer this lookup until
path materialization with a compatible fallback.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]