Hi Yuxia, Thank you for the clarifying comments, narrowed scope looks much better. Best Regards, Mehul Batra
On Sat, Aug 8, 2026 at 6:08 PM Jark Wu <[email protected]> wrote: > Hi Yuxia, > > Thanks for revisiting and refining the proposal. Narrowing the scope > to removing the mandatory system columns makes the FIP much clearer > and more focused. > > The current FIP looks good to me. LGTM. > > Best, > Jark > > On Thu, 6 Aug 2026 at 18:46, Yuxia Luo <[email protected]> wrote: > > > > Hi, all. > > If no any further comments, I'd like to start vote in next monday. > > > > On 2026/08/03 02:28:22 Yuxia Luo wrote: > > > Hi Lorenzo and Mehul, > > > > > > Thank you for the detailed feedback. My sincere apologies for the > delayed response. I’d like to revive this FIP and move it forward again. > > > > > > I have narrowed the FIP to focus only on removing mandatory system > columns. PARTITION_TIMESTAMP has been removed and can be discussed > separately if needed. > > > > > > Replies to Lorenzo > > > ------------------ > > > > > > 1. Scope > > > > > > I agree that the previous scope was misleading. > > > > > > Removing the system columns does not remove lake access. FULL mode > will > > > continue reading a lake snapshot before switching to Fluss, and > $lake will > > > continue using the native Paimon or Iceberg source. > > > > > > The lake reader and tiering writer will support both legacy tables > with > > > system columns and new tables with only user-defined columns. > > > > > > We may still provide an offset system column for offset-ordered lake > reads > > > in the future. Since adding such a column changes the user-visible > physical > > > schema and sort order, it should require an explicit table option. > > > > > > 2. Union reads > > > > > > This FIP does not change the existing FULL union-read behavior. > > > > > > The lake-to-Fluss transition uses snapshot-level per-bucket offsets > stored > > > with the lake snapshot metadata. It does not depend on the __offset > value > > > of each lake row. > > > > > > 3. Auto-partitioned table restriction > > > > > > This restriction belonged to PARTITION_TIMESTAMP and is no longer > part of > > > this FIP. > > > > > > Replies to Mehul > > > ---------------- > > > > > > 1. Schema and legacy detection > > > > > > The schema inspection will remain inside each lake-format plugin. > When an > > > existing lake table is opened, the Paimon or Iceberg implementation > will > > > classify it as: > > > > > > - legacy, when all three expected system columns exist; or > > > - clean, when none of them exists. > > > > > > The mode will be derived from the physical lake schema rather than > stored > > > separately in Fluss metadata. A partial set of system columns or > unexpected > > > types will cause schema validation to fail. Manual changes to a > > > Fluss-managed lake schema remain unsupported. > > > > > > 2. PARTITION_TIMESTAMP and partition boundaries > > > > > > Questions about timestamp-to-partition mapping, validation of > manually > > > partitioned tables, and CDC boundary handling are no longer part of > this > > > FIP. > > > > > > A separate proposal for partition-based historical reads should > define > > > those semantics explicitly, including duplicate handling. > > > > > > 3. Existing union-read boundary > > > > > > FULL mode will continue using the snapshot-level per-bucket offsets > reported > > > through the existing tiering metadata path. Removing per-row system > columns > > > does not change this transition point. > > > > > > 4. Backward compatibility > > > > > > Existing lake tables retain the legacy schema. A table dropped and > recreated > > > after the upgrade will use the clean schema, so users depending > directly on > > > __bucket, __offset, or __timestamp must migrate first. > > > > > > The safe upgrade order is: > > > > > > 1. Flink connector and lake plugin > > > 2. Tiering service > > > 3. Fluss cluster > > > > > > The cluster is upgraded last so it cannot create clean tables while > old > > > writers or lake readers are still running. > > > > > > 5. Paimon and Iceberg > > > > > > Yes, the change applies to both Paimon and Iceberg. Their catalog, > writer, > > > and reader implementations must both support the legacy and clean > layouts. > > > > > > Regarding the FIP name, the broader name made sense when > > > PARTITION_TIMESTAMP was included. Since that feature has been > removed, I > > > propose keeping the current title. > > > > > > Thanks again for the feedback. > > > > > > > > > On 2026/03/03 19:35:23 Mehul Batra wrote: > > > > Also I think the current name of the FIP doesn't do justice to the > scope, > > > > can we rename it something like: > > > > *Optional System Columns and Partition-Based Historical Reads* > > > > > > > > On Wed, Mar 4, 2026 at 12:41 AM Mehul Batra < > [email protected]> > > > > wrote: > > > > > > > > > Hi Yuxia, > > > > > > > > > > First of all thank you for leading this, It's an important aspect > as this > > > > > is non-trivial storage cost in Parquet/ORC files for columns that > most > > > > > consumers never read and schema introspection gets polluted too. > > > > > I've been going through FIP-27 in detail and have a few questions > I'd like > > > > > to clarify before implementation begins. Grouping them by area: > > > > > > > > > > *1. Schema & Legacy Detection* > > > > > > > > > > 1a. When datalake is re-enabled on an existing table and the lake > table > > > > > already exists with system columns, where does the schema > inspection > > > > > happen do we add a new method to the LakeCatalog interface (e.g., > > > > > getTableSchema(TablePath)), or is this handled at the Fluss server > > > > > metadata level outside the plugin boundary? > > > > > > > > > > 1b. Is the legacy/clean mode decision persisted in Fluss table > metadata > > > > > (e.g., as a property like fluss.lake.schema.mode = legacy | > clean), or is > > > > > it re-derived by inspecting the lake table schema each time? If > re-derived, > > > > > what happens if someone manually alters the lake table schema > externally? > > > > > > > > > > *2. PARTITION_TIMESTAMP Mode* > > > > > > > > > > 2a. The FIP shows a day-granularity example for > timestamp-to-partition > > > > > mapping. Can we document the exact mapping for all supported > time-unit > > > > > values (hour, day, month, quarter, year)? I assume it follows the > same > > > > > DateTimeFormatter patterns in PartitionUtils, but it would be good > to > > > > > make this explicit. > > > > > > > > > > 2b. Should the Flink connector fail fast at job submission time > (via > > > > > ValidationException) if PARTITION_TIMESTAMP is used on a > > > > > non-auto-partitioned table? Or do we allow it for manually > partitioned > > > > > tables as well? > > > > > > > > > > 2c. For PK tables with CDC, how are duplicates at the partition > boundary > > > > > resolved during the union read? Is it the same > snapshot-then-changelog > > > > > pattern that FULL mode uses today? The FIP mentions "downstream > > > > > idempotency" but CDC duplicate handling is non-trivial it would > help to be > > > > > more specific here. > > > > > > > > > > *3. Union Read Boundary* > > > > > > > > > > 3a. How is the exact transition point from lake historical reads > to Fluss > > > > > log reads determined per-partition is it the per-partition tiering > > > > > watermark stored in Fluss server metadata? > > > > > > > > > > *4. Backward Compatibility* > > > > > > > > > > 4a. If a user drops and recreates a table with the same name > post-upgrade, > > > > > the new lake table will not have system columns. Should we warn > users about > > > > > this schema change, especially if they have downstream jobs that > depend on > > > > > __offset or __bucket? > > > > > > > > > > *5. Scope* > > > > > > > > > > 5a. The changes apply to both Paimon and Iceberg lake catalogs, > correct? > > > > > Both PaimonLakeCatalog and IcebergLakeCatalog currently append > system > > > > > columns independently. > > > > > > > > > > Thanks for the FIP, happy to help with the implementation once > these are > > > > > clarified. > > > > > > > > > > > > > > > Best Regards, > > > > > Mehul Batra > > > > > > > > > > On Mon, Mar 2, 2026 at 7:26 PM Lorenzo Affetti via dev < > > > > > [email protected]> wrote: > > > > > > > > > >> Hello! I went through the FIP another time as I did not remember > doing it > > > > >> already :) > > > > >> > > > > >> I have additional questions beyond the first 2. > > > > >> > > > > >> Let me paste those here and add: > > > > >> > > > > >> 1. Isn't the scope of the FIP misleading? > > > > >> This FIP seems to be about removing system columns, but it > primarily > > > > >> proposes a new read mode named PARTITION_TIMESTAMP. > > > > >> Is this because removing those columns prevents users from > accessing data > > > > >> on the lake? > > > > >> If so: > > > > >> - how do user are supposed to do that now > > > > >> - What would change > > > > >> > > > > >> 2. How does this relate to union reads? > > > > >> I am quite new to the community and Fluss. Could you explain how > the new > > > > >> PARTITION_TIMESTAMP mode relates to union reads? > > > > >> If the answer is not obvious, perhaps this warrants a section in > the FIP. > > > > >> > > > > >> 3. Why *"*Only auto partitioned table is supported in this mode"? > > > > >> Why only for partitions generated by Fluss, and not for any > partition that > > > > >> represents a timestamp? > > > > >> > > > > >> On Wed, Feb 4, 2026 at 4:50 PM Lorenzo Affetti < > > > > >> [email protected]> wrote: > > > > >> > > > > >> > Hello Yuxia! > > > > >> > Thanks for the great FIP! > > > > >> > I have some questions: > > > > >> > > > > > >> > 1. Isn't the scope of the FIP misleading? > > > > >> > It seems this FIP is about removing system columns, but it > primarily > > > > >> > proposes a new read mode named PARTITION_TIMESTAMP. > > > > >> > > > > > >> > 2. How does this relate to union reads? > > > > >> > I am quite new to the community and Fluss. Could you explain > how the new > > > > >> > PARTITION_TIMESTAMP mode relates to union reads? > > > > >> > If the answer is not obvious, perhaps this warrants a section > in the > > > > >> FIP. > > > > >> > > > > > >> > Thank you! > > > > >> > > > > > >> > On Tue, Jan 20, 2026 at 8:20 AM yuxia < > [email protected]> > > > > >> wrote: > > > > >> > > > > > >> >> Hi, all. > > > > >> >> > > > > >> >> Currently, every Fluss lake table is automatically provisioned > with > > > > >> three > > > > >> >> mandatory system columns, __bucket , __offset , __timstamp > (intended > > > > >> for > > > > >> >> bucket and offset-based subscription as well as addition > informartion > > > > >> >> check). > > > > >> >> While originally designed to allow clients to pinpoint > specific data > > > > >> >> offsets of specific buckets, the practical evolution of the > ecosystem > > > > >> has > > > > >> >> rendered this default behavior suboptimal for the dowstream > since the > > > > >> >> dowstream warehouse or BI tools do not expect these internal > metadata > > > > >> >> fields. > > > > >> >> > > > > >> >> > > > > >> >> So, I'd like to propose FIP-27: Remove Mandatory System > Columns From > > > > >> >> Fluss Lake Tables [1] to remove the three mandatory system > columns > > > > >> while > > > > >> >> still keep compability. > > > > >> >> > > > > >> >> Welcome your feedback and suggestions on this proposal. > Looking forward > > > > >> >> to a productive discussion! > > > > >> >> > > > > >> >> [1]: > > > > >> >> > > > > >> > https://cwiki.apache.org/confluence/display/FLUSS/FIP-27%3A+Remove+Mandatory+System+Columns+From+Fluss+Lake+Tables > > > > >> >> > > > > >> >> Best regards, > > > > >> >> Yuxia > > > > >> >> > > > > >> > > > > > >> > > > > > >> > -- > > > > >> > Lorenzo Affetti > > > > >> > Senior Software Engineer @ Flink Team > > > > >> > Ververica <http://www.ververica.com> > > > > >> > > > > > >> > > > > >> > > > > >> -- > > > > >> Lorenzo Affetti > > > > >> Senior Software Engineer @ Flink Team > > > > >> Ververica <http://www.ververica.com> > > > > >> > > > > > > > > > > > > >
