jackye1995 commented on a change in pull request #3883: URL: https://github.com/apache/iceberg/pull/3883#discussion_r791188702
########## File path: core/src/main/java/org/apache/iceberg/BaseUpdateSnapshotRefs.java ########## @@ -0,0 +1,128 @@ +/* + * Licensed to the Apache Software Foundation (ASF) under one + * or more contributor license agreements. See the NOTICE file + * distributed with this work for additional information + * regarding copyright ownership. The ASF licenses this file + * to you under the Apache License, Version 2.0 (the + * "License"); you may not use this file except in compliance + * with the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, + * software distributed under the License is distributed on an + * "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY + * KIND, either express or implied. See the License for the + * specific language governing permissions and limitations + * under the License. + */ + +package org.apache.iceberg; + +import java.util.Map; +import org.apache.iceberg.relocated.com.google.common.base.Preconditions; +import org.apache.iceberg.relocated.com.google.common.collect.Maps; + +public class BaseUpdateSnapshotRefs implements UpdateSnapshotRefs { + + private final TableOperations ops; + private final TableMetadata base; + private final Map<String, SnapshotRef> refs; + + BaseUpdateSnapshotRefs(TableOperations ops) { + this.ops = ops; + this.base = ops.current(); + this.refs = Maps.newHashMap(base.refs()); + } + + @Override + public UpdateSnapshotRefs tag(String name, long snapshotId) { + Preconditions.checkArgument(name != null, "Tag name must not be null"); + Preconditions.checkArgument(base.snapshot(snapshotId) != null, "Cannot find snapshot with ID: %s", snapshotId); + Preconditions.checkArgument(!refs.containsKey(name), "Cannot tag snapshot, ref already exists: %s", name); Review comment: yeah the current API tries to have just `create`, `remove` and `rename` and ref property update APIs, because it's a bit unclear to me if we need a replace given that there is no such operation in schema and partition spec update. The action `replace` indicates the tag already exists, and I imagine it should throw exception if the tag does not exist. In that case, user still need to first check existence and then decide to create or replace, and it has little difference from just updating properties of the ref. If we allow `replace` to mean `put` that either create or update, then my concern is that it might cause unintented use that people might override a tag without knowing it already exists. The current API is basically trying to force a check on user side because there is not really a compare-and-swap of the specific ref that we could enforce as of today on server side and we have to replace the entire metadata file. any thoughts? ########## File path: core/src/main/java/org/apache/iceberg/BaseUpdateSnapshotRefs.java ########## @@ -0,0 +1,128 @@ +/* + * Licensed to the Apache Software Foundation (ASF) under one + * or more contributor license agreements. See the NOTICE file + * distributed with this work for additional information + * regarding copyright ownership. The ASF licenses this file + * to you under the Apache License, Version 2.0 (the + * "License"); you may not use this file except in compliance + * with the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, + * software distributed under the License is distributed on an + * "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY + * KIND, either express or implied. See the License for the + * specific language governing permissions and limitations + * under the License. + */ + +package org.apache.iceberg; + +import java.util.Map; +import org.apache.iceberg.relocated.com.google.common.base.Preconditions; +import org.apache.iceberg.relocated.com.google.common.collect.Maps; + +public class BaseUpdateSnapshotRefs implements UpdateSnapshotRefs { + + private final TableOperations ops; + private final TableMetadata base; + private final Map<String, SnapshotRef> refs; + + BaseUpdateSnapshotRefs(TableOperations ops) { + this.ops = ops; + this.base = ops.current(); + this.refs = Maps.newHashMap(base.refs()); + } + + @Override + public UpdateSnapshotRefs tag(String name, long snapshotId) { + Preconditions.checkArgument(name != null, "Tag name must not be null"); + Preconditions.checkArgument(base.snapshot(snapshotId) != null, "Cannot find snapshot with ID: %s", snapshotId); + Preconditions.checkArgument(!refs.containsKey(name), "Cannot tag snapshot, ref already exists: %s", name); + + refs.put(name, SnapshotRef.tagBuilder(snapshotId).build()); + return this; + } + + @Override + public UpdateSnapshotRefs branch(String name, long snapshotId) { + Preconditions.checkArgument(name != null, "Branch name must not be null"); + Preconditions.checkArgument(base.snapshot(snapshotId) != null, "Cannot find snapshot with ID: %s", snapshotId); + Preconditions.checkArgument(!refs.containsKey(name), "Cannot create branch, ref already exists: %s", name); Review comment: oh yes good point. Currently the expectation is that `TableMetadata` always updates the `main` branch to be the current snapshot ID so that operation is not needed. But it will cumbersome for a user to do remove and recreate the branch just to update the branch head. Do you think this is something we can add in another PR together with updates to the snapshot producers, or better to be out in the same one? ########## File path: core/src/main/java/org/apache/iceberg/BaseUpdateSnapshotRefs.java ########## @@ -0,0 +1,128 @@ +/* + * Licensed to the Apache Software Foundation (ASF) under one + * or more contributor license agreements. See the NOTICE file + * distributed with this work for additional information + * regarding copyright ownership. The ASF licenses this file + * to you under the Apache License, Version 2.0 (the + * "License"); you may not use this file except in compliance + * with the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, + * software distributed under the License is distributed on an + * "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY + * KIND, either express or implied. See the License for the + * specific language governing permissions and limitations + * under the License. + */ + +package org.apache.iceberg; + +import java.util.Map; +import org.apache.iceberg.relocated.com.google.common.base.Preconditions; +import org.apache.iceberg.relocated.com.google.common.collect.Maps; + +public class BaseUpdateSnapshotRefs implements UpdateSnapshotRefs { + + private final TableOperations ops; + private final TableMetadata base; + private final Map<String, SnapshotRef> refs; + + BaseUpdateSnapshotRefs(TableOperations ops) { + this.ops = ops; + this.base = ops.current(); + this.refs = Maps.newHashMap(base.refs()); + } + + @Override + public UpdateSnapshotRefs tag(String name, long snapshotId) { + Preconditions.checkArgument(name != null, "Tag name must not be null"); + Preconditions.checkArgument(base.snapshot(snapshotId) != null, "Cannot find snapshot with ID: %s", snapshotId); + Preconditions.checkArgument(!refs.containsKey(name), "Cannot tag snapshot, ref already exists: %s", name); + + refs.put(name, SnapshotRef.tagBuilder(snapshotId).build()); + return this; + } + + @Override + public UpdateSnapshotRefs branch(String name, long snapshotId) { + Preconditions.checkArgument(name != null, "Branch name must not be null"); + Preconditions.checkArgument(base.snapshot(snapshotId) != null, "Cannot find snapshot with ID: %s", snapshotId); + Preconditions.checkArgument(!refs.containsKey(name), "Cannot create branch, ref already exists: %s", name); Review comment: From the snapshot producer perspective, I was thinking 2 cases: 1) Append to branch: ``` table.newRewrite() .rewriteFiles(Sets.newSet(FILE_A), Sets.newSet(FILE_B)) .atBranch("beta") ``` In this case, I would say both use cases make sense: 1. user might want to fail when committing to a non-existing branch, because that branch was deleted by someone else. This could be satisfied by the current API by removing old branch and update it with the same info, a bit cumbersome but still works. 2. user might want to create the branch with the new snapshot ID, this could also be done by the current API by first checking branch existence and then do a create after snapshot ID is committed. 2) Add tag: ``` table.newRewrite() .rewriteFiles(Sets.newSet(FILE_A), Sets.newSet(FILE_B)) .tag("2022-01-01-compacted") ``` In this case, tagging with an existing name is likely a system issue as there should only be a single process doing the commit and tag it. This could be identified by the current API. So from my perspective the current API is enough, although the branch interface is not straight forward for people to just update the branch head. I would imagine it to be a frequent operation, so might worth adding that. Regarding the builder API, currently it's following the pattern of property update. I think `set` and `remove` would be beneficial for performing conditional update for specific refs and avoid full metadata check as we develop the rest catalog, do you think there is any other APIs for the builder that would be more suitable? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
