danny0405 commented on code in PR #10180:
URL: https://github.com/apache/hudi/pull/10180#discussion_r1407567997


##########
hudi-flink-datasource/hudi-flink/src/main/java/org/apache/hudi/configuration/FlinkOptions.java:
##########
@@ -125,6 +125,20 @@ private FlinkOptions() {
       .withDescription("Payload class used. Override this, if you like to roll 
your own merge logic, when upserting/inserting.\n"
           + "This will render any value set for the option in-effective");
 
+  public static final ConfigOption<String> INSERT_PARTITIONER_CLASS_NAME = 
ConfigOptions
+      .key("write.insert.partitioner.class.name")
+      .stringType()
+      .defaultValue("")
+      .withDescription("Insert partitioner to use aiming to re-balance records 
and reducing small file number "
+          + "in the scenario of multi-level partitioning. For example 
dt/hour/eventID"
+          + "Currently support 
org.apache.hudi.sink.partitioner.DefaultInsertPartitioner");
+
+  public static final ConfigOption<Integer> DEFAULT_PARALLELISM_PER_PARTITION 
= ConfigOptions
+      .key("write.insert.partitioner.parallelism.per.partition")
+      .intType()
+      .defaultValue(30)
+      .withDescription("The parallelism to use in each partition when using 
DefaultInsertPartitioner.");

Review Comment:
   Can we default fallback to the number of `write.tasks`, previously, the 
minimum file number is the parallelism of the tasks, we introduce another 
option for the number of files which could be confsing, it could also induces 
more small files for small tables?



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to