ravwojdyla commented on code in PR #1240:
URL: https://github.com/apache/parquet-mr/pull/1240#discussion_r1430866723
##########
parquet-hadoop/src/main/java/org/apache/parquet/hadoop/MemoryManager.java:
##########
@@ -74,7 +75,7 @@ private void checkRatio(float ratio) {
* @param writer the new created writer
* @param allocation the requested buffer size
*/
- synchronized void addWriter(InternalParquetRecordWriter<?> writer, Long
allocation) {
+ void addWriter(InternalParquetRecordWriter<?> writer, Long allocation) {
Review Comment:
@ConeyLiu yes it does. One of our production Spark jobs took close to 2
hours with partitioned parquet output and about 45 minutes without. Also please
see the stats in JIRA issues:
* https://issues.apache.org/jira/browse/PARQUET-2412
* https://issues.apache.org/jira/browse/SPARK-44003
> I think it should only happen when you are writing a file with many small
partitions.
In that particular test the partitions were probably relatively small.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]