This is an automated email from the ASF dual-hosted git repository.

nicholasjiang pushed a commit to branch branch-0.5
in repository https://gitbox.apache.org/repos/asf/celeborn.git


The following commit(s) were added to refs/heads/branch-0.5 by this push:
     new a604b6c3a [CELEBORN-1457] Avoid NPE during shuffle data cleanup
a604b6c3a is described below

commit a604b6c3ae3e99c0706efbed71f5977b16c93570
Author: jiang13021 <[email protected]>
AuthorDate: Wed Jun 12 16:14:55 2024 +0800

    [CELEBORN-1457] Avoid NPE during shuffle data cleanup
    
    ### What changes were proposed in this pull request?
    Avoid NPE during shuffle data cleanup by checking for null LevelDB.
    
    ### Why are the changes needed?
    If the LevelDB in StorageManager fails to initialize, the db will be null. 
This will cause a java.lang.NullPointerException when 
storageManager.cleanupExpiredShuffleKey(expiredShuffleKeys) is called, and the 
shuffle data in expiredShuffleKeys will not be cleaned up. The worker's disk 
may be filled up as a result.
    
    ### Does this PR introduce _any_ user-facing change?
    No
    
    ### How was this patch tested?
    Manual Testing
    
    Closes #2553 from jiang13021/celeborn-1457.
    
    Authored-by: jiang13021 <[email protected]>
    Signed-off-by: SteNicholas <[email protected]>
    (cherry picked from commit 8e2fe74a60a690911d07db6744b167d66e4c45d1)
    Signed-off-by: SteNicholas <[email protected]>
---
 .../celeborn/service/deploy/worker/storage/StorageManager.scala      | 5 +++--
 1 file changed, 3 insertions(+), 2 deletions(-)

diff --git 
a/worker/src/main/scala/org/apache/celeborn/service/deploy/worker/storage/StorageManager.scala
 
b/worker/src/main/scala/org/apache/celeborn/service/deploy/worker/storage/StorageManager.scala
index 857939929..e767060c7 100644
--- 
a/worker/src/main/scala/org/apache/celeborn/service/deploy/worker/storage/StorageManager.scala
+++ 
b/worker/src/main/scala/org/apache/celeborn/service/deploy/worker/storage/StorageManager.scala
@@ -232,8 +232,9 @@ final private[worker] class StorageManager(conf: 
CelebornConf, workerSource: Abs
       reloadAndCleanFileInfos(this.db)
     } catch {
       case e: Exception =>
-        logError("Init level DB failed:", e)
-        this.db = null
+        throw new IllegalStateException(
+          "Failed to initialize db for recovery during graceful worker 
shutdown.",
+          e)
     }
     saveCommittedFileInfosExecutor =
       ThreadUtils.newDaemonSingleThreadScheduledExecutor(

Reply via email to