hgromer commented on code in PR #8694:
URL: https://github.com/apache/hbase/pull/8694#discussion_r4173793761


##########
hbase-backup/src/main/java/org/apache/hadoop/hbase/backup/impl/IncrementalBackupManager.java:
##########
@@ -97,15 +107,16 @@ private List<String> excludeProcV2WALs(List<String> 
logList) {
   }
 
   /**
-   * For each region server: get all log files newer than the last timestamps 
but not newer than the
-   * newest timestamps.
+   * Gather all log files that either: 1) are newer than the older timestamps, 
but not newer than
+   * the newest timestamps, or 2) are archived logs whose host name does not 
occur in the newest
+   * timestamps.
    * @param olderTimestamps  the timestamp for each region server of the last 
backup.
    * @param newestTimestamps the timestamp for each region server that the 
backup should lead to.
    * @param conf             the Hadoop and Hbase configuration
-   * @return a list of log files to be backed up
+   * @return the log files to be backed up, and the log files held back for a 
later backup
    * @throws IOException exception
    */
-  private List<String> getLogFilesForNewBackup(Map<String, Long> 
olderTimestamps,
+  private LogFileSelection getLogFilesForNewBackup(Map<String, Long> 
olderTimestamps,

Review Comment:
   If a server crashes as we're doing log rolls + backups, its WALs will sit in 
the `/WALs` directory until that split is finished. We need to have a concept 
of WALs that were "held back", which are set as `LogFileSelection#pending`. 
   
   We will keep the boundary at before the oldest held back WAL, ensuring that 
the log cleaner doesn't delete it. These WALs will be backed up in the next 
incremental.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to