hgromer commented on code in PR #8694:
URL: https://github.com/apache/hbase/pull/8694#discussion_r4174076141
##########
hbase-backup/src/main/java/org/apache/hadoop/hbase/backup/impl/FullTableBackupClient.java:
##########
@@ -248,6 +263,75 @@ private void performSnapshots(Admin admin) throws
IOException {
}
}
+ /**
+ * Computes the per-host log boundaries stored by this full backup, using
+ * {@link BackupUtils#computeLogBoundaries}. A host that took part in this
backup's log roll gets
+ * its roll result: WALs created after the roll are included in the next
incremental backup, even
+ * though some of their edits may already be in the snapshot, which is safe
because deletes are
+ * replayed along with the puts. For any other host, its archived WALs and
the WALs of dead
+ * servers are covered by the snapshot, so the next incremental backup does
not replay them. The
+ * WALs of a live server that did not take part in the roll (for example one
that started during
+ * it) are pending, because they can still receive edits that are not in the
snapshot. The live
+ * servers are read only after the WAL directories are listed: a region
server creates its WAL
+ * directory only after the master registers it, so every live server whose
directory was listed
+ * is found. If the live servers do not include every host that took part in
the roll, the
+ * master's server list is incomplete (for example right after a master
failover), and the backup
+ * fails instead of treating live servers as dead. It must run right after
the log roll, before
+ * the snapshot is taken, so that a host starting later gets no boundary and
has all of its WALs
+ * included in the next incremental backup.
+ */
+ @RestrictedApi(
+ explanation = "Package-private for test visibility only. Do not use
outside tests.",
+ link = "",
+ allowedOnPath =
"(.*/src/test/.*|.*/org/apache/hadoop/hbase/backup/impl/FullTableBackupClient.java)")
+ static Map<String, Long> computeLogBoundaries(FileSystem fs, Path walRootDir,
+ Map<String, Long> rolledHosts, Admin admin) throws IOException {
+ Path logDir = new Path(walRootDir, HConstants.HREGION_LOGDIR_NAME);
+ Path oldLogDir = new Path(walRootDir, HConstants.HREGION_OLDLOGDIR_NAME);
+
+ Map<ServerName, List<String>> logsByServer = new HashMap<>();
+ for (FileStatus serverLogDir : fs.listStatus(logDir)) {
+ ServerName serverName =
+
AbstractFSWALProvider.getServerNameFromWALDirectoryName(serverLogDir.getPath());
+ if (serverName == null) {
+ continue;
+ }
+ List<String> logs = logsByServer.computeIfAbsent(serverName, k -> new
ArrayList<>());
+ for (FileStatus log : fs.listStatus(serverLogDir.getPath())) {
+ if (!AbstractFSWALProvider.isMetaFile(log.getPath())) {
+ logs.add(log.getPath().toString());
+ }
+ }
+ }
+
+ Set<ServerName> live = new HashSet<>(admin.getRegionServers());
+ Set<String> liveAddresses =
+ live.stream().map(sn ->
sn.getAddress().toString()).collect(Collectors.toSet());
+ if (!liveAddresses.containsAll(rolledHosts.keySet())) {
Review Comment:
Sadly, I can't think of a much better way to accomplish consistency here.
The HMaster may be initializing when we query for live RS, and unfortunately
that means we may get a partial list of RS back.
This sanity check is a little brittle, but I believe that it's robust enough
to do what we need. Given that we just rolled WAL files, if we don't see any of
those hosts in the live RS list we got back from the HMaster, we can likely
assume the HMaster didn't return a comprehensive list, and we should abort.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]