Dan,
Perhaps my email wasn't clear. It says "also need to provide user
instructions to backup the entire repo". By "also" I meant "in
addition to" a gold backup/recovery.
Thanks for clarifying a couple points I could not get an answer to:
1) Our only backup devices are the cluster disks.
2) RFO is driving the requirements
I hate to continue this discussion since Vincent is only pursuing the
gold option, but when you say that we do not have enough disk space on
the rfo cluster, that is incorrect. The repo could grow ~4x it's
current size before we would run out of disk space to duplicate the
repo. We have roughly 5.7 TB on the cluster and the current repo is
671 GB. Also, no one ever commented on my previous email discussing
how to make the repo significantly smaller. That's fine, because we
should do the gold method anyway and can get more space efficient later.
To wrap this up,
1) Vincent you are looking into all the code changes we need to
implement the gold file backup. Is that correct?
2) In the mean time, we are not backing up the rfo repo. So if any
one of the 6 repo sections (the repository is randomly distributed
over 6 agent machines) crashes, then we will be left with random
missing files. Our options then would be to either wait for the code
changes (the changes for the gold recovery will recover from this
situation) or we simply trash the entire database and repo and start
over.
Is there anything else?
Bob Gobeille
[email protected]
On Aug 10, 2009, at 6:58 PM, Dan Stangel wrote:
Bob,
I agree with you, the best user experience would come from backup and
restore of the entire repository. In our case, however, there's a big
problem: We do not have enough unallocated disk space on the RFO
cluster to
implement this.
So I don't see a problem in recommending the full repo backup
method, but in
our case we will not be able to implement this, and will have to use
the
gold files-only method.
Dan
-----Original Message-----
From: [email protected]
[mailto:[email protected]] On Behalf Of Gobeille, Robert
Sent: Monday, August 10, 2009 2:18 PM
To: Ma, Dong (vinc...@gdcc-bj-most)
Cc: [email protected]
Subject: Re: [FOSSology] unpack size/time
Hi Vincent,
As per my previous email, this is probably not a great way to back up
a fossology database. But it's good functionality to have because we
could add a feature that allows users to remove all non-reused files
except gold to save disk space. I don't see any bad ramifications
except for the user experience (waiting for unpack). The code changes
needed in unpack are minimal. All the places in the UI that read
files will have to change. So you don't want to create this backup
solution until all the code changes are done.
I think you also need to provide user instructions to backup the
entire repo. That is:
stop scheduler
pg_dumpall > myfile
backup myfile
backup repo gold, files, and license
start scheduler
Bob
On Aug 10, 2009, at 1:31 AM, Ma, Dong (vinc...@gdcc-bj-most) wrote:
Hi Bob,
If no dependencies with the unpacked files at the backup point, I
think we should not do unpack on the fly at the restore time. We
just restore the gold files and give the user's opinion to unpack
the gold files as they needed.
Is this opinion make sense? Or this will bring some bad
ramifications?
Thanks,
Vincent
-----Original Message-----
From: Gobeille, Robert
Sent: Friday, August 07, 2009 10:56 PM
To: Ma, Dong (vinc...@gdcc-bj-most)
Cc: [email protected]
Subject: Re: unpack size/time
On Aug 7, 2009, at 2:20 AM, Ma, Dong (vinc...@gdcc-bj-most) wrote:
Hi Bob,
Also have 2 more questions talk with you:
1. Backup all repo files, the incremental disk backup time is not
an
issue, but the restore process will also cost so much time when
restore all the repo file every time(one day or more). I just
consider this situation:
a. only backup gold(and license) files
b. when restore, only unpack the files which have dependency with
backup point running jobs (if no dependency with running jobs, will
not unpack any gold files in restore)
See previous email. There are no dependencies from the interrupted
unpack.
c. unpack other files with user's demand after restore
I am not sure the percentage of 'unpack the files which have
dependency with backup point running jobs' with whole repo unpack
files, if very little unpacked files should be unpacked in restore
process, the time cost will not the issue. May be the time will
less
than restore all the repo file.
I hope this question not confuse you.
Restore only the gold (and license) files would typically be much
faster than restoring all the repo files. Se my previous email in
this thread for times.
2. About ' Allow users to "archive" uploads.', I have a question is
if users have this requirement? If user want this feature of delete
the unpacked files? My thought is the users don't care about
unpacked files occupied how many disk space and want do delete them
if disk spaces is enough.
So far, this has not come up as a user requirement. The only thing
I've seen from users is the need to delete the entire upload from
the
database and repository. And we already do that.
Bob Gobeille
_______________________________________________
fossology mailing list
[email protected]
http://fossology.org/mailman/listinfo/fossology
_______________________________________________
fossology mailing list
[email protected]
http://fossology.org/mailman/listinfo/fossology