On Wed, Aug 25, 2010 at 04:26:43PM +0200, Michael Hanselmann wrote: > This is an additional patch on top of my previous design for > workerpool priorities. > > Signed-off-by: Michael Hanselmann <[email protected]> > --- > The design has been changed to have a priority per opcode. Since this is > mostly > Ganeti-internal, the feature will still be called “job priorities”. > > doc/design-2.3.rst | 50 ++++++++++++++++++++++++++++++++++++++++++++++++-- > 1 files changed, 48 insertions(+), 2 deletions(-) > > diff --git a/doc/design-2.3.rst b/doc/design-2.3.rst > index e45a077..46162c5 100644 > --- a/doc/design-2.3.rst > +++ b/doc/design-2.3.rst > @@ -154,12 +154,58 @@ Job priorities > Current state and shortcomings > ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ > > -.. TODO: Describe current situation > +Currently all jobs and opcodes have the same priority. Once a job > +started executing, its thread won't be released until all opcodes got > +their locks and did their work. When a job is finished, the next job is > +selected strictly by its incoming order. This does not mean jobs are run > +in their incoming order—locks and other delays can cause them to be > +stalled for some time. > + > +In some situations, e.g. an emergency shutdown, one may want to run a > +job as soon as possible. This is not possible currently if there are > +pending jobs in the queue. > > Proposed changes > ~~~~~~~~~~~~~~~~ > > -.. TODO: Describe changes to job queue and potentially client programs > +Each opcode will be assigned a priority on submission. Opcode priorities > +are integers and the lower the number, the higher the opcode's priority > +is. Within the same priority, jobs and opcodes are initially processed > +in their incoming order. > + > +Submitted opcodes can have one of the priorities listed below. Other > +priorities are reserved for internal use. Opcodes submitted without a > +priority (e.g. by older clients) are assigned the default priority. > + > + - High (20) > + - Normal (40, default) > + - Low (60)
Whatever the values, let's also documents the minimum (described below as zero), and the maximum. > +As a change from the current model where executing a job blocks one > +thread for the whole duration, the new job processor must return the job > +to the queue after each opcode and also if it can't get all locks in a > +reasonable timeframe. This will allow opcodes of higher priority > +submitted in the meantime to be processed or opcodes of the same > +priority to try to get their locks. When added to the job queue's > +workerpool, the priority is determined by the first unprocessed opcode > +in the job. > + > +If an opcode is deferred, the job should stay in the "waitlock" status. > +Technically it's waiting to acquire its locks. As discussed offline, I disagree here. It's in the queue, and not being processed, hence it's just 'queued'. Nothing breaks today with jobs being marked as queued between opcodes, and I don't see a reason to change. My rationale against waitlock is that when debugging a cluster, I will ignore 'queued' jobs (I know those are completely idle), and only look at what's in waitlock status. > +If an opcode can not be processed after a certain number of retries or a > +certain amount of time, it should increase its priority. This will avoid > +starvation. > + > +A job's priority can never go above zero (value 0). If a job hits > +priority 0, it must acquire its locks in blocking mode. I think this should read "go below zero". "go above zero" means coming from below zero, which is impossible. > +Opcode priorities are synchronized to disk in order to be restored after > +a restart or crash of the master daemon. > + > +Priorities also need to be considered inside the locking library to > +ensure opcodes with higher priorities get locks first, but the design > +changes for this will be discussed in a separate section. iustin
