Okay, in that sense, I agree.  The more general point, however, is that if a decision 
comes down to availability vs. data integrity, integrity wins.

But that doesn't mean we don't make every effort to maintain availability whenever 
possible.

Unfortunately, the cost of this can be quite high, which is why you don't see very 
many Tandem systems except in places where there's no alternative.

That hospital I used to work for had an official policy of 24x7 availability, but they 
weren't willing to spend the money required to actually achieve it.  This caused all 
KINDS of headaches when
things didn't go exactly as planned.

The last 5% accounts for 90% of the cost.

> -----Original Message-----
> From: Linux on 390 Port [mailto:[EMAIL PROTECTED] Behalf Of
> Dale Strickler
> Sent: Wednesday, October 29, 2003 10:36 AM
> To: [EMAIL PROTECTED]
> Subject: Re: [LINUX-390] Memory access faults.
>
>
> At 10:10 AM 2003_10_29, you wrote:
> >Well, given that I work for a financial institution, I can
> say that in
> >many cases "stopping everything" is exactly what DOES
> happen.  "Charging
> >ahead", knowing you're dealing with potentially
> >corrupted data and not knowing the extent of the problem, is
> irresponsible.
>
> I guess I did not state my point well, as you mention next :
>
> >For the REALLY critical stuff, parallel redundant systems are used
> >(Tandem, etc.), on the theory that a single failure can't
> knock down more
> >than part of the application.
>
> This is what I mean by "charge ahead."  If the system detects
> a failure in
> a transaction, so the transaction can not be trusted, then a redundant
> system must re-try.  (Or must be trusted that it handled the
> transaction
> correctly, assuming no fault was detected in the redundant
> system.  I think
> NASA polls 5 systems with a majority wins approach?)
>
> Of course known data is essential and yes the 'fall back'
> positions need to
> be as safe as possible.  The one level of confusion that
> often arises in
> discussion of these systems is the level of the system that is being
> discussed.  Specific routines may be stopped and request
> human intervention
> but the system or sub-systems must be able to cope with that
> part being
> shut down it if the over all system is truly a critical system.  For
> instance a nuke plant run off of one Window NT box that could
> give a 'Blue
> Screen of death' would be a REALY bad system design...
> Thought that would
> be an example of a system deciding to 'stop' and alert the user when
> something went wrong.
>
>
>
> >In a previous job I worked for a hospital.  Most of the
> systems we managed
> >were NOT involved in direct patient care, and it's a good
> thing.  When we
> >DID start getting involved in that area, it became
> >VERY scary.
> >
> >
> > >
> > > This is good *IF* it is not a critical system.  If the
> application is
> > > moving billions of financial transactions around the world
> > > and it costs
> > > brokers millions of dollars for every minute of down time
> just "stop
> > > everything, and let someone fix it" is not a good answer.
> > > The application
> > > needs to identify the failure point, establish what is likely
> > > good or bad
> > > data and charge ahead.  (After leaving a solid trail of bread
> > > crumbs for
> > > someone to follow....)
> > >
> > > - Dale
> > >
>
> Dale Strickler
> Cole Software, LLC
> Voice: 540-456-8896
> Fax: 540-456-6658
> Web: http://www.colesoft.com/
>

Reply via email to