Jump to content

Recommended Posts

Posted

Hi all,

 

Before I start let me make it clear I'm no longer in an IT post, so I am effectively an end user at the moment!

 

How long would you consider it to be acceptable to have loss of access to your SAN? Personal documents etc are all accessible but shared areas where the majority of project work etc are kept are not.

 

Thoughts please?

Posted

Hard question really, if its broken then it'll be down until its fixed or until someone has restored backups somewhere.

 

If its planned downtime, then depending on the make and model and the number of controllers then an update for example can take anywhere from an hour to 10 hours.

Posted (edited)

That's a "how long is a piece of string?" question.

 

Without knowing the exact circumstances under which the SAN has become unavailable, it's impossible to say what an acceptable recovery time will be.

 

Ideally it won't become unavailable at all ... but things break, or need patching, and sometimes the Gods throw awful combinations together. And that's without throwing in external factors like contractors cutting cables, builders on site, etc.

 

And it also depends on the person/people supporting it and the hours they're being paid for. They're entitled to sleep, holiday, sickness and only to work when they're being paid to do so.

Edited by elsiegee40
  • Thanks 1
Posted

How long would you consider it to be acceptable to have loss of access to your SAN? Personal documents etc are all accessible but shared areas where the majority of project work etc are kept are not.

I think that not only is there a "how long is a piece of string" element, but the term SAN is now quite wide ranging (i.e. abused!) - encompassing anything from a cheap NAS doing iSCSI or even these days directly running SAMBA for file shares, through to complex multi node systems which should be entirely immune to ANY downtime (yea, in theory).

 

In a sense the question is the wrong way round, as the end user you are the one who can actually say what the value (loss) is to you of not having access to your data. That should be what drives the requirement for storage availability (how robust that 'SAN' needs to be) and dictating the Service Level whoever is providing the service should meet. At the minimum you should come away from an actual problem knowing either that you can survive such an outage or that you need to talk to the service provider about storing your stuff on a more robust system.

Posted

It also depends on how much whoever's paying for the SAN is (or was) willing to spend on service contracts, out-of-hours in-house support and a host of other things.

 

Unfortunately, sometimes that money only gets spent after it hits the fan and $person_who_saved_money loses a lot of face.

Posted

Personally I think its acceptable to be down while it is being fixed. Whether it takes 30 minutes or a week. Ideally it would be better if it was fixed quickly but sometimes this doesn't happen.

 

But in a company whilst it is down it is potentially losing the company money so there would be a difference of opinion there.

 

Also as @elsiegee40 says it depends on the person fixing it and being paid etc.

Posted

If things are that critical the actual departments should have another method of working whilst the system is being repaired. Here in the NHS for example our core switch tripped several times due to another faulty switch recently which totally knocked out VMWare as the initial failure it tried to kick in HA, but then with the networking tripping several times in quick succession it confused itself as to where the VMs actually were and took us nearly a whole day to resolve as we had to contact VMWare etc.

 

What was the solution for the affected departments? Good old pen and paper :)

 

In many places I've seen it being a case of 'It's fine we spent £948475472340 on disaster recovery and failover' but never actually test it works so when it comes down to it their fancy solution is about as much use as a chocolate fireguard!

Posted

Just repeating what was said really but depends a lot.

 

Is this expected loss e.g. maintenance/migration etc? Then as long as it was planned for, whether minutes or days. You have to work on systems to do bits to them and unless you're rich enough to have redundant everything, there's times stuff goes down.

 

Is it unexpected aka broken? Then who's job is it to fix it? What hours are they working? Is it major or minor? If as you mentioned it's not a dead SAN and "certain areas" it's not going to be something simple one way or another in terms of asking us how long :p

 

Steve

Posted
With the suitable funding and resources it's possible to have highly resilient systems but unfortunately in education funding is often squeezed this all has a knock on effect in lots of area's.
Posted

This was an unexpected loss. It's just come back up, went down Thursday morning. Apparently multiple RAID issues and backups not copying correctly - sounds like a right headache tbf. Diagnosis was quick, but rectification was slow. Hopefully problems have been addressed and fixed as opposed to "get it working". PITA side is deadlines have been missed and this will cost us money.

 

Thanks for the thoughts, now enjoy the rest of your week! :-)

Posted
This was an unexpected loss. It's just come back up, went down Thursday morning. Apparently multiple RAID issues and backups not copying correctly - sounds like a right headache tbf. Diagnosis was quick, but rectification was slow. Hopefully problems have been addressed and fixed as opposed to "get it working". PITA side is deadlines have been missed and this will cost us money.

 

Thanks for the thoughts, now enjoy the rest of your week! :-)

 

If it went down Thursday morning and up when you posted assuming the IT Department were only working week days then thats quite quick in my opinion.

 

Even if they were working the weekend I'm guessing they would have wanted to get it working right rather than have multiple times in future where it would go down again.

 

Its unfortunate its lost the company money but in the long run it would be a minimal amount in the financial year.

 

Here's to hoping it is fixed for the long term.

  • Thanks 2
Posted (edited)

This is what SLA's are for... both IT and "users" agree what is acceptable and do-able.

If management/users want a shorter SLA than IT can currently provide, money needs spent (either on redundancy or people)

SLA's are much more than timers for jobs to provide stats. They are very useful in getting funding if used and implemented correctly.

 

We lost one of our main SAN's for 2 and a half a day a a while ago due to massive RAID failure (multiple HD failures, battery died and RAID controller failed all within a few hours of each other). Said SAN was on 3rd party limited support due to being so old. It took 2 people about two days (two full days mind- they were working from home in the evenings) to fix/replace/restore everything to working order. Once a full report was given for the down time, we suddenly got money for new SAN's with full support from the manufacture and better redundancy.

Edited by arwen

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...