SamanthaD Posted August 21, 2017 Posted August 21, 2017 Hi all, Before I start let me make it clear I'm no longer in an IT post, so I am effectively an end user at the moment! How long would you consider it to be acceptable to have loss of access to your SAN? Personal documents etc are all accessible but shared areas where the majority of project work etc are kept are not. Thoughts please?
RobD Posted August 21, 2017 Posted August 21, 2017 Hard question really, if its broken then it'll be down until its fixed or until someone has restored backups somewhere. If its planned downtime, then depending on the make and model and the number of controllers then an update for example can take anywhere from an hour to 10 hours.
elsiegee40 Posted August 21, 2017 Posted August 21, 2017 (edited) That's a "how long is a piece of string?" question. Without knowing the exact circumstances under which the SAN has become unavailable, it's impossible to say what an acceptable recovery time will be. Ideally it won't become unavailable at all ... but things break, or need patching, and sometimes the Gods throw awful combinations together. And that's without throwing in external factors like contractors cutting cables, builders on site, etc. And it also depends on the person/people supporting it and the hours they're being paid for. They're entitled to sleep, holiday, sickness and only to work when they're being paid to do so. Edited August 21, 2017 by elsiegee40 1
pcstru Posted August 21, 2017 Posted August 21, 2017 How long would you consider it to be acceptable to have loss of access to your SAN? Personal documents etc are all accessible but shared areas where the majority of project work etc are kept are not. I think that not only is there a "how long is a piece of string" element, but the term SAN is now quite wide ranging (i.e. abused!) - encompassing anything from a cheap NAS doing iSCSI or even these days directly running SAMBA for file shares, through to complex multi node systems which should be entirely immune to ANY downtime (yea, in theory). In a sense the question is the wrong way round, as the end user you are the one who can actually say what the value (loss) is to you of not having access to your data. That should be what drives the requirement for storage availability (how robust that 'SAN' needs to be) and dictating the Service Level whoever is providing the service should meet. At the minimum you should come away from an actual problem knowing either that you can survive such an outage or that you need to talk to the service provider about storing your stuff on a more robust system.
pete Posted August 21, 2017 Posted August 21, 2017 It also depends on how much whoever's paying for the SAN is (or was) willing to spend on service contracts, out-of-hours in-house support and a host of other things. Unfortunately, sometimes that money only gets spent after it hits the fan and $person_who_saved_money loses a lot of face.
hardtailstar Posted August 21, 2017 Posted August 21, 2017 Personally I think its acceptable to be down while it is being fixed. Whether it takes 30 minutes or a week. Ideally it would be better if it was fixed quickly but sometimes this doesn't happen. But in a company whilst it is down it is potentially losing the company money so there would be a difference of opinion there. Also as @elsiegee40 says it depends on the person fixing it and being paid etc.
googlemad Posted August 21, 2017 Posted August 21, 2017 If things are that critical the actual departments should have another method of working whilst the system is being repaired. Here in the NHS for example our core switch tripped several times due to another faulty switch recently which totally knocked out VMWare as the initial failure it tried to kick in HA, but then with the networking tripping several times in quick succession it confused itself as to where the VMs actually were and took us nearly a whole day to resolve as we had to contact VMWare etc. What was the solution for the affected departments? Good old pen and paper In many places I've seen it being a case of 'It's fine we spent £948475472340 on disaster recovery and failover' but never actually test it works so when it comes down to it their fancy solution is about as much use as a chocolate fireguard!
Steve21 Posted August 21, 2017 Posted August 21, 2017 Just repeating what was said really but depends a lot. Is this expected loss e.g. maintenance/migration etc? Then as long as it was planned for, whether minutes or days. You have to work on systems to do bits to them and unless you're rich enough to have redundant everything, there's times stuff goes down. Is it unexpected aka broken? Then who's job is it to fix it? What hours are they working? Is it major or minor? If as you mentioned it's not a dead SAN and "certain areas" it's not going to be something simple one way or another in terms of asking us how long Steve
Davit2005 Posted August 21, 2017 Posted August 21, 2017 With the suitable funding and resources it's possible to have highly resilient systems but unfortunately in education funding is often squeezed this all has a knock on effect in lots of area's.
SamanthaD Posted August 21, 2017 Author Posted August 21, 2017 This was an unexpected loss. It's just come back up, went down Thursday morning. Apparently multiple RAID issues and backups not copying correctly - sounds like a right headache tbf. Diagnosis was quick, but rectification was slow. Hopefully problems have been addressed and fixed as opposed to "get it working". PITA side is deadlines have been missed and this will cost us money. Thanks for the thoughts, now enjoy the rest of your week! :-)
Popular Post Vegas Posted August 21, 2017 Popular Post Posted August 21, 2017 You're lucky its not a big bell, apparently the downtime for a bell is 4 years! 6
hardtailstar Posted August 21, 2017 Posted August 21, 2017 This was an unexpected loss. It's just come back up, went down Thursday morning. Apparently multiple RAID issues and backups not copying correctly - sounds like a right headache tbf. Diagnosis was quick, but rectification was slow. Hopefully problems have been addressed and fixed as opposed to "get it working". PITA side is deadlines have been missed and this will cost us money. Thanks for the thoughts, now enjoy the rest of your week! :-) If it went down Thursday morning and up when you posted assuming the IT Department were only working week days then thats quite quick in my opinion. Even if they were working the weekend I'm guessing they would have wanted to get it working right rather than have multiple times in future where it would go down again. Its unfortunate its lost the company money but in the long run it would be a minimal amount in the financial year. Here's to hoping it is fixed for the long term. 2
arwen Posted August 22, 2017 Posted August 22, 2017 (edited) This is what SLA's are for... both IT and "users" agree what is acceptable and do-able. If management/users want a shorter SLA than IT can currently provide, money needs spent (either on redundancy or people) SLA's are much more than timers for jobs to provide stats. They are very useful in getting funding if used and implemented correctly. We lost one of our main SAN's for 2 and a half a day a a while ago due to massive RAID failure (multiple HD failures, battery died and RAID controller failed all within a few hours of each other). Said SAN was on 3rd party limited support due to being so old. It took 2 people about two days (two full days mind- they were working from home in the evenings) to fix/replace/restore everything to working order. Once a full report was given for the down time, we suddenly got money for new SAN's with full support from the manufacture and better redundancy. Edited August 22, 2017 by arwen
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now