paulellam Posted October 10, 2016 Posted October 10, 2016 Not naming which of the big brands and it did not happen at my school, but..... so called 'experienced' engineer, walks into school, pulls a drive out and replaces with a newe one while system is live, then waits ten minutes pulls out another and replaces it with a new one on a a Raid 5, wonders why its all collapses..... he did all this then calmly walks away!
Oaktech Posted October 10, 2016 Posted October 10, 2016 Wow... He'd be on my 'not on my site' blacklist.
zag Posted October 10, 2016 Posted October 10, 2016 I still don't understand people who have these kind of service contracts with Dell or HP or whoever. I've always bought Dell servers with 1 year warranties and upgraded the disks myself. In the days of virtualization, its easier to just migrate a VM to another host then fix things offline. I would never let someone else who doesn't know the system touch them!
Michael Posted October 10, 2016 Posted October 10, 2016 If only it took 10 minutes to replicate! It can take hours depending on the volume of information. You'd think that would be the first thing you'd question - RAID type and which disk is faulty, but also whether it's hot swap. Not all are!
Popular Post KevinH Posted October 10, 2016 Popular Post Posted October 10, 2016 "Good Morning, mate, I'm here to test your disaster recovery plan" 11
Michael Posted October 10, 2016 Posted October 10, 2016 What doesn't make sense is why he/she removed/change two disks. If 2 x disk in a RAID5 config were faulty, you'd lose the array anyway. But also why were they able to do this without you there? It's something I'd be very wary of!
paulellam Posted October 10, 2016 Author Posted October 10, 2016 In the techie at the schools defence, he is quite new to the role and expected the 'branded' engineer to know his stuff and the engineer just did what he did, with no time to question it. Not a great experience for the techie but gets the worst one out the way early in the career.
Michael Posted October 10, 2016 Posted October 10, 2016 In the techie at the schools defence, he is quite new to the role and expected the 'branded' engineer to know his stuff and the engineer just did what he did, with no time to question it. Not a great experience for the techie but gets the worst one out the way early in the career. I agree, yet I'd expect the engineer to understand 'RAID' at least the most common forms.
paulellam Posted October 10, 2016 Author Posted October 10, 2016 I agree, yet I'd expect the engineer to understand 'RAID' at least the most common forms. Amazes me that an engineer would not know that, I spent a few years being a 'disaster recovery' engineer (a long time ago though) and RAID should not be taken lightly, especially on somebody elses system. Luckily the school had backups etc, but the hassle it caused was huge.
Michael Posted October 10, 2016 Posted October 10, 2016 It's one of the first things I educate about the differences between servers and workstations to apprentices - RAID, what it is and the many types available. I can imagine the hassle was huge!
KevinH Posted October 10, 2016 Posted October 10, 2016 (edited) This sums it up Edited October 10, 2016 by KevinH
paulellam Posted October 11, 2016 Author Posted October 11, 2016 [ATTACH=CONFIG]39083[/ATTACH] This sums it up looool
Geoff Posted October 11, 2016 Posted October 11, 2016 We stopped using RAID 5 in about 2010 due to rebuild times taking too long (I think when we did the maths, 12Tb+ of data and we are guaranteed a second drive failure during the rebuild). On physical servers at a minimum data is RAID 6 with a hot spare and the OS is RAID 1 with a hot spare. SANs are raid 10 under the hood with 4+ hot spares. Disks are cheap, data is priceless.
Blue_Cookeh Posted October 11, 2016 Posted October 11, 2016 (edited) Even hot spares are getting a little questionable. With RAID rebuilds being the most stressful on healthy disks you want to make sure you have backups and do other bits and pieces before the machine starts rebuilding the array. I have to say though, we've only had good service from onsite engineers Dell and Lenovo have sent out to us. I wouldn't buy thousands of pounds of server without a minimum of 3 year NBD on site warranty. You try explaining why the £6000 server is now not working and out of warranty only 13 months in to management! Edited October 11, 2016 by Blue_Cookeh
zag Posted October 11, 2016 Posted October 11, 2016 Even hot spares are getting a little questionable. With RAID rebuilds being the most stressful on healthy disks you want to make sure you have backups and do other bits and pieces before the machine starts rebuilding the array. I have to say though, we've only had good service from onsite engineers Dell and Lenovo have sent out to us. I wouldn't buy thousands of pounds of server without a minimum of 3 year NBD on site warranty. You try explaining why the £6000 server is now not working and out of warranty only 13 months in to management! Is it really that hard to buy a spare hard disk and change it yourself though? In all my 15 years in this industry the only thing that has ever gone wrong with a server is a dead hard disk, or RAID card.
Geoff Posted October 11, 2016 Posted October 11, 2016 (edited) Even hot spares are getting a little questionable. With RAID rebuilds being the most stressful on healthy disks you want to make sure you have backups and do other bits and pieces before the machine starts rebuilding the array. Well quite, our next SAN (less than 3 years away) is going to be clustered so we can get away from this raid business entirely. I'm keeping my eye on Ceph personally as that ticks all our boxes. But commercial SANS and Software solutions such as VSAN and Windows Storage Server will have to be contemplated too. The other option is to say 'Bugger it all' and go for full cloud hosting. The nature of the data we deal with makes this difficult though (I don't work in a school any more). Edited October 11, 2016 by Geoff
gshaw Posted October 11, 2016 Posted October 11, 2016 Is it really that hard to buy a spare hard disk and change it yourself though? In all my 15 years in this industry the only thing that has ever gone wrong with a server is a dead hard disk, or RAID card. Never had a power supply pop?
zag Posted October 11, 2016 Posted October 11, 2016 Never had a power supply pop? I must say I have not, we do have good UPS systems though and use Powershute to auto shutdown when there have been power outages. Another thing we do is always buy the same servers, they have been very reliable over the years.
localzuk Posted October 11, 2016 Posted October 11, 2016 Is it really that hard to buy a spare hard disk and change it yourself though? In all my 15 years in this industry the only thing that has ever gone wrong with a server is a dead hard disk, or RAID card. You've been lucky We've had PSUs die, Motherboards die, even the backplane for the drive bays. Warranties have been invaluable.
gshaw Posted October 11, 2016 Posted October 11, 2016 I must say I have not, we do have good UPS systems though and use Powershute to auto shutdown when there have been power outages. Another thing we do is always buy the same servers, they have been very reliable over the years. You've been very lucky then We run one PSU from mains and one from UPS otherwise our UPS costs would be through the roof. Power at places I've worked has been ropey at best, lots of spikes and blips that pop PSUs from time to time. One went on our firewall last week despite being on a clean UPS supply.
KevinH Posted October 11, 2016 Posted October 11, 2016 We stopped using RAID 5 in about 2010 due to rebuild times taking too long (I think when we did the maths, 12Tb+ of data and we are guaranteed a second drive failure during the rebuild). On physical servers at a minimum data is RAID 6 with a hot spare and the OS is RAID 1 with a hot spare. SANs are raid 10 under the hood with 4+ hot spares. Disks are cheap, data is priceless.Ditto this statement, at least on spinning rust and gigunda drives. The bathtub curve is not your friend here ... The probability of drives from the same batch, with the same number of hours on them, failing close to one another suggests that the odds of a URE at the wrong time are just too high to chance.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now