Jump to content

Dont let an engineer pull two drives on a raid 5!!!


Recommended Posts

Posted
Not naming which of the big brands and it did not happen at my school, but..... so called 'experienced' engineer, walks into school, pulls a drive out and replaces with a newe one while system is live, then waits ten minutes pulls out another and replaces it with a new one on a a Raid 5, wonders why its all collapses..... he did all this then calmly walks away!
Posted

I still don't understand people who have these kind of service contracts with Dell or HP or whoever.

 

I've always bought Dell servers with 1 year warranties and upgraded the disks myself.

 

In the days of virtualization, its easier to just migrate a VM to another host then fix things offline.

 

I would never let someone else who doesn't know the system touch them!

Posted

If only it took 10 minutes to replicate! :D It can take hours depending on the volume of information.

 

You'd think that would be the first thing you'd question - RAID type and which disk is faulty, but also whether it's hot swap. Not all are!

Posted

What doesn't make sense is why he/she removed/change two disks. If 2 x disk in a RAID5 config were faulty, you'd lose the array anyway.

 

But also why were they able to do this without you there? It's something I'd be very wary of!

Posted
In the techie at the schools defence, he is quite new to the role and expected the 'branded' engineer to know his stuff and the engineer just did what he did, with no time to question it. Not a great experience for the techie but gets the worst one out the way early in the career.
Posted
In the techie at the schools defence, he is quite new to the role and expected the 'branded' engineer to know his stuff and the engineer just did what he did, with no time to question it. Not a great experience for the techie but gets the worst one out the way early in the career.

 

I agree, yet I'd expect the engineer to understand 'RAID' at least the most common forms.

Posted
I agree, yet I'd expect the engineer to understand 'RAID' at least the most common forms.

 

Amazes me that an engineer would not know that, I spent a few years being a 'disaster recovery' engineer (a long time ago though) and RAID should not be taken lightly, especially on somebody elses system. Luckily the school had backups etc, but the hassle it caused was huge.

Posted
It's one of the first things I educate about the differences between servers and workstations to apprentices - RAID, what it is and the many types available. I can imagine the hassle was huge!
Posted

We stopped using RAID 5 in about 2010 due to rebuild times taking too long (I think when we did the maths, 12Tb+ of data and we are guaranteed a second drive failure during the rebuild). On physical servers at a minimum data is RAID 6 with a hot spare and the OS is RAID 1 with a hot spare. SANs are raid 10 under the hood with 4+ hot spares.

 

Disks are cheap, data is priceless.

Posted (edited)

Even hot spares are getting a little questionable. With RAID rebuilds being the most stressful on healthy disks you want to make sure you have backups and do other bits and pieces before the machine starts rebuilding the array.

 

I have to say though, we've only had good service from onsite engineers Dell and Lenovo have sent out to us. I wouldn't buy thousands of pounds of server without a minimum of 3 year NBD on site warranty. You try explaining why the £6000 server is now not working and out of warranty only 13 months in to management!

Edited by Blue_Cookeh
Posted
Even hot spares are getting a little questionable. With RAID rebuilds being the most stressful on healthy disks you want to make sure you have backups and do other bits and pieces before the machine starts rebuilding the array.

 

I have to say though, we've only had good service from onsite engineers Dell and Lenovo have sent out to us. I wouldn't buy thousands of pounds of server without a minimum of 3 year NBD on site warranty. You try explaining why the £6000 server is now not working and out of warranty only 13 months in to management!

 

Is it really that hard to buy a spare hard disk and change it yourself though?

 

In all my 15 years in this industry the only thing that has ever gone wrong with a server is a dead hard disk, or RAID card.

Posted (edited)
Even hot spares are getting a little questionable. With RAID rebuilds being the most stressful on healthy disks you want to make sure you have backups and do other bits and pieces before the machine starts rebuilding the array.

 

Well quite, our next SAN (less than 3 years away) is going to be clustered so we can get away from this raid business entirely. I'm keeping my eye on Ceph personally as that ticks all our boxes. But commercial SANS and Software solutions such as VSAN and Windows Storage Server will have to be contemplated too.

 

The other option is to say 'Bugger it all' and go for full cloud hosting. The nature of the data we deal with makes this difficult though (I don't work in a school any more).

Edited by Geoff
Posted
Is it really that hard to buy a spare hard disk and change it yourself though?

 

In all my 15 years in this industry the only thing that has ever gone wrong with a server is a dead hard disk, or RAID card.

 

Never had a power supply pop?

Posted
Never had a power supply pop?

 

I must say I have not, we do have good UPS systems though and use Powershute to auto shutdown when there have been power outages.

 

Another thing we do is always buy the same servers, they have been very reliable over the years.

Posted
Is it really that hard to buy a spare hard disk and change it yourself though?

 

In all my 15 years in this industry the only thing that has ever gone wrong with a server is a dead hard disk, or RAID card.

 

You've been lucky

 

We've had PSUs die, Motherboards die, even the backplane for the drive bays. Warranties have been invaluable.

Posted
I must say I have not, we do have good UPS systems though and use Powershute to auto shutdown when there have been power outages.

 

Another thing we do is always buy the same servers, they have been very reliable over the years.

 

You've been very lucky then :p We run one PSU from mains and one from UPS otherwise our UPS costs would be through the roof. Power at places I've worked has been ropey at best, lots of spikes and blips that pop PSUs from time to time. One went on our firewall last week despite being on a clean UPS supply.

Posted
We stopped using RAID 5 in about 2010 due to rebuild times taking too long (I think when we did the maths, 12Tb+ of data and we are guaranteed a second drive failure during the rebuild). On physical servers at a minimum data is RAID 6 with a hot spare and the OS is RAID 1 with a hot spare. SANs are raid 10 under the hood with 4+ hot spares.

 

Disks are cheap, data is priceless.

Ditto this statement, at least on spinning rust and gigunda drives. The bathtub curve is not your friend here ... The probability of drives from the same batch, with the same number of hours on them, failing close to one another suggests that the odds of a URE at the wrong time are just too high to chance.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...