Jump to content

Recommended Posts

Posted

We have 8 drives in our D2D2T unit- SATA - 3.5 inch Enterprise - 7.2K with LSI 92608i controller (latest firmware/drivers etc)

 

Originally they were Toshiba 2TB drives and after two years they started to become really problematic - usually the error reading seemed to correct itself - but not before the backup process hiccupped and failed.

 

Tried replacing them under the 3 (or was it 5) year warranty these drives carry - but discovered that because they were sold as part of the appliance they were OEM - and only covered for one year.

 

Running really hot to touch - despite fans drawing air over them.

 

Anyhow - bit the bullet and replaced with 4TB DELL Enterprise disks -(Made by Seagate) supposedly new. All ran nice and cool - and problem free until....

 

 

Aain 2 years later - I am back with backups failing and "recovered" sector errors on one drive. Bought a new drive - but before I could replace it a different drive actually failed. Replaced failed drive, rebuilt OK. And ordered another drive and replaced the one originally giving problems...and now...a third drive produced some "recovered" sector errors. And unsurprisingly I now have a bad block - so will need to reformat - or something. Aghh.....

 

Again drives feel hot to touch.

 

Am I just unlucky? Apparently unless the drive actually fails I can't replace it under warranty because some remapping of sectors is expected (frankly I start getting really nervous when this happens ...). And why do all drives run fine for two years and then within the space of a week or so do several start to look as if they are going to fail? In fact the drives have no warranty - as although I bought them new and packaged - they had been sold as OEM drives (how would I have known this?)

Posted
When drives are raided, they encounter similar wear, this means there is a higher chance of them failing at a similar time
Posted
Yes...but 3 ...in a same week....after perfect performance for almost two years...I would be thinking maybe a mains electric blip...although the server is fed from a UPS....possibly warmer weather....software update being less tolerant....
Posted

Did you buy all of the hard disks at the same time from the same supplier (like most of us do) if so then on average they will all have been thrown around a couriers van the same amount of time and had roughly the same wear and tear as each other so it doesnt surprise me that multiple disks are failing are approx the same time.

 

I'm not sure what's recommended when buying disks you're going to RAID, should they all come from the same batch or should you order several from different people (ensuring they hold them in stock in their own warehouses as if they're all getting them from the same wholesaler then they're going to be from the same batch anyway) or should we even be buying the same size drives from a couple of different manufacturers so multiple disks in the future wont all be affected by any possible defects/issues that only appear in 3,4 or 5+ years time?

Posted
In fact the drives have no warranty - as although I bought them new and packaged - they had been sold as OEM drives (how would I have known this?)

These are the replacement drives? If so, I'd speak to whoever your supplier is about that. I try to be a little careful with server discs in terms of sourcing them from suppliers we know and have a continuing relationship with but we have never had any issues with manufacturer warranty.

 

Drives failing at the same time might indicate : environmental problems, manufacturing/design issues or dodgy firmware. The former would be my suspect if they are running so hot but it could be a combination of the first two (elevated block failure can then cause a lot of seek head movement essentially due to fragmentation and multiple retries as blocks fail).

Posted

As the HDDs have mechanical parts they will always eventually fail. To respond with a definitive answer on why they fail within a certain time frame, there are too many variables to give an accurate answer. BUT there are certain steps you can take to protect yourself, ensure the environment they work within is ideal - correct Air Con, No Vibration, etc.

 

Expect all hardware you rely on to fail once out of warranty, that way you will take NO RISK. Only purchase hardware with acceptable warranty, some companies will also provide extended warranty when the manufacturer will not (Great for Servers). Finally get yourself a good supplier you trust and a single point of contact ideally. The problem with purchasing stuff from different suppliers or even worse as cheap as possible off the net, this doesn't provide you with the singe point of contact when things go wrong.

Posted
Oxidisation of the contacts on the PCB is far more common than internal hard drive failure. It results in all sorts of problems and is easily fixed with a pencil eraser. Worth taking a look on your failed drives before checking your RAID array I reckon.
  • 3 weeks later...
Posted
When drives are raided, they encounter similar wear, this means there is a higher chance of them failing at a similar time

 

I don't agree that drives tend to fail when running on RAID. I worked for LSI for almost 2 full years and there is nothing in the RAID controller that makes the sectors bad. The RAID controller just optimizes the read or write from/to the HDDbut don't harness to make the HDD survive less.

The problem may lie on the hard disks or it could be the continuous read/write usage that makes the magnetic storage disks wear out so early.

 

Coming to the original problem:

The drives feel hot to touch... there may be a serious problem here. This could be because of heavy usage.

Posted
Normally - Hard Disk Temperatures normally seem warm...but when a disk has failed - I have noticed that all the disks often seem hot. I am beginning to suspect that when there is a disk error - the raid controller/disk drive strive for a while to recover/re-read sectors from across the stripe...eventually this sector is marked as bad of course - and often there are a cluster of sectors that get remapped at once. I have wondered if this atypical activity on the RAID controller affects timings and triggers (possibly a false) alert than sectors are failing on another drive or two. It just seems "spooky" to me that several disk seem to fail in a relatively short period...and I have seen this a number of times....but perhaps its a power glitch....or something (although we use a UPS).

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...