Jump to content

Recommended Posts

Posted

The server is a Dell PowerEdge T110 II

2x SAS WD9002BKTG disks in RAID 1 on LSI controller

 

I've been called in to help a school that I don't regularly support so I've 'inherited' the following scenario/setup.

 

The server is a 2012 R2 HyperVhost and it hosts a single guest, 2008 R2 as a curriculum server.

First signs of weirdness began to show on the Guest OS - sluggish, services failing, VPN failing...

After spending a few hours checking what was wrong with the Guest I began to see the same sluggishness with the Host. It wouldn't even shutdown normally.

When the Server rebooted I was greeted with a message from the LSI controller to say that there was an error with the RAID array.

 

Entering the LSI config I noticed that the RAID was Degraded and the Slot 0 disk was Rebuilding, 26% synced. (Pred Fail = Yes)

Hmmm, it synced 10% in approx 3-4 hours.

 

I'm assuming that Slot 0 disk is the duff one, but I need help to know what order to go about swapping out this disk and inserting the replacement, or any other action you think would improve things.

Teachers are being very patient but that's dwindling. I need to give them an accurate idea of when the server will be back online.

 

At my estimate, if I didn't change anything and fired up the H-Vhost and then the Guest, then that would almost certainly be tomorrow AM at the earliest.

However, I thinking this disk needs to be replaced...

 

Whilst the RAID is syncing, I'm trying to extract their data from a Windows Backup Image (.vhdx) file.

 

Any help would be most appreciated.

Posted

Not a stupid question, that's a very good one. Yes, the warranty is still live (Phew!), and Dell should be sending us a replacement disk hopefully by tomorrow or Monday at the latest.

 

I'm just reading around the net and it seems that to rebuild the array I have to delete the current array and create a new one, making sure I have the new disk in and the possibly faulty disk taken out. If for whatever reason this whole thing goes south, I'm just hoping that the backup will work as expected.

 

Thanks for your reply.

Posted
id be surprised if you needed to delete the array when ive had disks fail in the past ive just pulled the dead one and replaced it with a replacement drive and left the raid software to sync the drive it should be pretty much automatic. If its struggling to rebuild to an iffy disk personally id be tempted to remove it so it stops trying
Posted
You might find the good disk dies if it is having problems syncing with the faulty disk but can't, I am no raid expert but if it was me I would pull the faulty one out, just make sure you check if its hot swap or not first.
Posted

Thanks 3s-gtech, Sted

 

I'm inclined to agree with you both, especially as the info I'm reading up (Dell Community posts) isn't the LSI controller.

How to rebuild Raid1 on Poweredge T110 II - PowerEdge OS Forum - Servers - Dell Community

 

I was wondering whether to stop the build process, but a nervous twitch kicks in when I look at that menu option.

As it turns out, they have a Belkin kvm switch attached which no longer registers the keyboard or mouse now I've switched back to the server.

Question, if I reboot the server after I've plugged in the KVM directly, will I upset the RAID building?

 

Many thanks

Posted (edited)

Hot-swapping a non hot-swappable disk is bad, but I know for a fact that I've done it :laugh:

 

along with some RAM :doh:

 

I recommend neither. I do agree that pulling the dead disk should help with the reliability of the array, while you wait for the new one.

 

Question, if I reboot the server after I've plugged in the KVM directly, will I upset the RAID building?

 

Pretty much (it will usually start again), the good disk should be unaffected by this but the duff one may have a fit.

Edited by 3s-gtech
Posted (edited)
Thanks 3s-gtech, Sted

 

I'm inclined to agree with you both, especially as the info I'm reading up (Dell Community posts) isn't the LSI controller.

How to rebuild Raid1 on Poweredge T110 II - PowerEdge OS Forum - Servers - Dell Community

 

I was wondering whether to stop the build process, but a nervous twitch kicks in when I look at that menu option.

As it turns out, they have a Belkin kvm switch attached which no longer registers the keyboard or mouse now I've switched back to the server.

Question, if I reboot the server after I've plugged in the KVM directly, will I upset the RAID building?

 

Many thanks

honestly atm the kvm isnt something id spend any time on just plug a keyboard and mouse into a usb port and fix it later

Hot-swapping a non hot-swappable disk is bad, but I know for a fact that I've done it

 

 

 

along with some RAM

 

 

 

I recommend neither. I do agree that pulling the dead disk should help with the reliability of the array, while you wait for the new one.

 

 

i did once accidently hot swap a video card it was a testbed so basically a motherboard mounted to a piece of wood it had gone to sleep so no fans or owt unplugged the video card and plugged in a new one to test and it woke up lol

Edited by sted
Posted
You might find the good disk dies if it is having problems syncing with the faulty disk but can't, I am no raid expert but if it was me I would pull the faulty one out, just make sure you check if its hot swap or not first.

 

That's a fair point. Makes me want to stop this rebuild even more now.

 

I had to come away from the school for a couple of hours but I'm on my way back now. The sync should be in the 40% range. Because I can't move the options in the menu, is it okay to reboot the server - I won't hold anyone responsible :)

Posted
That's a fair point. Makes me want to stop this rebuild even more now.

 

I had to come away from the school for a couple of hours but I'm on my way back now. The sync should be in the 40% range. Because I can't move the options in the menu, is it okay to reboot the server - I won't hold anyone responsible :)

again in an ideal world i wouldnt reboot the server but it should survive it again based on my experience

Posted

Don't go deleting the array or you will be testing that backup way sooner than you wanted to!

 

The first thing you need to find out is if there's a legitimate hardware raid controller in the machine or if it's one of the cheapo software raid (a.k.a "FakeRaid") setups. The recovery procedure is going to depend on the answer to that question.

Posted
Hot-swapping a non hot-swappable disk is bad, but I know for a fact that I've done it :laugh:

 

along with some RAM :doh:

 

I recommend neither. I do agree that pulling the dead disk should help with the reliability of the array, while you wait for the new one.

 

 

 

Pretty much (it will usually start again), the good disk should be unaffected by this but the duff one may have a fit.

 

I think that pretty much makes up my mind now, but I'll just make sure I have all their data available (in progress - 70% extracted of 550 GB)

 

Thanks again everyone

Posted

Thanks Kevin

 

I'll have more info on the controller when I return to the school - approx 45 minutes.

I know the make is LSI, what else should I look for - a physical card, or does this tower have an onboard controller?

Posted

Just a thought, is there anything I can do to improve the reliability of this array? Possibly add a 3rd disk as a hot spare, although the school would have to purchase a new disk.

 

If that is a good idea, are there restrictions surrounding what disks I can add?

The current disks are Dell WD9002BKTG 900GB, can I add a 1TB Seagate Constellation disk, as listed here.

https://www.scan.co.uk/products/1tb-seagate-constellation-st91000640ss-7200rpm-sas-6gb-s-64mb-25

Posted
depends on the controller my experience of lsi controllers is they are pretty basic and raid 0/1 is probably all it supports so you couldnt say add a 3rd disk and make it raid 5 (even if it supports raid 5 it may not support conversion from raid1) and i doubt it would support raid 6/10
Posted (edited)

As the others have said, definately do not delete the array (unless you have a full backup!).

A new disk should simply rebuild itself (you may have to access software/webmin to do this once you have put the new disk in).

I suspect you had read the instructions on rebuilding an array, not rebuilding a new disk.

 

Edited to add - Ah, I hadn't noticed you only had two disks in there! I hope you have backups then.

Edited by Dos_Box
Posted

I agree with the above, don't delete the array. This will wipe out the configuration and the good disk (and its contents).

 

Simply remove the faulty drive and add it to the existing array when prompted. The server will then re-build.

Posted
As it's on hyper-v, can you grab a backup using the free version of veeam? Then if things go south with with the RAID, you have a reasonably quick way to get it back up and running.
Posted
Check the event logs on the host - if you're getting a ton of disk errors, the file system on the good disk might have been corrupted. We've done quite a few restores to raid 1's in the last year or so, and invariably the bad disk has corrupted the file system, requiring a restore from backup.
Posted

Thank you all very much for your helpful posts.

 

In essence, I've shutdown the server - 'rebuilding' was at 51%. Had a call to say that Dell are delivering, and hopefully fitting the disk tomorrow. Yay!

Point taken about the disk corruptions, I think I'm going to prepare as if this system has been affected and use the backup if I don't see signs that all is well (event logs...)

 

Funnily enough, the point about the two disks was raised by the non-tech head. I tactfully said that it would have been a decision when you bought the server, but assured him that I'll look into reducing the possibility of this happening again - add a Hot Spare, or Backup, add extra disks, reconfig RAID and restore...

 

I'm just hoping that the good disk doesn't pack up during the resync, which I'm guessing is going to be the best part of 24 hours :nerd:

Posted
900GB Disk sounds like 10k or 15k SAS, on a good drive, I wouldn't expect it to take 24hrs (depending on usage of the server during that time) - your slow rebuild was potentially due to a failing disk. In terms of options moving forwards, as it's a vhost, I'd advise against RAID5/6 as the performance would not be good and look to raid 10 if possible, or stick with raid 1 and hotspare it.
Posted

Thanks Willott

 

Still waiting for the Dell Engineer to arrive. I think you're right about the slowness. Can you think of any reasons why the controller decided to rebuild the failing drive?

I don't know how long the server had been trying to indicate the failing drive, I'll perhaps know more when the OS is back up and running.

 

I like the RAID 10, and RAID 1+spare idea. I'm assuming I can add a spare that isn't an identical disk as long as the capacity is at least 900GB.

I found this drive online, but the speed is slower 7500rpm.

https://www.scan.co.uk/products/1tb-seagate-constellation-st91000640ss-7200rpm-sas-6gb-s-64mb-25

 

Would that work, or should I look at the higher priced disks?

Posted

 

Thank you, I'll take a look at that asap.

New disk now installed, Dell engineer simply took out the duff disk, replaced it and went back into the bios LSI config - the disk just started rebuilding all by itself, exactly as others had stated above (thanks for the warning about not removing the array).

 

By the way, I saw that the RAID adapter is in fact a PERC H200A.

It is indeed rebuilding much faster than previously - should be done within 4 hours.

 

I'll keep you posted

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...