Jump to content

Recommended Posts

Posted

Hi guys,

 

I'm in desperate need of some ideas on how to fix a dying server. Story so far...

 

We have 3 x HP ProLient DL380 G4 Rackmount servers

 

One of the servers is randomly restarting every 20 minutes or so with an unknown hardware fault. We've caught it once on a blue screen saying just that "Hardware fault" (no other info) before ASR kicked in a rebooted the server.

 

When It started happening I used an Ultimate Boot CD to run a CPU stress test and MemTest86+. Both came back fine, so I started looking at hard drives. The controller reports everything AOK and no probs with any of the drives.

 

Server under warrenty so I've been in touch with HP and after a week of getting no where they agreed to replace the entire motherboard (the RAID controller is integrated).

 

A week later and it's started doing it again, rebooting every 20 minutes. HP error logs show the event - unkown hardware fault caused ASR to reboot unexpectedly.

 

I have brand new RAM which I'm putting in today as the servers where due for a RAM upgrade anyway.

 

So I can be pretty sure the problem is not RAM (if it happens again) or motherboard/RAID controller related.

 

Two questions:

 

  1. Does anyone know of a really good CPU stress test utility that test the CPU only (no activity on RAM/HDD)? Preferably one that can test 4-cores to destruction?
  2. How to work out which hard drive of 5 in an array is faulty when none of the usual diagnostic tools or warning lights are showing any suspected fault even occurs?

 

And maybe (3) anywhere else I can try looking to that I haven't considered yet?

 

Thanks in advance

 

Terry.

Posted

Random re-boots have been attributed to power supplies in the past. Do you have redundant PSU's in the server? if so try running on one and see if it still does it, then if not try with the other one out.

 

As long as you ahve one connected you can disconnect the other one (you'll get an amber warning light but that's about all)

 

Steve

  • Thanks 1
Posted

Hi

 

Have the server got an ups on it. I have had a problem with old old ups and also if they are over heating.

 

Richard

Posted (edited)

Update:

 

It's not the mobo or raid controller as they were replaced by HP

 

and

 

It's not the RAM as I've just put in 4 new sticks and got the same result.

 

Off to do a 'Prime95' CPU stress test. If it's not that then it's either a randomly faulty NIC or on of the HDD's.

 

edit: also it's not over heating - I've got half the school fighting me for access to the server room it soo cold in there (and so very hot everywhere else!). I find it hard to blame the UPS as 3 other servers are connected into that beast. Surely if that was causing the prob other servers would have been effected?

Edited by tmcd35
Posted
Try and find a copy of virtualPC online, this is a full hardware test program - boot from the CD and run a full system test - A colleague of mine uses it to test refurbished pcs and it often finds faults that dont always show up otherwise.
  • Thanks 1
Posted

actually, no I haven't! I'll get on to those tests straight away. Didn't think of the PSU :o

 

I've run 'Prime95' for around 10 min with 8 thread CPU stress only followed by 20+ min with 8 thread CPU and RAM stress - nearly 2Gb or Ram was committed. Server stayed solid through both tests.

 

I'm more convinced the problem is not mobo/raid/ram/cpu related.

 

I'll try the PSU check next.

 

I still find it suspicious that it mostly does it at the start of lessons - when everyone is logging on.

 

I need a good NIC/HDD stress tester.

 

I'll try virtualPC after the PSU tests.

 

Cheers guys. Please keep the ideas coming...

Posted

@tmcd35:

 

What temp are the processors running at and do they have thermal cutout settings via the bios?

Obviously as you say it seems when the server is fully loaded at log in or log off time that the problem occurs if the processors are heating up quickly without proper air flow around them then this could be the case. ;) :)

Posted

@Bossman, thanks for the suggestion. I really don't think the problem is heat/processor related. Mainly because it should have rebooted during the prime95 tests. The load was at 100% across all 8 cores/threads (1xdual core with HT) for around half hour total - no issues.

 

Of course neither the HDD's or the NICS were being accessed during the Prime'95 tests.

 

My gut instinct (as it as all along) still says a faulty HDD - I just can't work out how to determine which of the 5 buggers is causing the problem.

 

Also have 4 NICs total, It could be one of the three add-in cards.

 

@AdamGent - looked up Virtual PC Check/PC Check - $300? Good suggestion but not something I can get any time soon. Anyone know of any good FOSS that'll do the same/similar job?

 

@Plexar/@MGSTech - Currently running on one PSU - so far so good. Next mass logon is around 1:10pm. I'll switch PSU's at around 1:40pm ready for the last logon of the day (2:10ish). We'll see if the problem re-occurs during either periods? If it does I'll look at recreating the problem tomorrow.

Posted

If I turn ASR off then the server just freezes up but no reboot. I end up having to manually reboot the server. Only once have I seen a blue screen for an error which I think may have been thrown up by ASR before it restarted.

 

The Windows event logs are not showing anything obvious at around the same time as the reboots :(

Posted
I has a similar problem, and it did turn out to be one of the HDDs. I used an HDD Regenerator program which took hours, but the server seems OK now. (Though I'd be happier if SMT would pony up for new disks or preferably a whole new server).
Posted
@LeMarchand - how did you work out which HDD was at fault? I *could* spend £1000 on 5 new HDD's and swap them out one at a time. Letting the array rebuild to the hot spare each time before pulling the next. But that could take days and I'm worried about what happens if the faulty drive gives during the rebuild process. I'd rather find a way of determining which drive is at fault and just replace that.
Posted

Afternoon all..

If it is raid any way can you not just remove drives on a one by one basis and if the server stops restarting you will know which drive it is! If it is a drive?

Cheers Mark

  • Thanks 1
Posted
Afternoon all..

If it is raid any way can you not just remove drives on a one by one basis and if the server stops restarting you will know which drive it is! If it is a drive?

Cheers Mark

 

That's what I did, but it was at the "I'll try anything to get it back up" stage and may not have been the best course of action.

Posted

Actually, there is an idea there! I've been to worried about the system rebuilding on the hot-spare to try it. Some times you just can't see the wood for the trees. Yes, I'll had it to the list of tests to do (after I've finished the PSU checks).

 

If I pull the hot spare first, it can't rebuild the array. If I then pull the other drives one at a time I'll soon stumble across the culprit - if it is indeed a drive problem.

Posted

A faulty drive should not cause any reboots, as the whole point of RAID is to have redundancy - ie. if one drive dies, the other 4 should handle it fine.

 

This definitely sounds like a PSU issue.

Posted

Okay,

 

Whats the likelihood of both PSU's being faulty? Should I consider swapping them out, one at a time to the other two servers? See if the reboot problems follows both the PSU's?

 

As It stands. I've tried running off of 1 PSU and then running off the other. Reboots occurred both times. So it's either not a PSU problem or both the PSU's are at fault...

Posted

From all the tests and replacements you've made, I would have to agree that the probability of the fault being related to your RAID array is unlikely. All RAID controllers usually come with software so you can see various bits of information, including up-to-date status within Windows itself. If there was a fault I'm sure something would be picked up.

 

If you're using your PSU near its limit, this could trigger a reboot in theory. Your server should be connected to a UPS and this also can give clues with regards to power problems.

Posted
Thanks TMCD35... Cool that was my first post her too!:D

 

No prob! You made me re-think how I was looking at the problem - always worth a thank-you in my books ;)

Posted
If you're using your PSU near its limit, this could trigger a reboot in theory. Your server should be connected to a UPS and this also can give clues with regards to power problems.

 

Here in lies the problem. I have three identical servers plugged into one uber-UPS. I need to check if the PSU's are hot swappable. If they are I'm going to live swap them with another server. If it's a problem with the PSUs then the problem should follow them onto another server.

 

As all three systems are configured identically, If one PSU was near it's limit, then surely the others would be as well?

Posted
As all three systems are configured identically, If one PSU was near it's limit, then surely the others would be as well?

 

Very true, but who knows, the PSUs might be different brands or have different power ratings. It is possible but generally speaking they should be the same you're right.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...