tmcd35 Posted June 29, 2009 Posted June 29, 2009 Hi guys, I'm in desperate need of some ideas on how to fix a dying server. Story so far... We have 3 x HP ProLient DL380 G4 Rackmount servers One of the servers is randomly restarting every 20 minutes or so with an unknown hardware fault. We've caught it once on a blue screen saying just that "Hardware fault" (no other info) before ASR kicked in a rebooted the server. When It started happening I used an Ultimate Boot CD to run a CPU stress test and MemTest86+. Both came back fine, so I started looking at hard drives. The controller reports everything AOK and no probs with any of the drives. Server under warrenty so I've been in touch with HP and after a week of getting no where they agreed to replace the entire motherboard (the RAID controller is integrated). A week later and it's started doing it again, rebooting every 20 minutes. HP error logs show the event - unkown hardware fault caused ASR to reboot unexpectedly. I have brand new RAM which I'm putting in today as the servers where due for a RAM upgrade anyway. So I can be pretty sure the problem is not RAM (if it happens again) or motherboard/RAID controller related. Two questions: Does anyone know of a really good CPU stress test utility that test the CPU only (no activity on RAM/HDD)? Preferably one that can test 4-cores to destruction? How to work out which hard drive of 5 in an array is faulty when none of the usual diagnostic tools or warning lights are showing any suspected fault even occurs? And maybe (3) anywhere else I can try looking to that I haven't considered yet? Thanks in advance Terry.
Ex-MGSTech Posted June 29, 2009 Posted June 29, 2009 Random re-boots have been attributed to power supplies in the past. Do you have redundant PSU's in the server? if so try running on one and see if it still does it, then if not try with the other one out. As long as you ahve one connected you can disconnect the other one (you'll get an amber warning light but that's about all) Steve 1
ricki Posted June 29, 2009 Posted June 29, 2009 Hi Have the server got an ups on it. I have had a problem with old old ups and also if they are over heating. Richard
ricki Posted June 29, 2009 Posted June 29, 2009 Hi Have a look at Help! My PC keeps rebooting every 10 to 20 minutes! - CNET Community Newsletter: Q&A Forums this might help too. Richard 1
tmcd35 Posted June 29, 2009 Author Posted June 29, 2009 (edited) Update: It's not the mobo or raid controller as they were replaced by HP and It's not the RAM as I've just put in 4 new sticks and got the same result. Off to do a 'Prime95' CPU stress test. If it's not that then it's either a randomly faulty NIC or on of the HDD's. edit: also it's not over heating - I've got half the school fighting me for access to the server room it soo cold in there (and so very hot everywhere else!). I find it hard to blame the UPS as 3 other servers are connected into that beast. Surely if that was causing the prob other servers would have been effected? Edited June 29, 2009 by tmcd35
AdamGent Posted June 29, 2009 Posted June 29, 2009 Try and find a copy of virtualPC online, this is a full hardware test program - boot from the CD and run a full system test - A colleague of mine uses it to test refurbished pcs and it often finds faults that dont always show up otherwise. 1
plexer Posted June 29, 2009 Posted June 29, 2009 Have you ruled out the servers own psu's though as previously mentioned? Ben
tmcd35 Posted June 29, 2009 Author Posted June 29, 2009 actually, no I haven't! I'll get on to those tests straight away. Didn't think of the PSU I've run 'Prime95' for around 10 min with 8 thread CPU stress only followed by 20+ min with 8 thread CPU and RAM stress - nearly 2Gb or Ram was committed. Server stayed solid through both tests. I'm more convinced the problem is not mobo/raid/ram/cpu related. I'll try the PSU check next. I still find it suspicious that it mostly does it at the start of lessons - when everyone is logging on. I need a good NIC/HDD stress tester. I'll try virtualPC after the PSU tests. Cheers guys. Please keep the ideas coming...
bossman Posted June 29, 2009 Posted June 29, 2009 @tmcd35: What temp are the processors running at and do they have thermal cutout settings via the bios? Obviously as you say it seems when the server is fully loaded at log in or log off time that the problem occurs if the processors are heating up quickly without proper air flow around them then this could be the case.
tmcd35 Posted June 29, 2009 Author Posted June 29, 2009 @Bossman, thanks for the suggestion. I really don't think the problem is heat/processor related. Mainly because it should have rebooted during the prime95 tests. The load was at 100% across all 8 cores/threads (1xdual core with HT) for around half hour total - no issues. Of course neither the HDD's or the NICS were being accessed during the Prime'95 tests. My gut instinct (as it as all along) still says a faulty HDD - I just can't work out how to determine which of the 5 buggers is causing the problem. Also have 4 NICs total, It could be one of the three add-in cards. @AdamGent - looked up Virtual PC Check/PC Check - $300? Good suggestion but not something I can get any time soon. Anyone know of any good FOSS that'll do the same/similar job? @Plexar/@MGSTech - Currently running on one PSU - so far so good. Next mass logon is around 1:10pm. I'll switch PSU's at around 1:40pm ready for the last logon of the day (2:10ish). We'll see if the problem re-occurs during either periods? If it does I'll look at recreating the problem tomorrow.
DMcCoy Posted June 29, 2009 Posted June 29, 2009 What happens if you turn asr off? I've always found it to be more trouble than any potential use it could be.
tmcd35 Posted June 29, 2009 Author Posted June 29, 2009 If I turn ASR off then the server just freezes up but no reboot. I end up having to manually reboot the server. Only once have I seen a blue screen for an error which I think may have been thrown up by ASR before it restarted. The Windows event logs are not showing anything obvious at around the same time as the reboots
LeMarchand Posted June 29, 2009 Posted June 29, 2009 I has a similar problem, and it did turn out to be one of the HDDs. I used an HDD Regenerator program which took hours, but the server seems OK now. (Though I'd be happier if SMT would pony up for new disks or preferably a whole new server).
tmcd35 Posted June 29, 2009 Author Posted June 29, 2009 @LeMarchand - how did you work out which HDD was at fault? I *could* spend £1000 on 5 new HDD's and swap them out one at a time. Letting the array rebuild to the hot spare each time before pulling the next. But that could take days and I'm worried about what happens if the faulty drive gives during the rebuild process. I'd rather find a way of determining which drive is at fault and just replace that.
swgeek Posted June 29, 2009 Posted June 29, 2009 Afternoon all.. If it is raid any way can you not just remove drives on a one by one basis and if the server stops restarting you will know which drive it is! If it is a drive? Cheers Mark 1
LeMarchand Posted June 29, 2009 Posted June 29, 2009 Afternoon all.. If it is raid any way can you not just remove drives on a one by one basis and if the server stops restarting you will know which drive it is! If it is a drive? Cheers Mark That's what I did, but it was at the "I'll try anything to get it back up" stage and may not have been the best course of action.
matt40k Posted June 29, 2009 Posted June 29, 2009 Pay to get HP onsite 4 hrs response and get them to sort it. (Not sure how HP support compares to Dells)
tmcd35 Posted June 29, 2009 Author Posted June 29, 2009 Actually, there is an idea there! I've been to worried about the system rebuilding on the hot-spare to try it. Some times you just can't see the wood for the trees. Yes, I'll had it to the list of tests to do (after I've finished the PSU checks). If I pull the hot spare first, it can't rebuild the array. If I then pull the other drives one at a time I'll soon stumble across the culprit - if it is indeed a drive problem.
localzuk Posted June 29, 2009 Posted June 29, 2009 A faulty drive should not cause any reboots, as the whole point of RAID is to have redundancy - ie. if one drive dies, the other 4 should handle it fine. This definitely sounds like a PSU issue.
swgeek Posted June 29, 2009 Posted June 29, 2009 Thanks TMCD35... Cool that was my first post her too!
tmcd35 Posted June 29, 2009 Author Posted June 29, 2009 Okay, Whats the likelihood of both PSU's being faulty? Should I consider swapping them out, one at a time to the other two servers? See if the reboot problems follows both the PSU's? As It stands. I've tried running off of 1 PSU and then running off the other. Reboots occurred both times. So it's either not a PSU problem or both the PSU's are at fault...
Michael Posted June 29, 2009 Posted June 29, 2009 From all the tests and replacements you've made, I would have to agree that the probability of the fault being related to your RAID array is unlikely. All RAID controllers usually come with software so you can see various bits of information, including up-to-date status within Windows itself. If there was a fault I'm sure something would be picked up. If you're using your PSU near its limit, this could trigger a reboot in theory. Your server should be connected to a UPS and this also can give clues with regards to power problems.
tmcd35 Posted June 29, 2009 Author Posted June 29, 2009 Thanks TMCD35... Cool that was my first post her too! No prob! You made me re-think how I was looking at the problem - always worth a thank-you in my books
tmcd35 Posted June 29, 2009 Author Posted June 29, 2009 If you're using your PSU near its limit, this could trigger a reboot in theory. Your server should be connected to a UPS and this also can give clues with regards to power problems. Here in lies the problem. I have three identical servers plugged into one uber-UPS. I need to check if the PSU's are hot swappable. If they are I'm going to live swap them with another server. If it's a problem with the PSUs then the problem should follow them onto another server. As all three systems are configured identically, If one PSU was near it's limit, then surely the others would be as well?
Michael Posted June 29, 2009 Posted June 29, 2009 As all three systems are configured identically, If one PSU was near it's limit, then surely the others would be as well? Very true, but who knows, the PSUs might be different brands or have different power ratings. It is possible but generally speaking they should be the same you're right.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now