beany1 Posted November 7, 2011 Posted November 7, 2011 Hi All, Have been experiencing some strange issues with one of our Servers a HP Proliant ML350 running 2008, it randomly switches off and starts up again. This isn't a clean shutdown and happens mainly at night, in fact its only happened twice or there about in the last 6 months during the day, but has happened 19 times during the night in the same period. Checking event viewer / system logs I'm left with "The previous system shutdown at 05:06:51 on 24/10/2011 was unexpected." looking at all other logs there is non that are even close to the shutdown times. After this, last Thursday during the night the server switched off but this time was stuck on a blank screen with "internal health problem" and "external (power supply) health problem" I was unaware of the internal health LEDs so didn't check but switching the UPS off and back on and the server booted fine. This is the first time its done this so as you can understand - quite worried! Checked temperature both with software and a temp probe. It seems to be fine, nothing over 40 degrees. All fans are running. Everything is working OK even when the server is under load. There is no scheduled tasks, but as I said before the times are completely random. I was thinking it could be a UPS problem but there is another server attached to this which doesn't reboot, and its just had a new battery. So any ideas what I could check or do? Cheers!
glennda Posted November 7, 2011 Posted November 7, 2011 are there any blue screen notices? event id 1001 Also i have had a few problems recent with HP servers and drivers - so worth updating all the drivers and if they are the latest version maybe roll back to the previous release and then see if it carries on.
beany1 Posted November 7, 2011 Author Posted November 7, 2011 Hi Glennda, Did a filter for 1001 in system logs all it showed was events for SNMP starting. I also tried for id 1003. I literally haven't touched this server for ages apart from making changes in GP, but I will try update drivers. Cheers!
IanT Posted November 14, 2011 Posted November 14, 2011 Is it fully windows updated, proliant service packed, firmware updated?
beany1 Posted November 17, 2011 Author Posted November 17, 2011 Cheers for the ideas. The reboots are really intermittent, like some have been a week apart. So I'm trying to do it slowly to try find out the problem. Still nothing showing in logs. Did a windows update on Friday and so far so good, but it could just be waiting!
kevin_lane Posted November 17, 2011 Posted November 17, 2011 what changes did you make for the gp also do you have windows automatically install windows updates
beany1 Posted November 17, 2011 Author Posted November 17, 2011 Sorry I was meaning I've only gone on that server to make some changes to GP for normal workstations or AD, I haven't touched the domain controller gpos. Auto updates were off, but it's now completely up to date. Not sure what to check next. Hopefully the windows updates fixed it?!
Dave84 Posted November 17, 2011 Posted November 17, 2011 Not familiar with the ML350, but does it come with ILO? the ILO management log should tell you if you have had a hardware issue or not.
plexer Posted November 17, 2011 Posted November 17, 2011 Have you done any hardware diagnostics on it disk scan, memtest? Ben
beany1 Posted November 17, 2011 Author Posted November 17, 2011 Dave - it does come with ILO, wasn't aware it could do that / never got round to setting it up so I will have a look at that tomorrow! Thanks Ben - I haven't no, the first time it appeared to be a hardware issue was the other week and I haven't been able to power down the server since then. I did upgrade the RAM last Christmas, and ran a memtest on the new and old ram and it was all fine. So I'm hoping that it hasn't only lasted 6months! I'll make sure that these are my next things to check though, hopefully ILO will indicate if there is problems. Cheers
beany1 Posted November 18, 2011 Author Posted November 18, 2011 Damn - Last night another restart. Literally the server sits there. I haven't made any changes to it. I've had remote desktop open to it, and literally checked its still their every 15mins like a mad man. I've checked ILO - I'm assuming you mean the ILO2 log on the System Status page? Informational iLO 2 11/17/2011 22:51 11/17/2011 22:51 1 Server power restored. Informational iLO 2 11/17/2011 22:51 11/17/2011 22:51 1 Server power removed. So this refers to the reboot Yesterday. The only log before this is: Informational iLO 2 11/11/2011 19:01 11/11/2011 19:01 1 Server power restored. Caution iLO 2 11/11/2011 19:01 11/11/2011 19:01 1 Server reset. Which is when I restarted the server for Windows Updates. In the IML the last entry is on the 6th: Caution POST Message 11/06/2011 05:56 11/06/2011 05:56 1 POST Error: 1778-Drive Array Resuming Automatic Data Recovery Process Which coincides with another crash. Everything in System information is OK. I've upgraded the ILO firmware to the latest. So does this mean I don't have hardware issues? Or could I still but their not registering?
plexer Posted November 18, 2011 Posted November 18, 2011 What about the PSU in the server itself? Ben
glennda Posted November 18, 2011 Posted November 18, 2011 Is it connected to a UPS? Could also be the UPS failing. Or is it directly connected into mains? Toby
beany1 Posted November 18, 2011 Author Posted November 18, 2011 Not sure how to check Plexer. According to ILO its OK, surely if it was on its way out, when the server is under load it would cut out? However the hardware lights (that have only happened once) did indicate internal problem and external problem. Apparently an external problem is the PSU. It is connected to a UPS - which has just got a new battery, according the software the UPS is fine. Running the tests it can keep the servers powered up. Connected to the same UPS is another server - which isn't rebooting so I scrapped the idea of it being the UPS??
SYNACK Posted November 18, 2011 Posted November 18, 2011 Could be RAM and dodgey PSU, other things to try: Install latest driver Install the latest firmware for all components (NICs, BIOS, Power managment controler, RAID controller firmware) you can use the firmware update CD from HP or do it manually in Windows with the HP downloads. Use the Insite diagnostics from the latest smartstart CD to run a memory test that allows for ECC RAM and other avalible tests. Ramp the CPU up to 100% and leave it there for a few hours (folding@home SMP is a good one for this) to check for CPU overheat/point overheating (areas of CPU not near the temp sensor overheating before the temp sensor registers it) Swap the PSU to the other PSU bay, swap the PSU with another one from another identical server.
beany1 Posted November 18, 2011 Author Posted November 18, 2011 Cheers for that Synack. So far I've updated all the firmware's. Updated the drivers. I'm going to run diagnostics like the insite ones or memtest ASAP. But as no one goes home I'm finding it difficult! I only have 1 PSU for this server but I will try it in the redundant bay. See if I can purchase a redundant one next week. As for the CPU I will try it but as I said previously the server is pretty heavily loaded at the moment and even at peek times it doesn't reboot / get too hot. The times it has rebooted is times in which I'm assuming the server is not doing anything. During the night, I don't have any scheduled tasks so I really don't think much will be going on. Currently letting it make another backup to a removable drive so I am sure I've got everything!
SYNACK Posted November 18, 2011 Posted November 18, 2011 Don't trust memtest on a server it is not designed to handle ECC ram so although it may pass it may be throwin lots of ECC errors behind the scenes that only very rarely cause a glitch despite considerable damage.
beany1 Posted November 18, 2011 Author Posted November 18, 2011 Cheers Synack, So on the insight dvd their should be memory diagnostics? Ill try that ASAP
m25man Posted November 18, 2011 Posted November 18, 2011 You need to disable ASR using the HP tools or within the BIOS! ASR is a nightmare if you dont manage it properly, I have seen where a NIC looses its internet connection for a while due to excessive traffic the ASR (HP's Automated Server Recovery) decides that the best course of action is to perform a reset! Dont get me wrong, this isnt a fix it's just to stop the server from randomly restarting whilst you find out why it does it! My first experience of this was actually caused by a Backup exec job on another server on the same switch. The BE job would run each night for around 2 hours during which time the Slam Dunking Server would decide that because there was too much latency on that NIC it would do an ASR. Another was due to a known issue with HP Power Supplies and APC UPS's the APC UPS would go into Brown out or Black out due to an over or under voltage state, the APC kicks out a nasty Sawtooth or Square Wave rather than a nice smooth Sine Wave the HP Power Supplies complain bitterly about this and the good old ASR kicks in and reboots the server. The OS doesnt have a clue whats happening and all you will find in the windows logs is the restart was unscheduled! Google HP ASR Reboot - there is your homework for the evening....
beany1 Posted November 19, 2011 Author Posted November 19, 2011 (edited) Cheers for that m25man I will have a look at that ASAP on the server. Surely HP could make it so that the server shutdown gracefully, or at least attempt to then force a reset. sounds like a good feature just implemented in a strange way. Hopefully this will be the reason. Looking at some posts though it implies that a log was created about ASR in ILO logs? Is this not always the case? Edited November 19, 2011 by beany1
SYNACK Posted November 19, 2011 Posted November 19, 2011 The ASR is kind of like an advanced watchdog timer. If the OS is not responding to that it is almost certainly crashed hence no need to try a shutdown. Do you have all the HP services installed on it to report stuff correctly back to the ASR system. You are right though, it should log it in the IML.
mtillbrook Posted November 20, 2011 Posted November 20, 2011 the only time ive seen symptoms like that before was from a memory fault. We put some new memory in a server and it ran quite happily for about a month. For some reason after this, the server would shutdown and restart itself randomly - usually at night. It turns out the memory had a fault which, according to Crucial, was probably caused by an ESD resulting from proper ESD precautions not being taken when the memory was installed! This can take any amount of time to start causing problems from straight away to a few years down the road!
beany1 Posted November 21, 2011 Author Posted November 21, 2011 OK so an update... Using the bios utility MEMBIST - the status of the four sticks of RAM is OK. I then ran the HP insight diagnostics, the initial summary page displays all as fine, however running a custom test on Total memory, which tests 8 things, all of which pass apart from - ECC Test In the Error Log under the ECC Test description is 'Correctable ECC Events logging limit reached in SEL log Device. Ran on CPU 0' in the recommended repair 'Please refer IPMI Sensor Event Log for ECC events' Error code 021279. I read that this error can be because the RAM isn't seated properly so I removed them and reset them. I re ran the test with still a fail, and tried running the test with just the original RAM in and then the newer upgrade RAM, all the times failing the ECC test. I disabled ASR by entering bios then going to server availability followed by ASR status, I exited bios and booted the server, however it rebooted about 6 times, I couldn't even login before it reset. So I've had to enable it again. I tried the PSU in the redundant bay but that didn't make any difference. So is this all being caused by the RAM? the 2 original HP sticks as well as the new crucial / corsair ( can't remember which) both of which appear to have ECC problems? I assumed this type of problem could only occur from power shorts, but the server is going through an apc ups which apparently regulates voltage? As for ESD I don't see how I could of, I always make sure I ground my self, and I don't touch contacts etc, and the HP sticks of RAM had not been touched and they failed the ECC test. Still no error lights or any other indication of a problem. Help!
SYNACK Posted November 21, 2011 Posted November 21, 2011 You may need to try and clear the ECC log if there is an option for it. If there are too many ECC errors the system decides that the RAM is not trstworthy and reboots, sometimes disabling the faulty stick. As to the RAM itself, some RAM is just junk from the day it leaves the factory, you may have been unluck in the brand or the batch that you purchased.
beany1 Posted November 22, 2011 Author Posted November 22, 2011 Do you reckon if I clear the ECC log that the results of the test will be different? I really don't see how I could be so unlucky that both the original RAM and the upgraded RAM from crucial/corsair could also be faulty. Their is a matching pair of HP RAM - factory installed, which haven't been moved till yesterday when I reseated them. As well as a matching pair from crucial/corsair that I added just before last Christmas. How would I know which stick is faulty? I have read that on other websites about ECC problems, but their is no indication - from what I could see to which stick was faulty. The error lights on the main board for the RAM are all clear, and there is no indication within insight to which is causing the problem.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now