Jump to content

Recommended Posts

Posted
ECC log should tell you if you can access it. Yes clearing the log could make the results different as it could be failing based on historical errors. If you don't look at it beforehand though you loose all the debugging goodness of the log.
Posted

Cheers Synack,

 

Apparently to clear it:

 

"To solve the Problem with the full SEL Log we did the following actions:

 

1.) Remove the CMOS Battery on the Systemboard off the Server

 

2.) Set the System Maintenance Dipp Switch Number 6 to "On"

 

3.) Start the Server and let it Run for 3 minutes

 

4.) Set the System Maintenance Dipp Switch Number 6 to off

 

5.) Leave the Server without Power for 30 seconds

 

6.) Put the CMOS Battery back to the Server on the System board.

 

FINISH.

 

This Action clears everything, inclusing the SEL Log .

 

Attention .... the ILO Board is cleared too with this process"

 

This was the response on the HP forums for clearing the ECC log. Hopefully this is the correct procedure?

Posted

This was the response on the HP forums for clearing the ECC log. Hopefully this is the correct procedure?

 

Never done it before myself but it sounds convincing. Had simmilar issues but got around them by replacing the RAM with some of a better pedigree than we tried initially.

Posted
If the ECC tests fail after clearing this log, do you think my next step would be to replace all the RAM?

 

Yank the new ram, do the reset then retest. It is probably the new RAM at which point you can warrenty it or replace it with something else.

Posted

So I followed them instructions to clear the ECC - their wrong, it just cleared the bios settings. Apparently the log within insight diagnostics is what it was complaining about. Tested at first with the original HP RAM - ECC 100% pass, put the new RAM in 100% pass! No ECC Errors! I ran it a couple of times and each time results were fine. Did a complete test - everything passed. Looked in the log where the previous ECC error was to find "Malformed NVRAM detected. Device: HP Smart Array Controller. Slot 0 Property name: World Wide ID" So I've gone from a RAM problem to a array controller RAM problem?!

 

I've searched for that error and theirs a couple of sites mentioning clearing the NVRAM - which is apparently the same procedure as I mentioned before by switching 6 to on and removing the mother board battery. This didn't work - I tried it a couple of times and still the same error message.

 

Also read this could be another seating issue so tried re-seating it.

 

The server also decided to reset during the day today! which was great!

 

Any suggestions?

Posted

Well I'm assuming so, this error wasn't there till I resolved the RAM ECC problem.

 

In the array diagnostics theirs no errors, in insight no errors, just that "Malformed NVRAM detected. Device: HP Smart Array Controller. Slot 0 Property name: World Wide ID" in the log, and no error lights are lighting up.

Posted
Yes, you may end up needing to replace the RAID controller if that is infact the fault. How old is it, some servers can get really out of wack after 6 or so years and end up with rather complicated and difficult to track errors. Had some really old DL380 servers that had dodgey drive backplains and that was not fun to diagnose or repair. That was nto in a school though, rather a seporate business that was using it for testing.
Posted

Interesting, we have (or did have) exactly the same problem with the exact same model of server, never really 100% got to the bottom of it though it does seem to have stopped after all the WD HDs that shipped with the server (and replaced over and over again with other WD disks after many issues) were replaced with Seagate ones.

Hasn't happened in quote a while now.

Posted

So apparently the server isn't as old as I thought, haven't been working here too long, turns out the server had only been installed in 2009 so we do still have warranty on it (Of course all the really important documents that came with the server were misplaced.)

 

So have spent over three hours today talking to HP support...repeating everything on here over and over. The best suggestion has to be upgrading the firmware to a different smart array firmware for a different model, great suggestion hp! he suggested something that I haven't tried which was a "Power Drain" in which the server is unplugged and the power key is pressed for 20seconds, this sounds like a bit of rubbish to me but worth a try. Will try it tomorrow evening.

 

Hopefully they will sort it, I've found a few people mentioning on the net ml350 with the same issues.

 

I'll look tomorrow at the make of the drives not sure off the top of my head.

Posted

Update if anyone is interested!

 

After trying the power drain on Friday, was called this morning to hear that both hardware failure lights on red again, so the site manager restarted the server. Spoke to HP who have told me they think it is either a UPS problem or the power supply and if its neither of them they will replace the mainboard which will also replace the Smart Array.

 

Kinda getting the sense they don't know what the problem is either!

Posted

Thanks for all the help everyone.

 

So HP finally gave in and let an engineer come out. Apparently the NVRAM problem is nothing, apparently the power supply has had issues - I sent them the part number and revision and was told that it has problems. So they replaced the power supply and back plane. So far so good, no restarts, they have told me if there is any further issues they will replace the motherboard.

 

Cheers!

Posted
Thanks for all the help everyone.

 

So HP finally gave in and let an engineer come out. Apparently the NVRAM problem is nothing, apparently the power supply has had issues - I sent them the part number and revision and was told that it has problems. So they replaced the power supply and back plane. So far so good, no restarts, they have told me if there is any further issues they will replace the motherboard.

 

Cheers!

 

HP support can be good! you just sort of have to play them abit! Obviously with the more complex stuff they tend to hang around on it but otherwise they tend to be pretty top notch when it comes to replacing failed/faulty parts. Only time i had trouble was with the core switch - but the replacement parts for it where around 4k!

  • 1 month later...
Posted

Hi

I have the same problem, my server every few weeks shuts down, after reading the error report, i get a blue screen erorr. The problem is we have only just had this server installed about a year ago. Im really worried that this may be aserious problem.

 

What steps should i take to try and resolve this isse?

 

Thanks

 

Aaqib

Posted

Prolly should start a new thread for that aaqib.

 

That's a different problem from me as I never had any blue screen errors.

 

Start a new thread and give details of the blue screen errors etc.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...