Jump to content

Recommended Posts

Posted

STATUS UPDATE:

 

Okay, After plugging both PSUs back in the server entered boot windows and immediately reboot cycle. After a while I got the very (un)helpful screenshot attached!

 

I've have now removed BOTH PSU's. Swapped one PSU with the hot-spare in one server and the other with the hot-spare on another server.

 

IF it's a PSU prob then I should now get either or both the other servers randomly rebooting. Also, since I'm sure the two PSUs now in the first server a fine, it should no longer reboot every chance it gets?

 

If it does reboot I suppose I'm onto the HDDs next :(

IMG_0033.JPG

Posted

Looks like a CPU error. I assume it was correctly installed into the replace board?

 

Am I correct in thinking you've had this problem for over a week? Can't you just get HP out and let them sort it rather then spending you time fixing it? I tented to do this will Dell alot, soon as I know it's hardware related I just phone them up and basically say come fix it. That way I've spend about 1 hr fixing it and I know it'll be sorted by the next day. Ok you pay a little extra, but when you add up the hours you save it pays for it's self when stuff like this happens. Plus it keeps SMT happy :D

Posted

I'm still not convinced by a CPU error.

 

  1. The system board was replaced by an HP engineer last week
  2. The same said HP engineer fitted the CPUs in the new system - all 8 threads appear in windows
  3. A 30minute run of the Prime95 stress test on all 8 threads, 100% CPU usage, showed no signs of error at all

 

I'm not counting it out completely but given how much as been tested/changed I'm still routing for a HD problem (or maybe NIC related).

 

I could pull the CPU's one at a time and see what gives?

 

EDIT: I'm actually (in a perverse way) enjoying trying to hunt down the problem. I'm lucky enough to have a relatively free jobs list at the moment. This has top number 1 high priority. I'm here till it's fixed ;)

Posted
In my experience you don't get intermittent CPU problems. It either works or it doesn't, unless you count overheating, but this can create a whole range of problems.
Posted

There is an outside chance that one of the cores has a problem with a given set of instructions that neither of the CPU stress testing programs used. And it just so happens that the light from my torch bounced of Mars at the correct angle to make those instructions run through that core at that time and cause the problems.

 

I just don't think it's very likely - personally. I'd hate to be wrong as getting old of replacement processors now (three years old Dual Core, Hyper Threaded Xeons) would probably be difficult and expensive :(

Posted
I'm actually (in a perverse way) enjoying trying to hunt down the problem. I'm lucky enough to have a relatively free jobs list at the moment. This has top number 1 high priority. I'm here till it's fixed ;)

 

Just need the van with the spare parts in ;)

 

 

Umm... I guess you could remove one CPU, check the manual as you'll need to ensure you keep one in the 1st slot. Make sure it's coated in thermal paste. Personally I'll be phoning HP and moaning my arse off cause there engineer didn't fix it.

 

Might be worth disabling\remove the NIC if you think it could be that, also fireware the BIOS and RAID controller (careful you don't lose any data!!)

Posted (edited)

Updating all the various Firmwares was one of the things HP got me to do before they agreed to send an engineer out. It didn't fix the problem on the old board and although yes It's likely hp installed a new board with older firmware that which I used to update the last board, I don't think mobo was ever the problem so a firmware upgrade won't solve this.

 

Still if the problem keeps reoccurring desperation will inevitably lead me to try this again!

 

Make sure it's coated in thermal paste. Personally I'll be phoning HP and moaning my arse off cause there engineer didn't fix it.

To be fair to HP - they sent a bloke out (after a week of over the phone/internet diagnostics) to replace a working Mobo with another working mobo!

 

STATUS UPDATE: After switching round PSU's everything appears to be fine on all three servers. But then, of course, non of the three servers have come under any king of load at all over the past hour.

 

So, the question is - whats going to happen tomorrow?

 

Either - it's a PSU prob thats now fixed (potentially waiting to kill another server), or One (or both) of the other servers are going to reboot under load, or (most likely) this server will reboot itself at 9am tomorrow morning.

 

Place your bets now...

Edited by tmcd35
Posted

If it's a memory fault I've got plenty of stick to swap it out with!

 

It's so, so, so very unlikely to be a memory fault that I'm more than happy to offer up very very good odds and be pleased to take your money when proofed to be something else ;)

Posted

After all the posts etc of going back n forth I would either go with a hard drive or the NIC ( one of ) being under a load when everyone tries to login.

 

What make / model of NICS are they ? Also am guessing they are on gigabit ?

Posted

I've been avoiding the NIC's for a reason ;)

 

These server were built and installed by the last guy and I'm not overly happy with some of his set up choices - personal opinion and all.

 

There are two on board NICs and three PCI NICs. I believe (but have yet to check) they are all HP. I know 100% they are all Gigabit. I know each of the three servers are set up the same way. I know each server has 5 IP's and theres some NIC teaming going on.

 

I think I'd sooner rule out the HDDs before getting my hands dirty and working out which NIC is which IP and what services rely on which IPs and how exactly the teaming is configured.

Posted

I think it'll re-occur, it could be one of the following.

 

- PSU issue

- Mobo could have been fitted incorrectly

- RAM could be incompatible

- RAID controller issue

- BIOS issue

- NIC issue

 

I would try disabling all the extra stuff in the BIOS, such as HT etc. Check the UPS for current load, try with no extras and the NIC disconnected.

  • Thanks 1
Posted

STATUS UPDATE:

 

  • Swapped PSU's with other servers - problem server still rebooting
  • installed brand new RAM - problem server still rebooting
  • Prime95 8 thread stress test - no restarts during test
  • Prime95 tests RAM and all CPU cores - server still reboots, but not during tests
  • Brand new Mobo installed - server still reboots
  • Onboard/Integrated SCSI RAID controllet, replaced with mobo - server still reboots
  • Limited info in error logs show same non-descript error code despite above tests/changes

 

As you can see I really am left with just NIC's and HDD's to test. While I'm not going to totally discount any other possibility -

  • new mobo with same fault as last
  • incorrectly fitted cpus
  • problem cpu core
  • cpu overheating
  • incompatible ram
  • firmware/driver issues

 

The test done so far, and the state of the machine when this first started happening, suggests that these are all extremely remote and unlikely. Oh hum, another day of digging...

Posted (edited)

Here is a novel idea, hp RAID sets are portable between smartarray adapters. You could simple shut down the server and a good one then swap all of the hard drives between them. This would rule out the drives and the OS from the list of causes. If possible do it with two machines that are not DCs as I am not 100% on how the machine SID change would affect them.

 

You need to move all of the disks at once while the servers are off then when you boot them they will just read the raid config off the transposed drives. Be sure to put them in in the right order though.

 

Edit: Oh and also upgrade all the firmware if you have not already done so. :)

Edited by SYNACK
  • Thanks 1
Posted

If the server stays up long enough for me to work out the NIC config I think I'm going to start by pulling the three additional NICs and re-introducing them one at a time.

 

Sleeping on it overnight I think I agree with the consensus here. The next most likely place is one of the NICs. Thinking about it in all honesty a randomly dodgy drive is the least likely of causes.

 

I like the idea @Synack, but it does mean downing another server - even temporarily - to do it. Also all three servers are DC's. And I'd have to do the firmware updates on all servers first. Don't want firmware missmatch causing probs if I go down this route.

 

It'd be a quicker test than pulling the drives one at a time - but potentially riskier as doing any firmware updates on a PC rebooting as often as this one now is is not exactly a wise move.

Posted

Ooo, very good question. TBH I don't rightly know :o

 

There is a 256mb SoDIMM on the motherboard. When I first saw it I thought it may have something to do with onboard graphics (although why a server may need 256mb dedicated graphics ram is beyond me). Thinking about it, it's more likely this is the RAID RAM.

 

I would have thought its the same RAM from the previous mobo. I'm pretty sure HP only replaced the actual mobo itself.

 

I'm currently sitting here waiting for the server to reboot (or not). Next period starts in about half hour. I've taken out all three additional NIC's - which on investigation appear to be totally redundant.

 

Since taking the NICs out I've not had a reboot - but then server hardly been under any load. So I'm sitting here playing the waiting game...

Posted
I would thank you Bossman but thankfully our LEA's virus checker stopped your evil plan to infect my already poorly server with a bad case of swine flu ;) ...

virus.jpg

Posted
There are two on board NICs and three PCI NICs.

 

I have to ask why so many? I highly doubt it's the onboard, as of course they've been changed since replacing the motherboard. It could well be one of these PCI NICs causing the problems.

Posted

Don't worry, I already guessed before I posted that the LEA's bluecoat server was to blame. We blame it for just about everything else ;)

 

Looks like a useful tool, I'll downloaded it at home tonight and give it a try tomorrow.

 

STATUS UPDATE: The three additional NIC's have been out all morning. So far 2 periods down and the server HAS NOT REBOOTED ONCE!

 

Of course we're probably just going through another quiet patch and whatever event is triggering the reboots hasn't happened in the past two and half hours. Bit suspicious that it was rebooting every 5 minutes before then though...

 

I'm going to give it till half way through period 4 then put the three cards back in ready for period 5. In theory - if its one of the NICs - I shouldn't get another reboot until after I put them back in, and most likely at the start of P5.

 

If this works I'll run the same test tomorrow to confirm, then add the cards one at a time to see if it's a particular card or a NIC config problem.

Posted
I have to ask why so many? I highly doubt it's the onboard, as of course they've been changed since replacing the motherboard. It could well be one of these PCI NICs causing the problems.

 

God knows. All three servers are the same. In all three cases the onboard NICs are bonded 2Gbps and are the only NICs actually being used.

 

I have 9 unrequired NICs taking up valuable IP space!?!

 

I think the last guy was starting to look into server Virtualisation. The three machines are of a good enough spec to use a s a jumping off point into this tech, and is something I'm looking into introducing later this year.

 

Even in a VM environment I couldn't see a need for more tha 4 NICs per server - 2Gbps bond data, 1Gbps SAN, 1Gbps (or even 100mbps) for IT management traffic.

 

The last guy put a lot of effort into redundancy, and I think he was spending the money on the kit while he had the chance.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...