penfold Posted May 10, 2013 Posted May 10, 2013 I've got a problem for the last few days that we have a Server2008 R2 which keeps experiencing unexpected shutdowns. The error in the event log is Event ID41 Kernel-Power "the system has rebooted without cleanly shutting down first. This error could be caused if the system stopped responding, crashed, or lost power unexpectedly." I keep reading that this could be caused by a physcial or driver problem but as it's a VM (and the only one with problems) has anyone got any suggestions of where I can start to troubleshoot? I've ran a Virus Scan and windows Updates but non of which has made any difference.
Duke Posted May 10, 2013 Posted May 10, 2013 What hypervisor are you using, and have you got all the guest additions (e.g. VMware Tools) fully installed and up to date. Was this a fresh virtual build or was it a P2V?
penfold Posted May 10, 2013 Author Posted May 10, 2013 latest VM tools are installed and fresh build. Has been working fine for the last couple of months. We did have a trojan detected on one of the shares but althouh the system scans clean I cant help think it is related. Scans show it's clear so I am trying to troubleshoot the restarts.
m25man Posted May 10, 2013 Posted May 10, 2013 Is it an HP by any chance? If so check and replace your power supplies. I had exactly the same problem (Hyper-V though) the Hosts DC PSU was randomly tripping and as it turned out the tiny fan on the PSU had become clogged and sluggish and the PSU had been slowly overheating. Nothing in the logs other than a similar entry to those you have described. Popped in a new PSU and its been up ever since. 1
penfold Posted May 13, 2013 Author Posted May 13, 2013 Is it an HP by any chance? If so check and replace your power supplies. I had exactly the same problem (Hyper-V though) the Hosts DC PSU was randomly tripping and as it turned out the tiny fan on the PSU had become clogged and sluggish and the PSU had been slowly overheating. Nothing in the logs other than a similar entry to those you have described. Popped in a new PSU and its been up ever since. Wouldn't this affect more than 1 VM though? I only see this problem on 1 machine. I actually thought the problem had gone away as it seemed to be ok over the weekend but I've just noticed it happen again this morning:(
Duke Posted May 13, 2013 Posted May 13, 2013 Nothing else in the logs prior to it crashing? Any scheduled tasks that run around that time? Could you clone the VM, power it up without networking, then see if they both crash or whether just the 'live' on crashes? This might tell you if it's something wrong with the VM itself, or whether it's triggered by interaction with something else on the network.
penfold Posted May 13, 2013 Author Posted May 13, 2013 Nothing in the logs to indicate anything wrong (as far as I can see) I have had the replica up and running in a sandbox without issue. I've looked for user connections just before it shuts down but nothing stands out. My next step is to failover to the replica when I get a chance to do a clean shutdown and see how that goes.
Duke Posted May 13, 2013 Posted May 13, 2013 Hmm, if the replica runs fine then it does suggest it's being caused by interaction with something on the network, weird one though. What does the server do? Seen bad printer drivers and print spoolers and things crash servers in the past.
penfold Posted May 13, 2013 Author Posted May 13, 2013 It's just a file Server. I did think it was possibly being caused by a virus or something of that ilk stored on a users area, but a scan comes back clean. The replica runs fine, although as it isn't interacting with the network, I'm tempted to failover to that and see if it still reboots.
m25man Posted May 13, 2013 Posted May 13, 2013 Wouldn't this affect more than 1 VM though? I only see this problem on 1 machine. I actually thought the problem had gone away as it seemed to be ok over the weekend but I've just noticed it happen again this morning:( I had incorrectly interpreted that this was the physical box not just one VM... in which case my focus would be on the Hosts RAM or the storage subsystem the VM is stored on.
penfold Posted July 23, 2013 Author Posted July 23, 2013 (edited) OK, I have re-installed the SP1 on this box and we're still having problems and now I'm out of ideas. I've also realised that I had said that it runs ok in the sandbox, but when we ran it from the replica, the shutdown problem still occurred. My next action consists of either doing an upgrade to 2012 or rebuilding a new VM and transferring the files over. Neither of which I really want to do so I'm open to any suggestions that anyone may have in case I've missed something obvious in fixing this. I have also noticed that some time errors have occurred but I'm not sure if these are caused by the unexpected shutdown or if they are related to the same problem. Edited July 23, 2013 by penfold
Duke Posted July 29, 2013 Posted July 29, 2013 How much work would building a new VM and transferring data across be? If it's not giving you any indication why it's crashing then it might be the quickest fix. Have you got anything else you can rule in/out? Remove any unneeded virtual hardware from the server, upgrade VMware and VMware tools, build an identical server (from scratch, not cloned) and see if it crashes?
penfold Posted July 29, 2013 Author Posted July 29, 2013 That's what I'm doing now. Bit of a pain as it seems I was only doing this recently to get it from physical to virtual, but at least it should be quicker to solve the problem. Although now I'm putting in 2012 instead of 2008.
Duke Posted July 29, 2013 Posted July 29, 2013 Cool, maybe just leave the non-production one running but not in use for a while to make sure it doesn't suffer the same problem. If it does crash too then it definitely points to something bad in the environment.
penfold Posted July 29, 2013 Author Posted July 29, 2013 No other servers are suffering from the same symptoms so I assume something has gone corrupt on the server. Hopefully transferring over to a new one will solve it, but it does mean the backups are going to get a little skewed for a couple of weeks while I keep a copy of both servers going.
glipford Posted October 24, 2013 Posted October 24, 2013 What was your final solution? I have 2 VM's that are randomly shutting down/restarting with kernel-power eventid 41. I am in a cluster so plan to migrate to a different server but again the VM's are getting this error not the physical server
penfold Posted October 25, 2013 Author Posted October 25, 2013 I ended up building a new server and transferring the files over. I used it as an excuss to move from 2008 to 2012 but nothing else seemed to work. Probably not the answer you were after 1
glipford Posted October 25, 2013 Posted October 25, 2013 Thanks for the reply, I am running 2012 host and the VM's are 2008R2SP1, I moved all VM's to the second host to see what happens this today. Thanks again.
penfold Posted October 25, 2013 Author Posted October 25, 2013 @glipford - I did try migrating the VM's but it didn't make a difference. The only thing I didn't try was restoring the OS files from a replica/backup as by the time I realised what was happening the backups would have had the same files as the production VM. Building a new VM and transferring the files was the only thing that worked as otherwise the problem always came back.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now