Jump to content

Recommended Posts

Posted

Hi I think I may have posted this before but I cannot find it, so apologies.

I keep having random problems where some of our Windows Server 2012 R2 instances will randomly "grind to a halt". (more below)

This has happened to one of our print servers, to our exchange server, to our SCCM server etc. It seems to happen once every 6 months or so.

We are on a Server 2012 R2 forest (not server 2012 functional level though as we have 2008 legacy servers), hosted on VSphere 5.5 on 3 physical hosts. Other servers on the same physical host remain working.

It seems to happen after we have physical layer issues - our LEA ran switches are not on a UPS so the network goes down if we have a power failure.

 

Basically what happens - the affected server will take hours (12+) to boot up. Different servers seem to get stuck on different sections, but the last one was stuck on a Group Policy section. It will get stuck for hours but will eventually progress to the logon screen. When it hits the login screen however, half of the services have not started - e.g. you can not browse to the servers fileshares; nor connect remotely to Services as the RPC server has not started. You can log in as an Administrator but it will take hours (again, 12+) to reach the desktop. Once in, the machine will be very unresponsive e.g. the UI will not load various sections such as System.

If you disable the NIC, it will boot up and log in very quickly as expected; but the services still will not start / stop properly. trying to stop a service will just hang services.msc at 100% core usage.

Doing a group policy model shows nothing unusual as there will be another server with the same policy working fine.

You can connect to the event viewer but there is nothing in any of the logs that shows any errors; it is as if the server is working fine just extremely slowly. e.g. the latest error for hours will be that the login provider is taking a while (134524 seconds). and in system all you will see is random services starting/stopping.

 

If I restore the server from a Veeam backup, it wil work fine for a number of days and then it will display the exact same problem again.

The only way to fix it, is to boot the server up in Safe Mode with Networking; disjoin it from the domain; reboot it again in safe mode with networking and rejoin it to the domain with the same name. Then the server will work fine as if nothing had happened.

This isn't just a Server 2012 R2 issue as our Exchange server suffered exactly the same thing, running Server 2008.

 

Anyone else seen anything like this? It is very hard to search for unfortunately as there are no error codes and Windows Server can run slowly for many, many different reasons so finding out what causes this is a needle in a haystack :-(

I mean generally it is not considered a good idea to rely on Microsoft software for anything important so at the moment I am planning around it but if this happens to our primary domain controller I am not quite sure what I'll do.

 

Thanks

Posted

Hello Mate,

 

I have a very similar infrastructure to you (vSphere 5.5 etc) and I have a 2012R2 server which randomly locks up withthe only option to perform a reset on the guest from vSphere. Nothing at all like your experience i'm afraid.

 

Could I pose a couple of things to try though which may help to pin down the underlying problem:

 

(1) Are you using DRS on vSphere?

(2) Have you tried migrating the guest to another host?

(3) Have you tried migrating the datastore?

(4) How do the vm guest performance figure look during normal use/slow startup etc (look for latency issues)

(5) Can you stand another DC up to help with the failure on the PDC emulater?

 

Sorry if you have tried all of this all ready.

Posted

I assume you have two domain controllers - Best practice and all that. If they are both 2012 then raise the functional level as this only relates to the Domain controllers not any member servers.

 

Also if this happens again I suggest you really get a UPS for the rack/switches to stop this issue and also have you tried just unplugging the network cable or at least plugging it into a spare unmanaged switch to get the system back up.

 

Just my five cents worth.

Posted (edited)

Hi no we are not using DRS on our cluster; we have tried migrating the host to a different physical but we haven't tried migrating it to a different datastore yet. It migrates across hosts fine (i.e. doesn't take ages), but the server itself still boots slowly/doesn't work regardless of what host it is on even if it is alone on a host.

I wasn't sure if it was a VMWare issue or if it was a Windows issue a "layer" up. Checking the performance figures for the server in VMware, the CPU usage, Mem usage, Network, Disk Latency etc all flatlines almost as if the server is off; it is barely above 0. once I have rejoined it to the domain it returns to normal. We have two DC's and they seem to serve other stations fine during this.

I have just fixed our SCCM server by rejoining it to the domain but part of me thinks I should have left it as it was for further testing as it doesn't matter so much if it is down (as opposed to the mail or print servers)

 

I have just found the VMWare Communities so I am having a trawl through there to see if anyone else has reported this.

but yeah i am going to see if they can spring for a UPS. I think the school will argue the LEA should pay for it and the LEA will argue the school should pay for it...

it doesn't help that they put 15 or so PoE capable switches on one 15a ring, so every time the power goes down I have to come in to unplug them and manually reset the fusebox switch...

Edited by mikes

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...