mdrabble Posted November 4, 2015 Posted November 4, 2015 I have a rather puzzling issue with my VMs and I have no idea if it is the Hosts, the Storage or something completely different. I have all my VMs stored on Storage Server and connected via SMB3 to 3 HyperV Hosts in a Cluster. Recently almost all of my VMs will suddenly reboot for no apparent reason. Rather worryingly it seems to be happening at least twice a day (once during the day and once in the evening). Backups are using MS DPM and only take place in the evenings. I have checked and no backups are running during the day. When I look at the evening logs, nothing jumps out at me, On the VM, I see the following, but I would expect to see these as it has lost connection. Event ID 6008 – The shutdown was unexpected at (time/date) Event ID 41 – Kernel Power (The system has rebooted without cleanly shutting down first. This error could be caused if the system stopped responding, crashed, or lost power unexpectedly.) On the HyperV Hosts mention Event iD 1069 – FailoverClustering failed Cluster resource 'VM Details’ failed. Based on the failure policies for the resource and role, the cluster service may try to bring the resource online on this node or move the group to another node of the cluster and then restart it. Check the resource and group state using Failover Cluster Manager or the Get-ClusterResource Windows PowerShell cmdlet.) If this all of the VMs in the cluster then I could say possibly the Storage Server but some VMs remain up and running!! On the last reboot at 10:15ish this morning - from my 16 VMs, 6 remained up and running. When the VMs reboot it isn't always the same VMs that stay up and running either - it is pot luck which VMs stay running and on which host. I do have Virtual Machine Manager - nothing looks out of the ordinary on that - although I could be wrong and missed something. Anyone any ideas or suggestions? I'm waiting on a call from VeryPC to see if they can help but thought would also ask on here incase anyone has had similar issues. Cheers
Achandler Posted November 4, 2015 Posted November 4, 2015 I assume that the physical computers stay on, but the Virtual machines on them are rebooted. It definitely isn't consistantly on two hosts, because the actual Virtual machines ofcourse could be on different hosts each reboot (especially if they try to failover) which would make the reboots appear totally incosistant while still being consistant. If the physical hosts say that clustering has failed, it implies that the hosts themselves are falling from the cluster, then re-entering it (because I assume you don't notice them dropping out, only that the VMs are rebooting). If the hosts are physically rebooting then I would firstly run the failvoer cluster validation just to double check the software configuration. Is there any messages in Failover Manager saying the host was removed from the cluster for example? Then I would look at the possibility of it being a network drop (since power is fine), there is a limit on a dropped connection to one of the HyperV network because the host is "offline" when its actually online, it might jsut be by the time you arrive at the problem the network is back again. Do all the hosts have mutliple channels to the Storage Server?
mdrabble Posted November 4, 2015 Author Posted November 4, 2015 I assume that the physical computers stay on, but the Virtual machines on them are rebooted. Yep the hosts/storage stay on. It definitely isn't consistantly on two hosts, because the actual Virtual machines ofcourse could be on different hosts each reboot (especially if they try to failover) which would make the reboots appear totally incosistant while still being consistant. VMs don't attempt to migrate, they just reboot as if you turn them off and on again. If the physical hosts say that clustering has failed, it implies that the hosts themselves are falling from the cluster, then re-entering it (because I assume you don't notice them dropping out, only that the VMs are rebooting). If the hosts are physically rebooting then I would firstly run the failvoer cluster validation just to double check the software configuration. Is there any messages in Failover Manager saying the host was removed from the cluster for example? Nope, no messages about hosts being removed from the cluster. Then I would look at the possibility of it being a network drop (since power is fine), there is a limit on a dropped connection to one of the HyperV network because the host is "offline" when its actually online, it might jsut be by the time you arrive at the problem the network is back again. Do all the hosts have mutliple channels to the Storage Server? All hosts are using Teamed NICs so have multiple connections. Having gone through the Event logs again and looking at the Cluster Event, here is an example for 1 VM Event ID:1069 @ 10:27:36 Cluster resource 'Virtual Machine DC-01' of type 'Virtual Machine' in clustered role 'DC-01' failed. Based on the failure policies for the resource and role, the cluster service may try to bring the resource online on this node or move the group to another node of the cluster and then restart it. Check the resource and group state using Failover Cluster Manager or the Get-ClusterResource Windows PowerShell cmdlet. Event ID: 1069 @ 10:28:43 Cluster resource 'Virtual Machine DC-01' of type 'Virtual Machine' in clustered role 'DC-01' failed. The error code was '0x20' ('The process cannot access the file because it is being used by another process.'). Based on the failure policies for the resource and role, the cluster service may try to bring the resource online on this node or move the group to another node of the cluster and then restart it. Check the resource and group state using Failover Cluster Manager or the Get-ClusterResource Windows PowerShell cmdlet. Event ID: 1025 @ 10:28:43 The Cluster service failed to bring clustered role 'DC-01' completely online or offline. One or more resources may be in a failed state. This may impact the availability of the clustered role. After which the VM is back up and running. I did think about a possible connectivity issue but as all VMs are stored in the same location I would have assumed it would have affected all VMs and not just some.
localzuk Posted November 4, 2015 Posted November 4, 2015 I have seen exactly this same thing. It has almost completely stopped now though, since I did the following: 1. Disabled VMQs on everything I could find. 2. Installed the following: https://support.microsoft.com/en-gb/kb/2919355 https://support.microsoft.com/en-gb/kb/3000850 https://support.microsoft.com/en-gb/kb/3013769 https://support.microsoft.com/en-gb/kb/3072380 https://support.microsoft.com/en-gb/kb/3068445 https://support.microsoft.com/en-gb/kb/3068444 https://support.microsoft.com/en-gb/kb/3031598
mdrabble Posted November 4, 2015 Author Posted November 4, 2015 I have seen exactly this same thing. It has almost completely stopped now though, since I did the following: 1. Disabled VMQs on everything I could find. 2. Installed the following: https://support.microsoft.com/en-gb/kb/2919355 https://support.microsoft.com/en-gb/kb/3000850 https://support.microsoft.com/en-gb/kb/3013769 https://support.microsoft.com/en-gb/kb/3072380 https://support.microsoft.com/en-gb/kb/3068445 https://support.microsoft.com/en-gb/kb/3068444 https://support.microsoft.com/en-gb/kb/3031598 Numpty question - install on hosts and VMs or just hosts?
mdrabble Posted November 4, 2015 Author Posted November 4, 2015 Will download and take a look on Friday and taking tomorrow off :-)
mdrabble Posted December 18, 2015 Author Posted December 18, 2015 Right...... installed updated from this MS tech doc - https://support.microsoft.com/en-us/kb/2920151 Thought it was sorted as had a period of time where all VM stayed up and no random reboots when suddenly the other day, it started again - the latest random reboot as 01:50 this morning. Think I may try and look through more event logs and then call Microsoft to see if they can see what is going on.
localzuk Posted December 18, 2015 Posted December 18, 2015 I have now not had any reboots for 12 days. This is after doing a thorough check and installing missing updates on one of my nodes. I also shifted a couple of the heavier VMs over to a second storage server. On Monday, I will be doing a thorough update and tidy of all our servers, hardware and VM. Any that are still on 2012 will move to R2 etc... So, we'll see if it stays stable!
mdrabble Posted December 18, 2015 Author Posted December 18, 2015 All but 2 of my servers are running 2012R2. Are you using Storage Spaces or iscsi?
MadIT Posted April 11, 2017 Posted April 11, 2017 I know this thread is old, but i have the same issue. My environment consist of 2 windows server 2012 r2 and have HyperV fail-over cluster. i have about 30 VMs and have iscsi connection to a Dell SAN. we use veeam to backup. it happen to different vm, on different host. and it is not every night. but only happen at night. i have done everything, patches, updates, turn off VMQ. Each host have about 384GB of RAM. any solutions?
mdrabble Posted April 11, 2017 Author Posted April 11, 2017 My issue turned out to be the storage card. Replace with a new card and no problems since. 1
Ricardo Posted December 12, 2017 Posted December 12, 2017 Please, I'm exactly with the same problem, how did they solve it?
Ricardo Posted December 12, 2017 Posted December 12, 2017 Hello mdrabble, could you detail this solution if possible? Thanks Hello mdrabble, could you detail this solution if possible?
Ricardo Posted December 12, 2017 Posted December 12, 2017 Hello mdrabble, could you detail this solution if possible? Thanks Hello mdrabble, could you detail this solution if possible?
mdrabble Posted December 12, 2017 Author Posted December 12, 2017 Hi @Ricardo The issue turned out to be the HBA Controller not being able to handle to IO traffic which caused the card to reset/resend data. We replaced with a RAID card so it could handle the IO traffic and not had any issues since. Downside is that no longer using Storage Spaces which means cannot use the HDD/SSD Tiers - I now have a seperate HHD Array and SSD Array. Hope that helps
stevensedory Posted May 4, 2018 Posted May 4, 2018 Hi @Ricardo The issue turned out to be the HBA Controller not being able to handle to IO traffic which caused the card to reset/resend data. How did you know it was your HBA? And what HBA were you using? I'm having similar issues, but I don't seem to see any indication on our SAN that the HBA is having issues.
mdrabble Posted May 4, 2018 Author Posted May 4, 2018 How did you know it was your HBA? And what HBA were you using? I'm having similar issues, but I don't seem to see any indication on our SAN that the HBA is having issues. Has support calls open with Microsoft and sent lots of log files as we originally thought it was a HyperV/Cluster issue - once Microsoft ruled that out we looked at the hardware. HBA was an intel card (cannot remember what model) - again Intel had various logfiles and it turned out the HBA card couldnt handle the IO traffic. Once we switch to a Raid Card we havent had a since issue since.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now