Jump to content

Recommended Posts

Posted

I have a rather puzzling issue with my VMs and I have no idea if it is the Hosts, the Storage or something completely different.

 

I have all my VMs stored on Storage Server and connected via SMB3 to 3 HyperV Hosts in a Cluster.

 

 

Recently almost all of my VMs will suddenly reboot for no apparent reason. Rather worryingly it seems to be happening at least twice a day (once during the day and once in the evening). Backups are using MS DPM and only take place in the evenings. I have checked and no backups are running during the day.

 

When I look at the evening logs, nothing jumps out at me,

 

 

On the VM, I see the following, but I would expect to see these as it has lost connection.

 

Event ID 6008 – The shutdown was unexpected at (time/date)

Event ID 41 – Kernel Power (The system has rebooted without cleanly shutting down first. This error could be caused if the system stopped responding, crashed, or lost power unexpectedly.)

 

 

On the HyperV Hosts mention

 

Event iD 1069 – FailoverClustering failed

 

Cluster resource 'VM Details’ failed.

 

Based on the failure policies for the resource and role, the cluster service may try to bring the resource online on this node or move the group to another node of the cluster and then restart it. Check the resource and group state using Failover Cluster Manager or the Get-ClusterResource Windows PowerShell cmdlet.)

 

 

If this all of the VMs in the cluster then I could say possibly the Storage Server but some VMs remain up and running!! On the last reboot at 10:15ish this morning - from my 16 VMs, 6 remained up and running.

 

When the VMs reboot it isn't always the same VMs that stay up and running either - it is pot luck which VMs stay running and on which host.

 

I do have Virtual Machine Manager - nothing looks out of the ordinary on that - although I could be wrong and missed something.

 

Anyone any ideas or suggestions? I'm waiting on a call from VeryPC to see if they can help but thought would also ask on here incase anyone has had similar issues.

 

Cheers

Posted

I assume that the physical computers stay on, but the Virtual machines on them are rebooted.

 

It definitely isn't consistantly on two hosts, because the actual Virtual machines ofcourse could be on different hosts each reboot (especially if they try to failover) which would make the reboots appear totally incosistant while still being consistant.

 

If the physical hosts say that clustering has failed, it implies that the hosts themselves are falling from the cluster, then re-entering it (because I assume you don't notice them dropping out, only that the VMs are rebooting). If the hosts are physically rebooting then I would firstly run the failvoer cluster validation just to double check the software configuration. Is there any messages in Failover Manager saying the host was removed from the cluster for example?

 

Then I would look at the possibility of it being a network drop (since power is fine), there is a limit on a dropped connection to one of the HyperV network because the host is "offline" when its actually online, it might jsut be by the time you arrive at the problem the network is back again. Do all the hosts have mutliple channels to the Storage Server?

Posted
I assume that the physical computers stay on, but the Virtual machines on them are rebooted.

 

Yep the hosts/storage stay on.

 

 

It definitely isn't consistantly on two hosts, because the actual Virtual machines ofcourse could be on different hosts each reboot (especially if they try to failover) which would make the reboots appear totally incosistant while still being consistant.

 

VMs don't attempt to migrate, they just reboot as if you turn them off and on again.

 

If the physical hosts say that clustering has failed, it implies that the hosts themselves are falling from the cluster, then re-entering it (because I assume you don't notice them dropping out, only that the VMs are rebooting). If the hosts are physically rebooting then I would firstly run the failvoer cluster validation just to double check the software configuration. Is there any messages in Failover Manager saying the host was removed from the cluster for example?

 

Nope, no messages about hosts being removed from the cluster.

 

Then I would look at the possibility of it being a network drop (since power is fine), there is a limit on a dropped connection to one of the HyperV network because the host is "offline" when its actually online, it might jsut be by the time you arrive at the problem the network is back again. Do all the hosts have mutliple channels to the Storage Server?

 

All hosts are using Teamed NICs so have multiple connections.

 

 

Having gone through the Event logs again and looking at the Cluster Event, here is an example for 1 VM

 

Event ID:1069 @ 10:27:36

Cluster resource 'Virtual Machine DC-01' of type 'Virtual Machine' in clustered role 'DC-01' failed.

 

Based on the failure policies for the resource and role, the cluster service may try to bring the resource online on this node or move the group to another node of the cluster and then restart it. Check the resource and group state using Failover Cluster Manager or the Get-ClusterResource Windows PowerShell cmdlet.

 

 

Event ID: 1069 @ 10:28:43

Cluster resource 'Virtual Machine DC-01' of type 'Virtual Machine' in clustered role 'DC-01' failed. The error code was '0x20' ('The process cannot access the file because it is being used by another process.').

 

Based on the failure policies for the resource and role, the cluster service may try to bring the resource online on this node or move the group to another node of the cluster and then restart it. Check the resource and group state using Failover Cluster Manager or the Get-ClusterResource Windows PowerShell cmdlet.

 

 

Event ID: 1025 @ 10:28:43

The Cluster service failed to bring clustered role 'DC-01' completely online or offline. One or more resources may be in a failed state. This may impact the availability of the clustered role.

 

 

After which the VM is back up and running.

 

I did think about a possible connectivity issue but as all VMs are stored in the same location I would have assumed it would have affected all VMs and not just some.

Posted
Posted

 

 

Numpty question - install on hosts and VMs or just hosts?

  • 1 month later...
Posted

Right...... installed updated from this MS tech doc - https://support.microsoft.com/en-us/kb/2920151

 

Thought it was sorted as had a period of time where all VM stayed up and no random reboots when suddenly the other day, it started again - the latest random reboot as 01:50 this morning.

 

Think I may try and look through more event logs and then call Microsoft to see if they can see what is going on.

Posted

I have now not had any reboots for 12 days. This is after doing a thorough check and installing missing updates on one of my nodes. I also shifted a couple of the heavier VMs over to a second storage server.

 

On Monday, I will be doing a thorough update and tidy of all our servers, hardware and VM. Any that are still on 2012 will move to R2 etc... So, we'll see if it stays stable!

  • 1 year later...
Posted
I know this thread is old, but i have the same issue. My environment consist of 2 windows server 2012 r2 and have HyperV fail-over cluster. i have about 30 VMs and have iscsi connection to a Dell SAN. we use veeam to backup. it happen to different vm, on different host. and it is not every night. but only happen at night. i have done everything, patches, updates, turn off VMQ. Each host have about 384GB of RAM. any solutions?
  • 8 months later...
Posted

Hi @Ricardo

 

The issue turned out to be the HBA Controller not being able to handle to IO traffic which caused the card to reset/resend data.

 

We replaced with a RAID card so it could handle the IO traffic and not had any issues since.

 

Downside is that no longer using Storage Spaces which means cannot use the HDD/SSD Tiers - I now have a seperate HHD Array and SSD Array.

 

Hope that helps

  • 4 months later...
Posted
Hi @Ricardo

 

The issue turned out to be the HBA Controller not being able to handle to IO traffic which caused the card to reset/resend data.

 

How did you know it was your HBA? And what HBA were you using? I'm having similar issues, but I don't seem to see any indication on our SAN that the HBA is having issues.

Posted
How did you know it was your HBA? And what HBA were you using? I'm having similar issues, but I don't seem to see any indication on our SAN that the HBA is having issues.

 

Has support calls open with Microsoft and sent lots of log files as we originally thought it was a HyperV/Cluster issue - once Microsoft ruled that out we looked at the hardware.

 

HBA was an intel card (cannot remember what model) - again Intel had various logfiles and it turned out the HBA card couldnt handle the IO traffic.

 

 

Once we switch to a Raid Card we havent had a since issue since.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...