Jump to content

Recommended Posts

Posted

We have a two node S2D Hyper-V cluster (Server 2016) and the cluster seems to fail if host 1 goes down. Yesterday we moved all the VMs from host 1 to Host 2, tried to shut down host 1 and then host 2 blue screened.

 

Even when i disconnect the network cable from host one, the cluster seems to fail but if i disconnect/shutdown host 2, host 1 happy keeps working. I can migrate VMs okay from one host to another and running all the cluster tests seems to return back fine and the Quorum is also fine. The only thing that we do have is both DCs are in the cluster as VMs which i know is not best practice but surely this wouldn't stop the whole cluster going down if one host is online should it?

 

Thanks

Posted (edited)

Hi, my first thought of reading this is that you've not got the quorum but you've said you do have it set correctly and you've said that it passes all the cluster tests - if it is blue screening it may point to a hardware/driver issue somewhere. Some likely candidates may be 1) BASP/QASP/Vendor Teaming, 2) Some Network settings (VM queues etc?) 3) iSCSI connections to storage. 4) Shared Storage Drivers/SAS drivers.

 

Could you drop the server from the cluster and do further diagnostics from there?

Edited by tubs
Posted

As stated in the post above, it sounds to me like quorum is your main issue here. You can deploy a cloud witness to maintain quorum if one of your servers goes down, an overview can be found at Deploy a cloud witness for a Failover Cluster | Microsoft Docs. Honestly I don't understand why Microsoft even offer S2D on two servers, to my mind 3 would be a bare minimum.

 

On a tangent, I honestly think that hyperconverged computing for anything smaller than Enterprise is on the way out. It's easier and often cheaper when you look at total cost just to move things to platforms like Azure or AWS. For me it has increased reliability, scalability and ease of use by migrating all workloads over to Azure. A necessary prerequisite is going to be a decent leased line but after that your set to go. Most servers cost around £300 a year in electricity costs alone and by switching them off, you've already put £600 towards your VM's.

Posted
As stated in the post above, it sounds to me like quorum is your main issue here. You can deploy a cloud witness to maintain quorum if one of your servers goes down, an overview can be found at Deploy a cloud witness for a Failover Cluster | Microsoft Docs. Honestly I don't understand why Microsoft even offer S2D on two servers, to my mind 3 would be a bare minimum.

 

On a tangent, I honestly think that hyperconverged computing for anything smaller than Enterprise is on the way out. It's easier and often cheaper when you look at total cost just to move things to platforms like Azure or AWS. For me it has increased reliability, scalability and ease of use by migrating all workloads over to Azure. A necessary prerequisite is going to be a decent leased line but after that your set to go. Most servers cost around £300 a year in electricity costs alone and by switching them off, you've already put £600 towards your VM's.

I think i have now managed to sort the issues out but i definately take your point about Azure. Hosting VMs in azure/AWS is something ive never done and wouldn't know where to start to be honest. We have 7 VMs, 2 of which are DCs with one print server and one SCCM server but ideally i would love to start hosting more things in the cloud.

Posted
It'd be interesting to know what the issue is that you managed to clear with the cluster? Clusters somtimes seem to be a law unto themselves ...
Posted
It'd be interesting to know what the issue is that you managed to clear with the cluster? Clusters somtimes seem to be a law unto themselves ...

 

The issue we had in the end was someone had set a static entry in the hosts file!

Posted (edited)

Before I moved to a stretch cluster, I looked at S2D but discounted it as the hardware requirements were very strict. Ironically I would have had to run our 2x SAN in HBA anyway as S2D needs direct access. So far though, ive found the stretch cluster to work so I imagine S2D (when setup) should work just as well. For "piece of mind" I would have preferred to have S2D for "disaster resiliency" but have each node allowing local drive RAID.

 

CAU is a clusterf*** though. I have never had that work. We are also still running our hosts on 2016DC I havent upgraded them to 2019 yet.

Edited by KK20

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...