Jump to content

Recommended Posts

Posted

Just making sure at this point that all your backups are in a good state and that your nightly backup is still functioning?

 

Also while I have a think, do you not have any vendor support?

Posted
Just making sure at this point that all your backups are in a good state and that your nightly backup is still functioning?

 

Also while I have a think, do you not have any vendor support?

 

Backups are happy via Veeam still. We do have some VMs not backed up, mainly test machines, MDT builds etc.

 

No vender support per say as these are refurb servers, but they are under warranty (ICT Direct)

Posted

A huge shout out and thank you to @Tefters who gave up his evening last night to go on a wild 'Hunt the Microsoft issue' expedition, which at one point was looking like the backups were going to get a 'Full Function Test'. However we kept at it and by late last night everything was back in a healthy state.

 

We thought the first host, HOST-A which had the IO Errors was the issue, but found later on that this server was in fact fine, the second host wasn't communicating with it. As this one was running the cluster the IO Error was between it, and the disks on Host A. Had some windows update problems as well, which do appear to have caused the issue, but which one we're not sure. So for now the servers have updates pause, mainly to give me chance for my stress level and heart rate to drop back down to normal. Easter they'll get a good going over.

 

Amazing of @Tefters to provide such fantastic support to myself, along with suggestions from others in this thread. Thank you very much!

  • Thanks 1
  • 2 weeks later...
Posted (edited)

not sure if I'm on the right post, but similar updates related glitch had caused lots of headaches/stress.

HC cluster with 2 nodes only 2 years old, was updated in January, after engineers checked all pre-requisites. On the day updates were completed, no problems, everything ticked green.

Next day, random VMs were hanging, accessible only in Hyper v manager, while their network card was unresponsive, any attempt to disable/enable in hyper v manager RDP was just not responding. Event logs pointed to RHS service, Event ID 1146. Such VMs could not be controlled, and only way to recover them was to put the node in standby, drain roles if possible, and reboot the node. So much for failover clustering. Then when node was up, the VM which was locked by RHS, would sometimes work, but had to be fixed with chkdsk, and 3 VMs had to be rebuilt using ISO, or partially restored from backups. This instability continued several weeks, sometimes VMs would lock up overnight while backups were running. Good response received from support engineers, sometimes working out of hours or weekends. In the meantime rental hyper v servers being used to host all VMS, as original cluster is unstable, even after full rebuild of 2019 datacenter OS. Anyone with helpful input welcome, because root cause of "The cluster Resource Hosting Subsystem (RHS) stopped unexpectedly" is still not known? There are some event log errors pointing to Mellanox firmware...

Edited by evasion
Posted (edited)

What hardware are we talking about here, make/model?

 

You say 2 years old, was this new or 2nd hand?

 

The storage IO errors seen in this post were only root caused down to a Windows patch but didn't pursue exactly which one as it was out of 2.

It could have been storage or storage network issues, looked like storage on the surface but this was a 2nd hand setup with no time for further analysis unfortunately, getting the cluster up was the priority.

 

The general consensus is something funny went into the January updates which has caused a bit of a stir with HCI clusters.

 

Have not spotted anything out there on the web specifically stating about this but haven't been hunting either so...

 

Who are the engineers you have been working with, are they or do they have direct lines into MS premium support?

Edited by Tefters
Posted (edited)

I will add that I do not trust any company when it comes to patching a cluster, Microsoft, Citrix, VMware, I have been burnt by them all in the past so I turn off auto update on cluster hosts as the "aware" in cluster aware updating stand for "we are aware something might F up" and cause a HUGE headache, IT nowadays is "it should" not "it will".

 

I manually patch my hosts quarterly during half terms that way I have adequate time to troubleshoot and if 1 goes wrong I don't hit the update button on the second.

 

If I was a data centre engineer nowadays I would either have a mirrored test environment both HW and SW or have CAU off, patch 1 manually and then re-enable it for the rest and then turn it straight back off after it's done.

Edited by Tefters
Posted

@robyholmes What are your storage NIC's on your servers, make/model?

 

Having just checked my cluster i'm running QLogic FastLinQ QL41262-DE 25GbE Adapters for storage NIC's so wonder if there is a Mellanox issue here in which case Nvidia have possibly F'd up somewhere as I believe they own Mellanox now.

Posted (edited)
Mellanox ConnectX-5 adapters. Back to back on 25 Gb link. HC cluster is now only used for testing in the hope it will reveal root cause of RHS crashing. When this service which is responsible for high availability is randomly crashing, it's a game over. Edited by evasion
Posted

Trying to find the root cause though is key to this all.

 

HCI does have a wide audience, maybe not as much as classic server/san but I do see HCI taking over whether that be 2 node setup or more.

 

Is this a freak occurrence, is it a Mellanox firmware update, is this a Windows update, those are questions that need answers.

 

Sans also suffer from similar issues, if a SAN header fails the other is supposed to take over, supposed to being the crucial word here.

 

I've had a storage network completely collapse due to a SAN header failure and the other one didn't register properly to take over.

 

I've also been in a split brain scenario with HA databases, that is not fun!

 

I've seen core switches both for normal network and VOIP horribly fail.

I remember an issue 2 years back from cross line VOIP calls which is supposed to be impossible due to encrypted end to end signalling and it took me 11 weeks to convince my SIP provider who had to convince BT, Virgin Media and O2 to all jump on a call while I explain this and everyone tells me no it's impossible we are not seeing this and then an old boy from BT in 60's who consulta for them now joined the call after coming back from holiday in Lanzarote, heard me explain the issue for the third time and he instantly knew what i was talking about and traced it in 6 test calls. It was a faulty module on a Virgin Media core VOIP switch in Northampton and yes calls were crossing over due to this and when calling our school people were slipping into doctors surgery calls and all sorts.

 

Anyways I digress but the main point is sh#t happens regardless and we need to investigate and eliminate to work out what happened and maybe fix or at least document it for future references.

  • Thanks 1
Posted

Sorry for the delay. Looked up our NICs and we have the following:

S2D Cluster Links: HP Ethernet 10Gb 2-port 530FLR-SFP+ Adapter

Hyper-V & Management: HPE Ethernet 1Gb 4-port 331i Adapter

Posted
Sorry for the delay. Looked up our NICs and we have the following:

S2D Cluster Links: HP Ethernet 10Gb 2-port 530FLR-SFP+ Adapter

Hyper-V & Management: HPE Ethernet 1Gb 4-port 331i Adapter

 

Thanks Rob.

So you have 1GB links to the main network from your cluster and 10GB between the 2 for storage sync?

Posted
re replacement, it's not a cluster but 2 hyper v servers. have set up failover, and testing currently. Yes, diagnostic work has limitations, and in case of suspected hidden firmware bug, even more difficult. All hardware tests on both cluster nodes have passed.
Posted
Thanks Rob.

So you have 1GB links to the main network from your cluster and 10GB between the 2 for storage sync?

Yes thats right. The 1Gbps links are teamed and spread across two core switches.
  • Thanks 1
  • 5 months later...
Posted (edited)

Yes with many thanks to @Tefters

 

The server reporting the issue we found in the end wasn't the server with the issue. We had to shut the cluster down by restarting the one remaining server. But once we did this, when both came back online they synced up again.

 

It seems to be linked to servers having different updates on them. One updates first, shows the I/O error but in truth, it's waiting on the other server updating as well. At least that's what we found with my system.

 

Not ideal shutting down the working server and thus cluster. But worked for me.

 

Just make sure you have completed backups first.

Edited by robyholmes
Posted
Did you ever figure this out? We're experiencing exactly the same symptoms you describe. Just started a few months ago.
How do you usually update your cluster?

 

Do you have CAU on?

 

I learnt very early on with Windows clustering that completely disabling and removing the CAU config was the right move.

 

I update my cluster servers manually every half term by pausing a node via PowerShell (failing over roles and VM's), pausing storage sync, patching a server, unpausing storage sync and then unpausing the node. Wait for the cluster to re-sync fully and then repeat for the other node.

 

This method has never steered me wrong in the last 5 years since migrating from ESXi.

Posted

We don't use CAU. We patch just like you do. Every few months or as needed for any critical updates. A node at a time. Pause/maintenance mode/patch/unpause/disable maintenance/wait for disks to be healthy and storage jobs to complete/rinse, repeat.

 

I seriously hate maintenance on S2D clusters. Microsoft and the hardware vendors really need to fix this. CAU works great on traditional failover clusters without cluster shared volumes. But CAU and WAC/OpenManage have never been consistent and/or safe enough for us to use on any of our S2D clusters. Hoping 2025 will be better, but I'm not holding my breath.

Posted
We used to have this issue too. If one of the cluster nodes rebooted or went offline, and the cluster witness could not be found, it would lock all the disks until both nodes were rebooted. The fix was moving the Cluster File Share Witness to a standalone storage device.
Posted
We used to have this issue too. If one of the cluster nodes rebooted or went offline, and the cluster witness could not be found, it would lock all the disks until both nodes were rebooted. The fix was moving the Cluster File Share Witness to a standalone storage device.
Ideally it would be a physical DC so you could also boot up the cluster properly after a full shutdown and all the FQDN's can resolve.
Posted
Ideally it would be a physical DC so you could also boot up the cluster properly after a full shutdown and all the FQDN's can resolve.

 

What's your experience with Cloud witnesses? We use them here on all our clusters and I've wondered if there may be some communication issue with them that's causing some of our issues. That said, everything tests fine and I've never seen any indication of firewall or web filtering interfering with anything.

Posted (edited)
What's your experience with Cloud witnesses? We use them here on all our clusters and I've wondered if there may be some communication issue with them that's causing some of our issues. That said, everything tests fine and I've never seen any indication of firewall or web filtering interfering with anything.
Personally, I would never fully trust or recommend a cloud witness for an on-prem cluster or vice-versa due to latency, isp, firewall, web filter, cloud provider having an off-day, code change by the provider behind the scenes etc...

 

Cloud witness for cloud (in the same region), on-prem witness for on-prem.

 

I've seen some weird setups that make no sense with witnesses as VM's etc which I don't agree with either.

 

A cluster witness is a basic file share that monitors multiple hosts, I usually recommend 2 scenarios, first scenario is also best practise by Microsoft and anyone who installs clusters at large scale (Dell, HP etc...).

 

1. Physical DC with logon attempts from all clients disabled except for the cluster hosts. Server also hosts DNS. This way if you have to shutdown the whole cluster or have prolonged power cut over UPS uptime you bring this server online first and then can bring up the cluster properly afterwards as AD and DNS are available. Trying to bring a cluster up without DNS while possible is a pain and can cause further issues where you might need to reboot hosts individually afterwards.

 

2. At the very least just make a cr#ppy workstation the cluster witness and shove it in the server rack.

 

My work setup I killed multiple birds with one stone, slight best practise while at the same time slight cringe.

 

I have a 2 node S2D cluster and then later purchase a physical server as a new backup server (Veeam).

I made this server a physical DC, disabled all AD logon requests except for the S2D hosts and installed DNS for name resolution. No DHCP needed as hosts are fixed IP.

Finally I made it the cluster witness.

I then also installed Veeam and use that to backup the 2 cluster hosts (off server backup, VM's only, to save processing power on the hosts) and then have a job to backup itself so then I have a backup of the Veeam server and the Physical DC itself.

 

I would never usually entertain installing something like a backup management software on a DC (or any other software outside basic AD tools) for that matter but it's only used as a physical DC for 2 machines just to logon if they ever have to power up/reboot.

 

In short - 1 box for physical DC, backup management and storage and finally cluster witness.

 

Ow and before anyone's comment on backups on the box they are also pushed to Wasabi on a 30 day immutable GFS setup with a 5 year retention.

 

*Taps brain and smiles* [emoji1]

 

45 minute full cluster backup, 10 minutes self backup (with C drive only as D is backup storage) and then lastly 22 minutes to upload it all to the cloud.

Edited by Tefters

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...