Jump to content

Recommended Posts

Posted

Morning all

 

I have a 3 node Hyper V cluster running Server 2019 and storage spaces

 

Is an in place upgrade to Server 2025 ok to do? or will it cause issues and best to do from scratch?

 

I'm thinking of migrating VM's from one sever to remaining 2 and then removing from cluster, upgrading and then adding back to cluster and moving VM's and repeat the process for the remaining servers.

 

Is that a sensible course of action?

 

Cheers

Posted

Sounds about right to me, that would be my approach. Although i'm sure you'll need to triple check the hardware is s2d 2025 compliant.

 

Also, back up back up back up!

 

i'd also back up each s2d hosts as part for the update process incase a speedy roll back is needed.  

  • Like 1
Posted

I can't speak for 2025 but I did an in-place on my hosts from 2019 to 2022. I migrated everything over to one host and removed the upgrading host from the cluster, did the upgrade as normal, then added it back and verified the cluster. Left it running for four weeks to make sure nothing was going to go wrong, then did the same procedure on the other host. I think then after I was able to raise the functional level of the cluster from the Failover Cluster Manager console, much like you would do with AD and upgrading/replacing your DCs.

  • Like 1
Posted

I was reminded when I did mine in summer - if you team the NICs in Windows on the host, delete the team first, then upgrade, then recreate. If you don't, the team doesn't seem to work and none of your VMs will get an IP address (and may not even start).

  • Like 1
  • 1 month later...
Posted

Careful if you are using S2D and an inplace upgrade.  I had a two node S2D fail on an inplace upgrade (2016 to 2022) and found it easier to just create a fresh cluster and restore from VEEAM.  I now look after a three node S2D on 2022 and will leave that for a while.  I believe the safest way is to evict a server, upgrade, rejoin, leave to resync, repeat, upgrade VM levels as appropriate.  2019 made the cluster upgrades easier, the S2D was the issue for me - it just barfed.

  • Like 2
Posted
On 13/10/2025 at 15:09, 3s-gtech said:

I was reminded when I did mine in summer - if you team the NICs in Windows on the host, delete the team first, then upgrade, then recreate. If you don't, the team doesn't seem to work and none of your VMs will get an IP address (and may not even start).

for the teaming, did you go from 2016?  Was your NIC team a legacy load balance?  That was possibly why, newer server uses switch embedded teaming.  I think you can still set up a legacy team in 2019+ via powershell but it will get fiddly.  I have inplace upgraded a 2019 with switch embedded team to 2022 (not a hypervisor, our onsite VEEAM local backup coordinator) and this upgraded just fine with the team intact.

  • 3 months later...
Posted (edited)

So I did this over the Christmas break and it went well and truly out the window! and broke things and have been limping along.

 

Checked and MS Documents say yes it is supported - but 

  • Quote

    Mixed Mode Limitation: While you can run a 2025 node in a 2019 cluster during the upgrade, you cannot add 2025 nodes to a 2019 cluster permanently; the rolling upgrade must be completed.

     

Did the correct thing, drained Node, Evicted it, in place upgrade and then when it came to rejoin it sat, and the sat and sat and eventually failed due to cluster mixed-version clustering!

 

Marked all the disks in that node as unhealthy, lost communication, blah, blah, blah, basically marked them for death!

 

Bit more googling and came across - https://learn.microsoft.com/en-us/answers/questions/5543921/add-windows-2025-servers-to-an-existing-hyper-v-wi

 

Quote

Migrating virtual machines from a Windows Server 2019 Hyper-V cluster to Windows Server 2025 nodes is a supported scenario, but it does require careful planning. Unfortunately, Windows Server 2025 nodes cannot be directly added to an existing 2019 cluster, as mixed-version clustering is not supported between these two versions

 

Which is what happened to me! - so why does some info say yes it is supported and other info saying not supported!

 

Did I miss something, have I misunderstood something, some self doubt creeping in.

 

I've installed Server 2022 on the now defunct node, but still had issues due the discs in S2D being stuck as removing from pool and basically not being able to remove them in anyway shape or form, as they were not been seen as able to pool, so had to totally wipe them.

 

Ordered some additional disks (which took forever to arrive as I needed some extra capacity in order for the S2D to repair itself!

 

Then Friday on the last day of term, another node has a hissy fit and the only fix was a bios update according to the Dell error report - as it rebooted over and over it took down the entire network for a couple of hours  as I was off on annual leave and had to cancel it and get an uber into work in order to flash the bios and fix the quorum issue in order to get the network back up and running.

 

During half term tried different things to get things back on track and I was seriously considering nuking the whole cluster last week during half term but as per usual, no notice, staff and kids in using the network which kicked those plans into touch.

 

Currently, the original discs are still in the server as they a healthy but S2D thinks they have issues and are still stuck being removed (for now.....))

 

I've added additional discs to each node and waited until my only free night so I could shutdown all VMs, reboot all nodes to try and force anything stuck to kick in, add the new discs into S2D - powered up the VMs and then kicked off a repair which is currently sat at 47 percent on the S2D Repair with approx. 3 to 4 hours to go.

 

If this repair fails, I will have to limp along until I have a free weekend in order to nuke the cluster, and set up from scratch using Server 2025 and restore from Veeam (and yes backups work as have tested them last week)

 

It is that old things, while IT works, no body notices or comments on uptime (was over 4 years) but as soon as it breaks, they all notice and complain and it is the worst thing ever! (to them) and then people start to question if I actually know what I am doing.

 

And now, time for a brew!

 

 

Edited by mdrabble
Posted

I am getting too old for this :censored:  bring back the days of the good old Econet networks and BBC Micro

 

S2D does not want to remove, fix, reset the "unhealthy" disks - which are actually healthy, but S2D for some reason does not want them.

 

I'm at a loss as to what to do next, so the options I think I have left.....

 

1. Pay someone to "fix it"

2.Nuke the cluster, fresh install to Server 2025 on all nodes and the restore data.

 

What are people thoughts?

 

Posted
1 hour ago, mdrabble said:

I am getting too old for this :censored:  bring back the days of the good old Econet networks and BBC Micro

 

S2D does not want to remove, fix, reset the "unhealthy" disks - which are actually healthy, but S2D for some reason does not want them.

 

I'm at a loss as to what to do next, so the options I think I have left.....

 

1. Pay someone to "fix it"

2.Nuke the cluster, fresh install to Server 2025 on all nodes and the restore data.

 

What are people thoughts?

 


Ouch and I thought I'd had full-on time!

I would go nuke and restore - with all the updates/upgrades that have happened, and changes on some nodes but not others, feels far safer to start from a clean sheet. That way if anything goes wrong at that point, you know exactly where everything stands. 


For your S2D are you using mirror/parity/combination of? Now is probably the only opportunity to change this should you wish - my preference is 3 way mirror due to performance.

VM-level restores are generally easy and trouble free. And at least with a clean cluster setup you can also spin up a new VM in advance to test to ensure if there is an issue, if its cluster or the restored VM. 


Also - 2025 Workgroup Cluster is a thing - having spent some time implementing it - not sure its 100% ready yet as has been quite tricky to get fully operational.

  • Like 1
Posted (edited)
45 minutes ago, StephenPink said:


Ouch and I thought I'd had full-on time!

I would go nuke and restore - with all the updates/upgrades that have happened, and changes on some nodes but not others, feels far safer to start from a clean sheet. That way if anything goes wrong at that point, you know exactly where everything stands. 


For your S2D are you using mirror/parity/combination of? Now is probably the only opportunity to change this should you wish - my preference is 3 way mirror due to performance.

VM-level restores are generally easy and trouble free. And at least with a clean cluster setup you can also spin up a new VM in advance to test to ensure if there is an issue, if its cluster or the restored VM. 


Also - 2025 Workgroup Cluster is a thing - having spent some time implementing it - not sure its 100% ready yet as has been quite tricky to get fully operational.

 

Was originally set as a 3 way mirror, but currently degraded with just the 2 nodes fully working.

 

Spoken to the head to let him know the situation and he happy with how I am handling things.

 

I think you are right with the nuke option - I am going to spin up and old server to restore some critical services, eg, VOIP, SIMS, making sure that happy before nuking the cluster - will be done over a weekend (not the next 2 as at Blackpool Brass Band contest this Sunday and then Wife's birthday the following weekend) - so will give me time to build and test the temporary server and do so test restores before proceeding with the nuking of the cluster.

 

So, would you option for 2022 LTSC or go Server 2025?

 

Then will need to think if better to PowerShell script the node configs or use Windows GUI to do each one......

 

 

Edited by mdrabble
Posted
On 26/02/2026 at 10:33, mdrabble said:

 

Was originally set as a 3 way mirror, but currently degraded with just the 2 nodes fully working.

 

Spoken to the head to let him know the situation and he happy with how I am handling things.

 

I think you are right with the nuke option - I am going to spin up and old server to restore some critical services, eg, VOIP, SIMS, making sure that happy before nuking the cluster - will be done over a weekend (not the next 2 as at Blackpool Brass Band contest this Sunday and then Wife's birthday the following weekend) - so will give me time to build and test the temporary server and do so test restores before proceeding with the nuking of the cluster.

 

So, would you option for 2022 LTSC or go Server 2025?

 

Then will need to think if better to PowerShell script the node configs or use Windows GUI to do each one......

 

 

 

Oh yeah of course - I would also recommend a file share witness to help with that! Or think if you have a domain joined cluster there are other options for witness... 

Yeah sounds like a plan - and good luck!

 

Ah sorry I was meaning the Workgroup Cluster specifically - otherwise Server 2025 seems fine. I did a mix of PowerShell and GUI and documented the lot - definitely safer in some ways to script it so definitely identical configs. 

Posted (edited)

Where are you with this at the moment, do you still have 2 nodes running 2019 and one running 2022?

If so is the 2022 node still failing to join the cluster because of the storage & if you run this command from an admin PowerShell on the 2022 node what does the output look like? 

Get-PhysicalDisk

 

I had issues with storage when I broke a node previously and might have a few cmdlets that got things back online.

 

If you're certain you're not going to lose data, this should reset the disks -maybe! I think you'd run this on the 2022 node to reset the storage and then attempt to join the node to the cluster again:

Get-PhysicalDisk | Reset-PhysicalDisk -ErrorAction SilentlyContinue

Get-Disk | ? Number -ne $null | ? IsBoot -ne $true | ? IsSystem -ne $true | ? PartitionStyle -ne RAW | ? BusType -ne USB | % {
    $_ | Set-Disk -isoffline:$false
    $_ | Set-Disk -isreadonly:$false
    $_ | Clear-Disk -RemoveData -RemoveOEM -Confirm:$false
    $_ | Set-Disk -isreadonly:$true
    $_ | Set-Disk -isoffline:$true
}

Get-Disk | Where Number -Ne $Null | Where IsBoot -Ne $True | Where IsSystem -Ne $True | Where PartitionStyle -Eq RAW | Group -NoElement -Property FriendlyName

 

Edited by ThomL
Posted

Cheers for the script, it won’t let me reset the disks as a repairs keeps trying to kick in.

 

I’m just waiting for the current job to finished which has been running for over 4 days!

Posted

So at the moment you've got an active storage job trying to repair the storage? what does get-storagejob return at the moment (powershell)?

 

What is the state of your cluster currently, do you still have 2 nodes running 2019 and one running 2022 with the 2022 node causing problems, and it thinks it's running a 3 node configuration?

Posted

Cluster is a 3 way mirror, nodes 2 and 3 running Server 2019 and are Running Degraded but stable.

 

Node 1 is now Server 2022

 

Of the 2 NVMe Drives that are Journal Disks, 1 is find but the other is flagged as {Removing From Pool, Transient Error} Unhealthy

 

The HDD are similar, in that some are same error, or Lost Connections or {Removing From Pool, OK} Healthy 

 

I suspect the repair/regeneration is down to the drives being flagged as Unhealthy and the Pool keeps trying to evaluate things (although I could be wrong)

 

Repair is currently on just over 5 days the CSV Regeneration is currently on under 7 hours (they were on the same timings previously)

 

I'm hesitant to use ChatGPT or another AI as they are not always correct - but it may be time to stick in what has happened and let AI see if there is a way to fix things.

 

Node 1 has no VMs, not CSV attached to it, so i am wondering if to evict again run Update-StoragePool and then reboot node 2 wait until back up and happy before rebooting node 3

 

Once Both happy (as happy can be with a missing node) wipe all the disks on Node 1 - using either diskpart of powershell to remove all data, to ensure any meta data is removed.

 

Check any pool flags on the pool and try and resolve/remove them before adding node 1 back into the mix

 

Or Contact either Microsoft or any company to come and fix the issue.

 

Either way I've not been able to relax as if anything happens to another node then it squeaky bottom time!

Posted (edited)

I think it depends on where the drives with the issues are - are there drives with issues on nodes 2 and 3?

 

I'd be looking to get all drives on those nodes healthy, then after reviewing storage jobs I'd evict node 1 (server 2022) - clean the disks with the script above, check network settings and then with the drive all checked and looking ready to join the cluster I'd add the node back to the cluster, add the drives to the storage pool etc. wait for storage jobs to complete and at this point  you should be stable again. Once stable evict node 2, get server 2022 installed, reintroduced to the cluster, wait for stability and follow the same process for node 3.

 

Then up the functional level of the cluster and start the update process again to get to server 2025 if desired.

 

This should be salvageable, keep calm and logical - no rash powershell commands or reboots!

Edited by ThomL
Posted

All the drives with issues are on node 1, cleaned the disks, added the node back and still stuck in the repair loop.

 

The issue is that the cluster still sees the old drives and is trying to remove them but won't remove them until it repairs.

 

Cannot remove the drives as the cluster says there is not enough capacity - (as it is still trying to reference the disks and therefore wouldn't have enough capacity to keep things repairable - is how i understand things)

 

Adding them back should say fresh disks and kicks in to repair mode - but then gets stuck in a repair loop. (is my understanding and take on things)

 

 

Decided to run things through ChatGPT and it more or less gave me similar commands to run, although it wanted me to try and remove disks using friendly names which would have affected all nodes 😁 which when I pointed out it kept saying good catch! - Good job I don't blindly trust the info it gave me.

 

It confirmed it might be in a repair loop but insisted it was fixable and kept giving me the same commands over and over.

 

Checked the StorageJob and it have processed 0% with 0 bytes processed for 15+ hours on the CSV regeneration and Cluster Performance Regeneration and then nearly 6 days on the CSV Repair.

 

Other than purchasing brand new drives is to nuke the cluster.

 

I have an old HP DL380p server with 128Gb Ram and 8TB of Storage, so will test restores on that and check the VMs run, so When I nuke the cluster, I can have the Main Systems online should I not have things up and running in time.

 

 

Posted (edited)

If you shut down node 1 does the cluster/VMs/CSV stay online?

 

If so, bring the node back online without networking so it doesn't join the cluster and you can mess with it: blast the storage to get it in a good state, I think this is the powershell code that removes the S2D metadata: 

Get-PhysicalDisk | Reset-PhysicalDisk

With this disks all clean you should be able to re introduce to the cluster, at which point it will see it has the storage it needs to rebuild as the disks will show as healthy.

 

I had an issue similar to this when first testing S2D years a go - as you've mentioned the cluster was already aware of the disks and wouldn't use them - I', pretty certain the script I posted earlier got the disks back to a healthy state and ready for reuse... Might have to nuke the node and reinstall windows server 2022 then run the script to clean the drives and then look to rejoin the cluster/get S2D things rolling

 

Get-PhysicalDisk | Reset-PhysicalDisk

Get-Disk | ? Number -ne $null | ? IsBoot -ne $true | ? IsSystem -ne $true | ? PartitionStyle -ne RAW | ? BusType -ne USB | % {
    $_ | Set-Disk -isoffline:$false
    $_ | Set-Disk -isreadonly:$false
    $_ | Clear-Disk -RemoveData -RemoveOEM -Confirm:$false
    $_ | Set-Disk -isreadonly:$true
    $_ | Set-Disk -isoffline:$true
}

Get-Disk | Where Number -Ne $Null | Where IsBoot -Ne $True | Where IsSystem -Ne $True | Where PartitionStyle -Eq RAW | Group -NoElement -Property FriendlyName
Edited by ThomL
  • Thanks 1
Posted

Yea the CSV stays online without Node 1 being available.

 

will have a play later and let you know how I get on.

 

thanks for the help 🙂

Posted

One time I broke an S2D cluster and we paid for a single incident ticket with Microsoft to resolve, it took a little while to get past the lower levels of support - but once escalated to the correct level you could see they knew the product well and have the cluster back online and repairing in maybe just over an hour. The cluster had totally fallen over at the time and was in a worse state than you describe your cluster.

 

I think we got the ticket going via this web portal: https://support.serviceshub.microsoft.com/supportforbusiness Make sure to sign in/setup an account that isn't your school Office 365 account - these don't work on the portal I've just read, might need to use personal email account or a google account or similar to get a ticket open. I also don't remember how much we paid - possibly £200-£400, but this was 7/8 years a go.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...