robyholmes Posted February 28, 2024 Posted February 28, 2024 Hi All, I've found one of our two Hyper-V S2D Cluster nodes is showing all of it's storage disks (SSDs) as in a IO Error state. The drives have a warning all of them on this server. VMs are running on the other server currently and won't migrate to the affected server. What's odd, is when I looking this error on this Microsoft page, it seems to imply the issue would only affect a single disk, and yet it's affecting all Storage Pool disks on this one server. There are 8 disks in each server, two are running the OS (RAID1) and 6 are in the storage pool. It's all these 6 that have this error. https://learn.microsoft.com/en-us/windows-server/storage/storage-spaces/storage-spaces-states#drive-health-state-warning Things I've tried so far are: - Restart - Reset Disk & Repair Virtual Disk (As shown in the above Microsoft guide) - Update server Drives & Firmware (But not the drives themselves, I can't get hold of the firmware for them) - Pull a drive out when in the OS - This is interesting, HPs Smart Storage tool shows it gone, but the Storage Pool does not. After a restart with the drive still out, Storage Pool still shows it with the same error, but the chassis / enclosure slot is blank. Inserting the drive again when in the OS doesn't change anything (Other than HP Smart Storage), a restart again returns the drive slot, but with the same error. - Tried to remove one drive from the storage pool, but this can't be done as the pool is in an unhealthy state so it won't allow you to remove it The server was purchase from ICT-Direct, with which it's still under warranty. They've have a look if they can get the firmware for the drivers but can't find it either. The drives are Samsung 7L34800, MZ7L3480HCHQ or PM893 I understand. Which seems to be a OEM drive so Samsung don't publish anything for it. HP's SUM updater doesn't see them either. We use cluster aware updating and I think this server has installed an update, restarted and then done this. The other host is waiting to restart but can't because it can't migrate the roles. I'm worried it might have the same issue once it's restarted. I'm not seeing much in event viewer for this either, these bits around but nothing that gives me anything via a google search. Google is also turning up very little other than the Microsoft link above (Plus a Lenovo copy of it) when I look up the IO Error for Storage Spaces. I was hoping to upload a screenshot of the drives and the error message, but it seems today is not my day and I'm getting a Failed to upload error. Help? Thanks, Rob
robyholmes Posted February 28, 2024 Author Posted February 28, 2024 Screenshot now attached (Thanks ZeroHour)
Tefters Posted February 28, 2024 Posted February 28, 2024 If you think it's occurred due to a recent update have you tried rolling back the update? 1
Tefters Posted February 28, 2024 Posted February 28, 2024 Also what are the errors in the event log and is device manager showing any anomalies for the host such as raid controller driver issues?
Tefters Posted February 28, 2024 Posted February 28, 2024 Ow and cluster validation as well. Worth running that as it might spit out some information that's of use.
5tu Posted February 28, 2024 Posted February 28, 2024 (edited) Hi All, I've found one of our two Hyper-V S2D Cluster nodes is showing all of it's storage disks (SSDs) as in a IO Error state. The drives have a warning all of them on this server. VMs are running on the other server currently and won't migrate to the affected server. What's odd, is when I looking this error on this Microsoft page, it seems to imply the issue would only affect a single disk, and yet it's affecting all Storage Pool disks on this one server. There are 8 disks in each server, two are running the OS (RAID1) and 6 are in the storage pool. It's all these 6 that have this error. https://learn.microsoft.com/en-us/windows-server/storage/storage-spaces/storage-spaces-states#drive-health-state-warning Things I've tried so far are: - Restart - Reset Disk & Repair Virtual Disk (As shown in the above Microsoft guide) - Update server Drives & Firmware (But not the drives themselves, I can't get hold of the firmware for them) - Pull a drive out when in the OS - This is interesting, HPs Smart Storage tool shows it gone, but the Storage Pool does not. After a restart with the drive still out, Storage Pool still shows it with the same error, but the chassis / enclosure slot is blank. Inserting the drive again when in the OS doesn't change anything (Other than HP Smart Storage), a restart again returns the drive slot, but with the same error. - Tried to remove one drive from the storage pool, but this can't be done as the pool is in an unhealthy state so it won't allow you to remove it The server was purchase from ICT-Direct, with which it's still under warranty. They've have a look if they can get the firmware for the drivers but can't find it either. The drives are Samsung 7L34800, MZ7L3480HCHQ or PM893 I understand. Which seems to be a OEM drive so Samsung don't publish anything for it. HP's SUM updater doesn't see them either. We use cluster aware updating and I think this server has installed an update, restarted and then done this. The other host is waiting to restart but can't because it can't migrate the roles. I'm worried it might have the same issue once it's restarted. I'm not seeing much in event viewer for this either, these bits around but nothing that gives me anything via a google search. Google is also turning up very little other than the Microsoft link above (Plus a Lenovo copy of it) when I look up the IO Error for Storage Spaces. I was hoping to upload a screenshot of the drives and the error message, but it seems today is not my day and I'm getting a Failed to upload error. Help? Thanks, Rob I saw a couple of people post in the Patch Tuesday Megathread on Reddit about the Jan 2024 cumulative server updates causing I/O errors on failover clusters. In response, I skipped these updates but have since applied the Feb 2024 ones without any issues. Edited February 28, 2024 by gybe78 1
robyholmes Posted February 29, 2024 Author Posted February 29, 2024 @Tefters - Sometimes a solution is staring you right in the face, rolling back the update would make sense to try wouldn't it! I'll give that a go especially in light of gybe78 post. Checked device manager, all drives are showing fine and not reporting any errors. The only errors I can find in Event Manager are deep down in StorageManagement-PartUtil which is reporting 'Failed to get disk properties'. History only goes back to 26th Feb for this. The last windows updates installed on the 20th so can't be sure it's linked with that. I have however found errors in another event log now, FailoverClustering-StorageBusClient which on the 20th February 2024 at 00:16, right when Windows Update was installing the 2024-02 Cumulative Update @gybe78 You might be on to something, but as you can read above, it appears to be the 2024-02 Cumulative update, not the 2024-02 one. I'll uninstall it and see what happens. The other host is likely sat waiting to install the same one right now.
robyholmes Posted February 29, 2024 Author Posted February 29, 2024 Well uninstalled KB5034770 (02-2024), restarted. Problem remains. Nothing appears to have changed either. It's now showing KB5034129 (01-2024) can be uninstalled, but when I try and uninstall it, it get gives me 'An error occurred, uninstall failed'. I'm guessing it doesn't want to roll back two updates worth.
robyholmes Posted February 29, 2024 Author Posted February 29, 2024 (edited) Still no look with uninstalling KB5034129 (01-2024), DSIM doesn't see it as installed. Control Panel>Installed Updates>Uninstall an update does show it as installed, but won't remove it. It's install date is today so it's showing it due to the rollback. I'll try another restart to see if that helps. Ran a cluster validation check. A number of warnings mainly about a VM that's offline (Expect) and that this one host/node is paused (Which I did so it didn't keep trying to move VMs to it). The network was in a error state, but this is because the network drivers are different between the servers. This is expected as I updated the drivers on this HOST to try and solve the IO error. The cluster is still communicating over the LAN & Cluster link so I think this is safe to ignore for now. I can't update the other host now anyway as it'll kick all VMs offline. The Storage Spaces area of the Validation report is all green, not even a mention of half the drives being in a warning state. So that's good isn't it. This is the thing, it didn't show these drives being in a warning state. I only found one because the other host wasn't able to restart for windows updates because it couldn't move the VMs to the first Host. EDIT: Another restart still no luck with uninstalling update. Drives all still show IO Error. Looking more and more like it's a full OS re-build. Edited February 29, 2024 by robyholmes
pablo007 Posted February 29, 2024 Posted February 29, 2024 EDIT: Another restart still no luck with uninstalling update. Drives all still show IO Error. Looking more and more like it's a full OS re-build. been a long time since i touched s2d failovers as we are vmware but can you take the disks offline in disk manger ( maybe format) and reattach or is the os on the same volume?
Tefters Posted February 29, 2024 Posted February 29, 2024 My next chain of thought here and note its only thought as I haven't done this in practice but its the route i would personally go down. If device manager shows the disks as being ok and so does everything else then chances are its just a MS storage pool f*** up ?! Based on this move all VM's onto the working server (if not already) and then remove the non-working server from the cluster and its disks from the storage pool and then try re-adding them. s2d cluster remove server from cluster pool - Search (bing.com)
robyholmes Posted February 29, 2024 Author Posted February 29, 2024 I've tried removing the disks from the pool, but it won't allow it as it's already degraded. I'll try removing the server itself from the pool. The other thought I had, was do I try physically move the disks from HOSTA to HOSTB (Working Host). It has bays to allow this and looking up, it should be supported in Storage Sense. That way I'd know if the disk/s came back online, that the disk was fine. Just risks upsetting the other host.
Tefters Posted February 29, 2024 Posted February 29, 2024 I've tried removing the disks from the pool, but it won't allow it as it's already degraded. I'll try removing the server itself from the pool. The other thought I had, was do I try physically move the disks from HOSTA to HOSTB (Working Host). It has bays to allow this and looking up, it should be supported in Storage Sense. That way I'd know if the disk/s came back online, that the disk was fine. Just risks upsetting the other host. I would strongly recommend AGAINST that idea, you do not want to affect the data on the working disks or the host for that matter! Not unless you are prepared for it all to possibly go Pete Tong and you left with an entire cluster rebuild and DR restore.
robyholmes Posted February 29, 2024 Author Posted February 29, 2024 I would strongly recommend AGAINST that idea, you do not want to affect the data on the working disks or the host for that matter! Not unless you are prepared for it all to possibly go Pete Tong and you left with an entire cluster rebuild and DR restore. Well that's it. That's not really something I want to be dealing with as well. I'll look to remove the host then, then re-add and see what happens. See if it'll even let me remove it in this state.
Tefters Posted February 29, 2024 Posted February 29, 2024 Well that's it. That's not really something I want to be dealing with as well. I'll look to remove the host then, then re-add and see what happens. See if it'll even let me remove it in this state. 1. Remove Server 2. Check Storage Pool to see if disks removed 3. Remove disks manually via PowerShell if possible That where i'm at in terms of thinking at the moment.
robyholmes Posted February 29, 2024 Author Posted February 29, 2024 I've removed the server from the cluster, it's gone from the Enclosures list in Failover Cluster Manager, but the disks are still present but with no enclosure shown (Which makes sense). In Server Manager>Storage Spaces however it's now showing double the amount of disks for the failed server when I view the Storage Pool, as shown here: I restarted the affected server, which is when the drives appeared twice. I'm thinking I might be best leaving it for a abit, see what it shows later. I'm worried about removing the disks for the pool now it's showing them twice.
robyholmes Posted February 29, 2024 Author Posted February 29, 2024 The status on the drives without a chassis shown, is 'Lost communication'. The status on the drives with the chassis shown is 'IO Error'.
Tefters Posted February 29, 2024 Posted February 29, 2024 (edited) The status on the drives without a chassis shown, is 'Lost communication'. The status on the drives with the chassis shown is 'IO Error'. Do the disks still show in the storage pool in cluster manager? Edited February 29, 2024 by Tefters
robyholmes Posted February 29, 2024 Author Posted February 29, 2024 (edited) Yes they do: It's why I keep thinking the disks aren't talking to the OS/Cluster. So it doesn't know they have gone, or are back etc. Especially as it's all of them at once. HP Smart Storage tool however is reading them just fine. Edited February 29, 2024 by robyholmes Correct screenshot
Tefters Posted February 29, 2024 Posted February 29, 2024 So the node is completely removed but it's disks are still in the storage pool. My best guess is it's a cluster issue not a server/disk issue as those disks shouldn't be showing. In cluster aware manager under all the side tabs and tabs at the bottom of each page has the old server references all gone? Does PowerShell if you do a node query on show 1 node in the cluster?
Tefters Posted February 29, 2024 Posted February 29, 2024 I'm also then thinking do you need to remove the disks from the storage pool via PowerShell, something like https://serverfault.com/questions/983824/how-to-remove-a-perfectly-good-disk-s2d-storage-spaces-direct-so-it-can-be-used
robyholmes Posted February 29, 2024 Author Posted February 29, 2024 The node isn't shown Failover Cluster Manager. If I look under Storage>Pool>Physical Disks, all the disks including those from the removed server are shown (But without an enclosure or slot number) If I look under Exclosures, only the remaining servers enclosures are shown. Powershell running the Get-ClusterNode command only shows the one HOST server now, which is correct.
Tefters Posted February 29, 2024 Posted February 29, 2024 Reading this https://learn.microsoft.com/en-us/windows-server/storage/storage-spaces/remove-servers did you use the cleanup disks flag half way down the page?
robyholmes Posted February 29, 2024 Author Posted February 29, 2024 No I didn't. I removed it from the Failover Cluster Manager UI. Running that command now doesn't work as the node is already gone. Running Get-PhysicalDisk on the working server shows all the drives, including those from the other removed server. With the IO Error shown again.
Tefters Posted February 29, 2024 Posted February 29, 2024 Ok so I just watched this to confirm my thoughts on CSV's So what you currently have is a 1 node compute cluster with a CSV storage pool across 2 servers. The state you ideally need to get to is a 1 node cluster running with a storage pool on 1 node so you can then tackle the faulty server/storage and then implement it again once fixed. Hopefully that seems logical. I'm trying to piece together the best way I would go about this myself so give me a few
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now