Jump to content

Recommended Posts

Posted

Hi All,

 

I've found one of our two Hyper-V S2D Cluster nodes is showing all of it's storage disks (SSDs) as in a IO Error state. The drives have a warning all of them on this server. VMs are running on the other server currently and won't migrate to the affected server.

 

What's odd, is when I looking this error on this Microsoft page, it seems to imply the issue would only affect a single disk, and yet it's affecting all Storage Pool disks on this one server. There are 8 disks in each server, two are running the OS (RAID1) and 6 are in the storage pool. It's all these 6 that have this error.

https://learn.microsoft.com/en-us/windows-server/storage/storage-spaces/storage-spaces-states#drive-health-state-warning

 

Things I've tried so far are:

- Restart

- Reset Disk & Repair Virtual Disk (As shown in the above Microsoft guide)

- Update server Drives & Firmware (But not the drives themselves, I can't get hold of the firmware for them)

- Pull a drive out when in the OS - This is interesting, HPs Smart Storage tool shows it gone, but the Storage Pool does not. After a restart with the drive still out, Storage Pool still shows it with the same error, but the chassis / enclosure slot is blank. Inserting the drive again when in the OS doesn't change anything (Other than HP Smart Storage), a restart again returns the drive slot, but with the same error.

- Tried to remove one drive from the storage pool, but this can't be done as the pool is in an unhealthy state so it won't allow you to remove it

 

The server was purchase from ICT-Direct, with which it's still under warranty. They've have a look if they can get the firmware for the drivers but can't find it either. The drives are Samsung 7L34800, MZ7L3480HCHQ or PM893 I understand. Which seems to be a OEM drive so Samsung don't publish anything for it. HP's SUM updater doesn't see them either.

 

We use cluster aware updating and I think this server has installed an update, restarted and then done this. The other host is waiting to restart but can't because it can't migrate the roles. I'm worried it might have the same issue once it's restarted.

 

I'm not seeing much in event viewer for this either, these bits around but nothing that gives me anything via a google search. Google is also turning up very little other than the Microsoft link above (Plus a Lenovo copy of it) when I look up the IO Error for Storage Spaces.

 

I was hoping to upload a screenshot of the drives and the error message, but it seems today is not my day and I'm getting a Failed to upload error.

 

Help?

 

Thanks,

Rob

Posted (edited)
Hi All,

 

I've found one of our two Hyper-V S2D Cluster nodes is showing all of it's storage disks (SSDs) as in a IO Error state. The drives have a warning all of them on this server. VMs are running on the other server currently and won't migrate to the affected server.

 

What's odd, is when I looking this error on this Microsoft page, it seems to imply the issue would only affect a single disk, and yet it's affecting all Storage Pool disks on this one server. There are 8 disks in each server, two are running the OS (RAID1) and 6 are in the storage pool. It's all these 6 that have this error.

https://learn.microsoft.com/en-us/windows-server/storage/storage-spaces/storage-spaces-states#drive-health-state-warning

 

Things I've tried so far are:

- Restart

- Reset Disk & Repair Virtual Disk (As shown in the above Microsoft guide)

- Update server Drives & Firmware (But not the drives themselves, I can't get hold of the firmware for them)

- Pull a drive out when in the OS - This is interesting, HPs Smart Storage tool shows it gone, but the Storage Pool does not. After a restart with the drive still out, Storage Pool still shows it with the same error, but the chassis / enclosure slot is blank. Inserting the drive again when in the OS doesn't change anything (Other than HP Smart Storage), a restart again returns the drive slot, but with the same error.

- Tried to remove one drive from the storage pool, but this can't be done as the pool is in an unhealthy state so it won't allow you to remove it

 

The server was purchase from ICT-Direct, with which it's still under warranty. They've have a look if they can get the firmware for the drivers but can't find it either. The drives are Samsung 7L34800, MZ7L3480HCHQ or PM893 I understand. Which seems to be a OEM drive so Samsung don't publish anything for it. HP's SUM updater doesn't see them either.

 

We use cluster aware updating and I think this server has installed an update, restarted and then done this. The other host is waiting to restart but can't because it can't migrate the roles. I'm worried it might have the same issue once it's restarted.

 

I'm not seeing much in event viewer for this either, these bits around but nothing that gives me anything via a google search. Google is also turning up very little other than the Microsoft link above (Plus a Lenovo copy of it) when I look up the IO Error for Storage Spaces.

 

I was hoping to upload a screenshot of the drives and the error message, but it seems today is not my day and I'm getting a Failed to upload error.

 

Help?

 

Thanks,

Rob

 

I saw a couple of people post in the Patch Tuesday Megathread on Reddit about the Jan 2024 cumulative server updates causing I/O errors on failover clusters. In response, I skipped these updates but have since applied the Feb 2024 ones without any issues.

Edited by gybe78
  • Thanks 1
Posted

@Tefters - Sometimes a solution is staring you right in the face, rolling back the update would make sense to try wouldn't it! I'll give that a go especially in light of gybe78 post. Checked device manager, all drives are showing fine and not reporting any errors. The only errors I can find in Event Manager are deep down in StorageManagement-PartUtil which is reporting 'Failed to get disk properties'. History only goes back to 26th Feb for this. The last windows updates installed on the 20th so can't be sure it's linked with that.

 

I have however found errors in another event log now, FailoverClustering-StorageBusClient which on the 20th February 2024 at 00:16, right when Windows Update was installing the 2024-02 Cumulative Update

@gybe78 You might be on to something, but as you can read above, it appears to be the 2024-02 Cumulative update, not the 2024-02 one. I'll uninstall it and see what happens. The other host is likely sat waiting to install the same one right now.

Posted

Well uninstalled KB5034770 (02-2024), restarted. Problem remains. Nothing appears to have changed either.

 

It's now showing KB5034129 (01-2024) can be uninstalled, but when I try and uninstall it, it get gives me 'An error occurred, uninstall failed'. I'm guessing it doesn't want to roll back two updates worth.

Posted (edited)

Still no look with uninstalling KB5034129 (01-2024), DSIM doesn't see it as installed. Control Panel>Installed Updates>Uninstall an update does show it as installed, but won't remove it. It's install date is today so it's showing it due to the rollback. I'll try another restart to see if that helps.

 

Ran a cluster validation check. A number of warnings mainly about a VM that's offline (Expect) and that this one host/node is paused (Which I did so it didn't keep trying to move VMs to it). The network was in a error state, but this is because the network drivers are different between the servers. This is expected as I updated the drivers on this HOST to try and solve the IO error. The cluster is still communicating over the LAN & Cluster link so I think this is safe to ignore for now. I can't update the other host now anyway as it'll kick all VMs offline.

The Storage Spaces area of the Validation report is all green, not even a mention of half the drives being in a warning state. So that's good isn't it. This is the thing, it didn't show these drives being in a warning state. I only found one because the other host wasn't able to restart for windows updates because it couldn't move the VMs to the first Host.

 

EDIT: Another restart still no luck with uninstalling update. Drives all still show IO Error. Looking more and more like it's a full OS re-build.

Edited by robyholmes
Posted

 

EDIT: Another restart still no luck with uninstalling update. Drives all still show IO Error. Looking more and more like it's a full OS re-build.

 

been a long time since i touched s2d failovers as we are vmware but can you take the disks offline in disk manger ( maybe format) and reattach or is the os on the same volume?

Posted

My next chain of thought here and note its only thought as I haven't done this in practice but its the route i would personally go down.

 

If device manager shows the disks as being ok and so does everything else then chances are its just a MS storage pool f*** up ?!

Based on this move all VM's onto the working server (if not already) and then remove the non-working server from the cluster and its disks from the storage pool and then try re-adding them.

 

s2d cluster remove server from cluster pool - Search (bing.com)

Posted

I've tried removing the disks from the pool, but it won't allow it as it's already degraded. I'll try removing the server itself from the pool.

 

The other thought I had, was do I try physically move the disks from HOSTA to HOSTB (Working Host). It has bays to allow this and looking up, it should be supported in Storage Sense. That way I'd know if the disk/s came back online, that the disk was fine. Just risks upsetting the other host.

Posted
I've tried removing the disks from the pool, but it won't allow it as it's already degraded. I'll try removing the server itself from the pool.

 

The other thought I had, was do I try physically move the disks from HOSTA to HOSTB (Working Host). It has bays to allow this and looking up, it should be supported in Storage Sense. That way I'd know if the disk/s came back online, that the disk was fine. Just risks upsetting the other host.

 

I would strongly recommend AGAINST that idea, you do not want to affect the data on the working disks or the host for that matter! Not unless you are prepared for it all to possibly go Pete Tong and you left with an entire cluster rebuild and DR restore.

Posted
I would strongly recommend AGAINST that idea, you do not want to affect the data on the working disks or the host for that matter! Not unless you are prepared for it all to possibly go Pete Tong and you left with an entire cluster rebuild and DR restore.

 

Well that's it. That's not really something I want to be dealing with as well.

 

I'll look to remove the host then, then re-add and see what happens. See if it'll even let me remove it in this state.

Posted
Well that's it. That's not really something I want to be dealing with as well.

 

I'll look to remove the host then, then re-add and see what happens. See if it'll even let me remove it in this state.

1. Remove Server

2. Check Storage Pool to see if disks removed

3. Remove disks manually via PowerShell if possible

 

That where i'm at in terms of thinking at the moment.

Posted

I've removed the server from the cluster, it's gone from the Enclosures list in Failover Cluster Manager, but the disks are still present but with no enclosure shown (Which makes sense).

 

In Server Manager>Storage Spaces however it's now showing double the amount of disks for the failed server when I view the Storage Pool, as shown here:

Screenshot 2024-02-29 154740.jpg

 

I restarted the affected server, which is when the drives appeared twice. I'm thinking I might be best leaving it for a abit, see what it shows later. I'm worried about removing the disks for the pool now it's showing them twice.

Posted (edited)
The status on the drives without a chassis shown, is 'Lost communication'. The status on the drives with the chassis shown is 'IO Error'.

 

Do the disks still show in the storage pool in cluster manager?

Edited by Tefters
Posted (edited)

Yes they do:

Screenshot 2024-02-29 161613.jpg

 

It's why I keep thinking the disks aren't talking to the OS/Cluster. So it doesn't know they have gone, or are back etc. Especially as it's all of them at once. HP Smart Storage tool however is reading them just fine.

Screenshot 2024-02-29 154740.jpg

Edited by robyholmes
Correct screenshot
Posted

So the node is completely removed but it's disks are still in the storage pool.

My best guess is it's a cluster issue not a server/disk issue as those disks shouldn't be showing.

 

In cluster aware manager under all the side tabs and tabs at the bottom of each page has the old server references all gone?

Does PowerShell if you do a node query on show 1 node in the cluster?

Posted

The node isn't shown Failover Cluster Manager. If I look under Storage>Pool>Physical Disks, all the disks including those from the removed server are shown (But without an enclosure or slot number)

If I look under Exclosures, only the remaining servers enclosures are shown.

 

Powershell running the Get-ClusterNode command only shows the one HOST server now, which is correct.

Posted

No I didn't. I removed it from the Failover Cluster Manager UI. Running that command now doesn't work as the node is already gone.

 

Running Get-PhysicalDisk on the working server shows all the drives, including those from the other removed server. With the IO Error shown again.

Posted

Ok so I just watched this to confirm my thoughts on CSV's

 

So what you currently have is a 1 node compute cluster with a CSV storage pool across 2 servers.

 

The state you ideally need to get to is a 1 node cluster running with a storage pool on 1 node so you can then tackle the faulty server/storage and then implement it again once fixed. Hopefully that seems logical.

 

I'm trying to piece together the best way I would go about this myself so give me a few

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...