Jump to content

Recommended Posts

Posted (edited)

Hey, thank you in advance for any comments, tips etc.

 

Facing a peculiar issue on two file servers, it was originally just one. However, the second one has started doing the same thing, whereby file access for reading/writing is very slow, and disk active time shows 100% in resource monitor. It feels like a hardware issue; however, the two file servers are at different schools, run on similar but different hardware, and both run Windows Server 2019 Std full patched.

 

The hardware the two VMs run on is similar, HPE Servers; one has an HPE MSA 1050 with 10Gb iSCSI connections (2 each in multipath active active, so 20Gb in total), and the other has 12Gb SAS (2 each in multipath active active so 24Gb in total) with the same HPE MSA 1050.

 

The MSA 1050s themselves each have 24 x 1.2Tb 10k SAS disks in a RAID10 volume.

 

The HV hosts run 2019 Datacenter and are also fully patched.

 

Ther other virtual machines on the failover cluster as well as other large file server arent affected at all.

 

Things we have tried.

  • Turning off Sophos
  • Uninstalling Sophos
  • Turning off Windows Defender and Real-time scanning
  • Deduped the Data to be none dedupe (one server is still deduped)
  • Re-Installed the OS
  • Recreated the VHDX and restored from backup all the data
  • Drivers on the Hyper-V hosts for storage controllers
  • Setup hyper-v replication to a spare none SAN-based server and replicated the server to this, then swapped the servers round live so the primary one was running from the spare server; this seemed to solve the issue strangely, so is this a hardware issue? I moved it back to the failover cluster, and it appeared to be fine but then started with the active time again shortly after.
  • Upgraded Veeam to the latest version as though it might be CBT causing the issue. However, I haven't yet turned off CBT as I can't find a guide on how to achieve this completely.

 

Screenshot of the error with sensitive data covered over.

 

active time.jpg

 

Any pointers would be greatly appreciated.

 

Thanks

Edited by danrhodes
added more info
Posted
We had an issue with a Hyper-V host where the caching had changed to write-through, from write-back, as the battery backup on our RAID controller had dropped below the required threshold, which had a big effect on our disk performance. We changed the battery and ren-enabled write-back and performance went back to normal. A slgihtly different scenario, as this was direct disk storage on the host and not a SAN but perhaps worth checking to see if that has changed in your environment(s) and ruling it out if nothing else.
Posted (edited)
Indexing might be causing high spikes. Compression (users using compression) might also be causing grief.

 

Nice idea, just checked but indexing is off.

Edited by danrhodes
Posted
We had an issue with a Hyper-V host where the caching had changed to write-through, from write-back, as the battery backup on our RAID controller had dropped below the required threshold, which had a big effect on our disk performance. We changed the battery and ren-enabled write-back and performance went back to normal. A slgihtly different scenario, as this was direct disk storage on the host and not a SAN but perhaps worth checking to see if that has changed in your environment(s) and ruling it out if nothing else.

 

Thanks, will check this out but suspect it might not be the case as all the other VM's on the cluster are fine.

 

Cheers

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...