Pyroman Posted May 18, 2021 Posted May 18, 2021 Hiya, I'm banging my head against the wall, our Main server keeps randomly freezing. The first time it got noticed was when it took all the papercut printers out because they couldn't get a connection to the server. So far, I've ruled out that it's not Veeam and I'm struggling to see any software that's pegging the CPU at 100% or using all the memory or anything obvious. I tried to download a file on chrome while its had this issue and it freezes and when it come back the download has failed so it's obviously affecting the network connection as well. Program windows randomly say Not Responding and I can't click anything on the taskbar and then 5 mins later it comes back to life and works for a short while. Installs hang, uninstalls hang. I'm going to take it down after work tonight and do a memory and HDD diagnosis to rule those out but it's been running along happily for a few years now and all of a sudden with no changes it's started being a massive pain. I'm working by myself so it's difficult not having someone to bounce ideas off!
3s-gtech Posted May 18, 2021 Posted May 18, 2021 Any RAID diagnostic tools installed? That could be a dying disk or controller. Any warning codes or lights on the front panel?
Pyroman Posted May 18, 2021 Author Posted May 18, 2021 Cheers, it's a DELL T130, no warning lights on the front. I'm just trying to install openmanage now but obviously having issues because everything keeps freezing!!
mavhc Posted May 18, 2021 Posted May 18, 2021 Don't forget to check the IPMI. Also in task manager you can check the Status column for Suspended, and right click and Analyse Wait Chain
southhamster Posted May 18, 2021 Posted May 18, 2021 Do you have an idrac you can connect to? May give useful info.
Pyroman Posted May 18, 2021 Author Posted May 18, 2021 So just to add more fun into the mix, we've just had a powercut!!! It did give me the chance to run the Dell BIOS diagnostics though and no HDD or RAM issues, or any issues at all for that matter were reported. Unfortunately I was half way through installing openmanage when it died so I'll try again!
Foresthippy Posted May 18, 2021 Posted May 18, 2021 The RAID on T130 is, by default, an S130 software RAID that gets particularly poor reviews. Just thought that was worth mentioning as could be part of the problem.
Pyroman Posted May 18, 2021 Author Posted May 18, 2021 So we have the PERC H330 adapter and according to iDRAC everything is hunky dory
Leemanator Posted May 18, 2021 Posted May 18, 2021 (edited) Are you getting anything in the event viewer like the below? Log Name: Application Source: ESENT Date: 17/05/2021 12:50:38 Event ID: 533 Task Category: General Level: Warning Keywords: Classic User: N/A Computer: ******* Description:svchost (3364,T,0) SRUJet: A request to write to the file "C:\WINDOWS\system32\SRU\SRU.chk" at offset 0 (0x0000000000000000) for 4096 (0x00001000) bytes has not completed for 36 second(s). This problem is likely due to faulty hardware. Please contact your hardware vendor for further assistance diagnosing the problem. We've had a handful of Dell Poweredge Servers start freezing up randomly in the past couple of weeks, with event's like these showing writes are taking a long time. No hardware errors in open manage or when running diagnostics. Suspect it's to do with the April cumulative but can't see a pattern between the servers affected and the servers unaffected as yet. Edited May 18, 2021 by Leemanator
Pyroman Posted May 18, 2021 Author Posted May 18, 2021 Are you getting anything in the event viewer like the below? Log Name: Application Source: ESENT Date: 17/05/2021 12:50:38 Event ID: 533 Task Category: General Level: Warning Keywords: Classic User: N/A Computer: ******* Description:svchost (3364,T,0) SRUJet: A request to write to the file "C:\WINDOWS\system32\SRU\SRU.chk" at offset 0 (0x0000000000000000) for 4096 (0x00001000) bytes has not completed for 36 second(s). This problem is likely due to faulty hardware. Please contact your hardware vendor for further assistance diagnosing the problem. We've had a handful of Dell Poweredge Servers start freezing up randomly in the past couple of weeks, with event's like these showing writes are taking a long time. No hardware errors in open manage or when running diagnostics. Suspect it's to do with the April cumulative but can't see a pattern between the servers affected and the servers unaffected as yet. YES!!! Any idea which KB that is off the top of your head? I'll try uninstalling it. svchost (1112) SoftwareUsageMetrics-Svc: A request to write to the file "C:\Windows\system32\LogFiles\Sum\Svctmp.log" at offset 0 (0x0000000000000000) for 4096 (0x00001000) bytes succeeded, but took an abnormally long time (19 seconds) to be serviced by the OS. This problem is likely due to faulty hardware. Please contact your hardware vendor for further assistance diagnosing the problem.
Pyroman Posted May 18, 2021 Author Posted May 18, 2021 Of course it's an update that can't be uninstalled! @Leemanator Are your servers up to date on the hardware side, BIOS etc? I'm just on with DELL at the moment and they're checking it's nothing hardware related
mavhc Posted May 18, 2021 Posted May 18, 2021 Hardware raid is usually terrible. Especially when the battery fails. RAID firmware upgrade required?
Leemanator Posted May 18, 2021 Posted May 18, 2021 We're in the process of running the SUU on the affected servers now, will let you know how it goes for us. I've found rebooting the server seems to alleviate the issue for a short period so we've managed to avoid any end user issues by rebooting when needed. All affected servers have been updated up to the April cumulative update (kb5501347), but it doesn't seem to correlate with when the issues occurred, some servers had this installed for several weeks before any faults started happening, and one server seemed to be having this issue before this update was applied looking into event viewer, so sort of suggests it wasn't this update at fault... we've started deploying the may cumulative to see if this has a fix on it but not hopeful Also we've had this issue on a couple of server 2019 boxes too so doesn't seem to be 2016 specific. Hopefully the bios & firmware updates will work, we are a few versions behind by the looks.
Pyroman Posted May 18, 2021 Author Posted May 18, 2021 We're in the process of running the SUU on the affected servers now, will let you know how it goes for us. I've found rebooting the server seems to alleviate the issue for a short period so we've managed to avoid any end user issues by rebooting when needed. All affected servers have been updated up to the April cumulative update (kb5501347), but it doesn't seem to correlate with when the issues occurred, some servers had this installed for several weeks before any faults started happening, and one server seemed to be having this issue before this update was applied looking into event viewer, so sort of suggests it wasn't this update at fault... we've started deploying the may cumulative to see if this has a fix on it but not hopeful Also we've had this issue on a couple of server 2019 boxes too so doesn't seem to be 2016 specific. Hopefully the bios & firmware updates will work, we are a few versions behind by the looks. DELL support pointed me in the direction of this link: https://www.dell.com/support/kbdoc/en-uk/000178586/update-poweredge-servers-with-platform-specific-bootable-iso which I've just burnt to a disk and run on the server so I'll update when it's finished rebooting. I had the same as you, the April update has been on for a while and issues only appeared on Monday, I already have the May update applied as well so not sure that fixes the issue. My firmware was quite a way out of date on the RAID controller, BIOS, Network etc
Pyroman Posted May 18, 2021 Author Posted May 18, 2021 DELL support pointed me in the direction of this link: https://www.dell.com/support/kbdoc/en-uk/000178586/update-poweredge-servers-with-platform-specific-bootable-iso which I've just burnt to a disk and run on the server so I'll update when it's finished rebooting. I had the same as you, the April update has been on for a while and issues only appeared on Monday, I already have the May update applied as well so not sure that fixes the issue. My firmware was quite a way out of date on the RAID controller, BIOS, Network etc Updated and rebooted, still having issues with our printers not connecting. All papercut printers are showing "Can't connect to server, contact administrator"
Guest Guest Posted May 18, 2021 Posted May 18, 2021 Updated and rebooted, still having issues with our printers not connecting. All papercut printers are showing "Can't connect to server, contact administrator"Are all the PaperCut services running? We had that issue after patch Tuesday, manually started the service and all was fine
Pyroman Posted May 19, 2021 Author Posted May 19, 2021 Are all the PaperCut services running? We had that issue after patch Tuesday, manually started the service and all was fine As far as I can see, yesah they're running, I've just restarted the services and worked on this until about 2am last night so will let you know when I get in this morning!
computer_expert Posted May 19, 2021 Posted May 19, 2021 Just throwing this out there, but do you have ESET installed on the server? If so, is it endpoint protection or file server security (FSS)? There was a thread on here a while ago and they were having freezing issues due to using endpoint protection on the server rather than file server security. Installing FSS resolved the issue.
Leemanator Posted May 19, 2021 Posted May 19, 2021 Updated and rebooted, still having issues with our printers not connecting. All papercut printers are showing "Can't connect to server, contact administrator" Yeah we've updated the bios/raid controller firmware to latest and still seeing issues, I'm starting to bang my head against the wall as well. All our affected servers are using the H330 Perc adapter, I suspect it's due to this controller having no cache. Are you using raid 5/6? One of the servers that was affected worst by it yesterday morning has gone back to normality after a reboot, and been fine now for the best part yesterday and this morning, whereas a few others have gotten worse and reboots did not help. What A/V you running? we've got Sophos Central just wondering if that correlates with you? disabling the real time scan didn't seem to resolve any issues, but it's the only third party software that's installed on all affected servers.
Pyroman Posted May 19, 2021 Author Posted May 19, 2021 (edited) @Leemanator We've got the H300 Perc Drives are just mirrored Toshiba 4TB drives So it turns out Papercut was a total red herring. I haven't previously had iDRAC installed on the server and installed it on Monday. This caused papercut to think that the virtual adapter it created was the main NIC so it pushed out a configuration to the printers to phone home to a 169.254.0.2 address. Which is the iDRAC virtual card. I rang our Papercut support people yesterday and they said it was the server slowdown causing the issue and not papercut and left it at that. After another call with them today they found the IP issue and forced paperxut to use an alternate IP (what it should have been) and that config pushed out and fixed the printer issue. I'm having issues with windows updates failing and windows defender also throwing an error code and wonder if that's what's causing it. We have sophos installed as well on the server so that's a correlation between us. Edited May 19, 2021 by Pyroman
Leemanator Posted May 19, 2021 Posted May 19, 2021 @Leemanator We've got the H300 Perc Drives are just mirrored Toshiba 4TB drives So it turns out Papercut was a total red herring. I haven't previously had iDRAC installed on the server and installed it on Monday. This caused papercut to think that the virtual adapter it created was the main NIC so it pushed out a configuration to the printers to phone home to a 169.254.0.2 address. Which is the iDRAC virtual card. I rang our Papercut support people yesterday and they said it was the server slowdown causing the issue and not papercut and left it at that. After another call with them today they found the IP issue and forced paperxut to use an alternate IP (what it should have been) and that config pushed out and fixed the printer issue. I'm having issues with windows updates failing and windows defender also throwing an error code and wonder if that's what's causing it. We have sophos installed as well on the server so that's a correlation between us. Had stopped sophos on one server for a bit, didn't seem to make a difference. Do you use either splashtop or datto by any chance? I've just noticed that the date of install of the latest update correlates exactly the day when the servers started experiencing issues. Might just be a coincidence but I've removed it from one server and following a reboot it's ran fine for the last half hour.
DGardiner Posted May 19, 2021 Posted May 19, 2021 Updated and rebooted, still having issues with our printers not connecting. All papercut printers are showing "Can't connect to server, contact administrator" Papercut can be a bit funny about the idrac if its connected as a network card - Gotta set the bind address manually iirc or disable the idrac>os interface
Pyroman Posted May 19, 2021 Author Posted May 19, 2021 Had stopped sophos on one server for a bit, didn't seem to make a difference. Do you use either splashtop or datto by any chance? I've just noticed that the date of install of the latest update correlates exactly the day when the servers started experiencing issues. Might just be a coincidence but I've removed it from one server and following a reboot it's ran fine for the last half hour. Unfortunately not, I've got teamviewer on there and I noticed above reader got installed via an update on Friday so I've removed that. Managed to fix the Windows updates issue (I think) by using the Windows update troubleshooter after every other way of fixing it didn't work
Pyroman Posted May 20, 2021 Author Posted May 20, 2021 Papercut can be a bit funny about the idrac if its connected as a network card - Gotta set the bind address manually iirc or disable the idrac>os interface Yeah that was a real spanner in the works, broke one thing trying to fix another thing. Spent hours trying to work that one out before we realised that the 169 address wasn't in fact a 'can't find the network' address.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now