Jump to content

Recommended Posts

Posted

Hiya,

 

I'm banging my head against the wall, our Main server keeps randomly freezing.

 

The first time it got noticed was when it took all the papercut printers out because they couldn't get a connection to the server.

 

So far, I've ruled out that it's not Veeam and I'm struggling to see any software that's pegging the CPU at 100% or using all the memory or anything obvious.

 

I tried to download a file on chrome while its had this issue and it freezes and when it come back the download has failed so it's obviously affecting the network connection as well.

 

Program windows randomly say Not Responding and I can't click anything on the taskbar and then 5 mins later it comes back to life and works for a short while.

 

Installs hang, uninstalls hang.

 

I'm going to take it down after work tonight and do a memory and HDD diagnosis to rule those out but it's been running along happily for a few years now and all of a sudden with no changes it's started being a massive pain.

 

I'm working by myself so it's difficult not having someone to bounce ideas off!

Posted
Cheers, it's a DELL T130, no warning lights on the front. I'm just trying to install openmanage now but obviously having issues because everything keeps freezing!!
Posted

Don't forget to check the IPMI.

 

Also in task manager you can check the Status column for Suspended, and right click and Analyse Wait Chain

Posted

So just to add more fun into the mix, we've just had a powercut!!!

 

It did give me the chance to run the Dell BIOS diagnostics though and no HDD or RAM issues, or any issues at all for that matter were reported.

 

Unfortunately I was half way through installing openmanage when it died so I'll try again!

Posted
The RAID on T130 is, by default, an S130 software RAID that gets particularly poor reviews. Just thought that was worth mentioning as could be part of the problem.
Posted (edited)

Are you getting anything in the event viewer like the below?

 

 

Log Name: Application

Source: ESENT

Date: 17/05/2021 12:50:38

Event ID: 533

Task Category: General

Level: Warning

Keywords: Classic

User: N/A

Computer: *******

Description:svchost (3364,T,0) SRUJet: A request to write to the file "C:\WINDOWS\system32\SRU\SRU.chk" at offset 0 (0x0000000000000000) for 4096 (0x00001000) bytes has not completed for 36 second(s). This problem is likely due to faulty hardware. Please contact your hardware vendor for further assistance diagnosing the problem.

 

 

We've had a handful of Dell Poweredge Servers start freezing up randomly in the past couple of weeks, with event's like these showing writes are taking a long time. No hardware errors in open manage or when running diagnostics. Suspect it's to do with the April cumulative but can't see a pattern between the servers affected and the servers unaffected as yet.

Edited by Leemanator
Posted
Are you getting anything in the event viewer like the below?

 

 

Log Name: Application

Source: ESENT

Date: 17/05/2021 12:50:38

Event ID: 533

Task Category: General

Level: Warning

Keywords: Classic

User: N/A

Computer: *******

Description:svchost (3364,T,0) SRUJet: A request to write to the file "C:\WINDOWS\system32\SRU\SRU.chk" at offset 0 (0x0000000000000000) for 4096 (0x00001000) bytes has not completed for 36 second(s). This problem is likely due to faulty hardware. Please contact your hardware vendor for further assistance diagnosing the problem.

 

 

We've had a handful of Dell Poweredge Servers start freezing up randomly in the past couple of weeks, with event's like these showing writes are taking a long time. No hardware errors in open manage or when running diagnostics. Suspect it's to do with the April cumulative but can't see a pattern between the servers affected and the servers unaffected as yet.

 

YES!!! Any idea which KB that is off the top of your head? I'll try uninstalling it.

 

svchost (1112) SoftwareUsageMetrics-Svc: A request to write to the file "C:\Windows\system32\LogFiles\Sum\Svctmp.log" at offset 0 (0x0000000000000000) for 4096 (0x00001000) bytes succeeded, but took an abnormally long time (19 seconds) to be serviced by the OS. This problem is likely due to faulty hardware. Please contact your hardware vendor for further assistance diagnosing the problem.

Posted

Of course it's an update that can't be uninstalled!

 

@Leemanator Are your servers up to date on the hardware side, BIOS etc? I'm just on with DELL at the moment and they're checking it's nothing hardware related

Posted

We're in the process of running the SUU on the affected servers now, will let you know how it goes for us. I've found rebooting the server seems to alleviate the issue for a short period so we've managed to avoid any end user issues by rebooting when needed.

 

All affected servers have been updated up to the April cumulative update (kb5501347), but it doesn't seem to correlate with when the issues occurred, some servers had this installed for several weeks before any faults started happening, and one server seemed to be having this issue before this update was applied looking into event viewer, so sort of suggests it wasn't this update at fault... we've started deploying the may cumulative to see if this has a fix on it but not hopeful

 

Also we've had this issue on a couple of server 2019 boxes too so doesn't seem to be 2016 specific. Hopefully the bios & firmware updates will work, we are a few versions behind by the looks.

Posted
We're in the process of running the SUU on the affected servers now, will let you know how it goes for us. I've found rebooting the server seems to alleviate the issue for a short period so we've managed to avoid any end user issues by rebooting when needed.

 

All affected servers have been updated up to the April cumulative update (kb5501347), but it doesn't seem to correlate with when the issues occurred, some servers had this installed for several weeks before any faults started happening, and one server seemed to be having this issue before this update was applied looking into event viewer, so sort of suggests it wasn't this update at fault... we've started deploying the may cumulative to see if this has a fix on it but not hopeful

 

Also we've had this issue on a couple of server 2019 boxes too so doesn't seem to be 2016 specific. Hopefully the bios & firmware updates will work, we are a few versions behind by the looks.

 

DELL support pointed me in the direction of this link: https://www.dell.com/support/kbdoc/en-uk/000178586/update-poweredge-servers-with-platform-specific-bootable-iso which I've just burnt to a disk and run on the server so I'll update when it's finished rebooting. I had the same as you, the April update has been on for a while and issues only appeared on Monday, I already have the May update applied as well so not sure that fixes the issue.

 

My firmware was quite a way out of date on the RAID controller, BIOS, Network etc

Posted
DELL support pointed me in the direction of this link: https://www.dell.com/support/kbdoc/en-uk/000178586/update-poweredge-servers-with-platform-specific-bootable-iso which I've just burnt to a disk and run on the server so I'll update when it's finished rebooting. I had the same as you, the April update has been on for a while and issues only appeared on Monday, I already have the May update applied as well so not sure that fixes the issue.

 

My firmware was quite a way out of date on the RAID controller, BIOS, Network etc

 

Updated and rebooted, still having issues with our printers not connecting. All papercut printers are showing "Can't connect to server, contact administrator"

Guest Guest
Posted
Updated and rebooted, still having issues with our printers not connecting. All papercut printers are showing "Can't connect to server, contact administrator"
Are all the PaperCut services running? We had that issue after patch Tuesday, manually started the service and all was fine
Posted
Are all the PaperCut services running? We had that issue after patch Tuesday, manually started the service and all was fine

 

As far as I can see, yesah they're running, I've just restarted the services and worked on this until about 2am last night so will let you know when I get in this morning!

Posted

Just throwing this out there, but do you have ESET installed on the server? If so, is it endpoint protection or file server security (FSS)?

 

There was a thread on here a while ago and they were having freezing issues due to using endpoint protection on the server rather than file server security. Installing FSS resolved the issue.

Posted
Updated and rebooted, still having issues with our printers not connecting. All papercut printers are showing "Can't connect to server, contact administrator"

 

Yeah we've updated the bios/raid controller firmware to latest and still seeing issues, I'm starting to bang my head against the wall as well.

 

All our affected servers are using the H330 Perc adapter, I suspect it's due to this controller having no cache. Are you using raid 5/6?

 

One of the servers that was affected worst by it yesterday morning has gone back to normality after a reboot, and been fine now for the best part yesterday and this morning, whereas a few others have gotten worse and reboots did not help.

 

What A/V you running? we've got Sophos Central just wondering if that correlates with you? disabling the real time scan didn't seem to resolve any issues, but it's the only third party software that's installed on all affected servers.

Posted (edited)

@Leemanator We've got the H300 Perc Drives are just mirrored Toshiba 4TB drives

 

So it turns out Papercut was a total red herring. I haven't previously had iDRAC installed on the server and installed it on Monday. This caused papercut to think that the virtual adapter it created was the main NIC so it pushed out a configuration to the printers to phone home to a 169.254.0.2 address. Which is the iDRAC virtual card.

 

I rang our Papercut support people yesterday and they said it was the server slowdown causing the issue and not papercut and left it at that. After another call with them today they found the IP issue and forced paperxut to use an alternate IP (what it should have been) and that config pushed out and fixed the printer issue.

 

I'm having issues with windows updates failing and windows defender also throwing an error code and wonder if that's what's causing it.

 

We have sophos installed as well on the server so that's a correlation between us.

Edited by Pyroman
Posted
@Leemanator We've got the H300 Perc Drives are just mirrored Toshiba 4TB drives

 

So it turns out Papercut was a total red herring. I haven't previously had iDRAC installed on the server and installed it on Monday. This caused papercut to think that the virtual adapter it created was the main NIC so it pushed out a configuration to the printers to phone home to a 169.254.0.2 address. Which is the iDRAC virtual card.

 

I rang our Papercut support people yesterday and they said it was the server slowdown causing the issue and not papercut and left it at that. After another call with them today they found the IP issue and forced paperxut to use an alternate IP (what it should have been) and that config pushed out and fixed the printer issue.

 

I'm having issues with windows updates failing and windows defender also throwing an error code and wonder if that's what's causing it.

 

We have sophos installed as well on the server so that's a correlation between us.

 

Had stopped sophos on one server for a bit, didn't seem to make a difference.

 

Do you use either splashtop or datto by any chance? I've just noticed that the date of install of the latest update correlates exactly the day when the servers started experiencing issues. Might just be a coincidence but I've removed it from one server and following a reboot it's ran fine for the last half hour.

Posted
Updated and rebooted, still having issues with our printers not connecting. All papercut printers are showing "Can't connect to server, contact administrator"

 

Papercut can be a bit funny about the idrac if its connected as a network card - Gotta set the bind address manually iirc or disable the idrac>os interface

Posted
Had stopped sophos on one server for a bit, didn't seem to make a difference.

 

Do you use either splashtop or datto by any chance? I've just noticed that the date of install of the latest update correlates exactly the day when the servers started experiencing issues. Might just be a coincidence but I've removed it from one server and following a reboot it's ran fine for the last half hour.

 

Unfortunately not, I've got teamviewer on there and I noticed above reader got installed via an update on Friday so I've removed that. Managed to fix the Windows updates issue (I think) by using the Windows update troubleshooter after every other way of fixing it didn't work

Posted
Papercut can be a bit funny about the idrac if its connected as a network card - Gotta set the bind address manually iirc or disable the idrac>os interface

 

Yeah that was a real spanner in the works, broke one thing trying to fix another thing. Spent hours trying to work that one out before we realised that the 169 address wasn't in fact a 'can't find the network' address.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...