TheRobins Posted June 17, 2014 Posted June 17, 2014 There was a previous post pointing to HDD firmware, HDD is seeming more likely to be the cause. I wonder what level of RAID it is running. I had a similar issue years ago on a Raid 5 system, and remove each member drive one at a time until the server became stable, of course that pointed to the drive being the culprit. How old is the server?
witch Posted June 17, 2014 Author Posted June 17, 2014 Server is an ML350 running Server 2008R2. It is about 4 years old now. It has a Raid 5 system - all drives have checked OK with every sort of diagnostics I could find. I will consider removing a drive at a time if the problem persists.
abillybob Posted June 17, 2014 Posted June 17, 2014 Server is an ML350 running Server 2008R2. It is about 4 years old now. It has a Raid 5 system - all drives have checked OK with every sort of diagnostics I could find. I will consider removing a drive at a time if the problem persists. I'm sorry for just jumping in here and I have to confess I don't want to sit reading through 5 pages of posts bout we had a similar issue. I did fix it though can't remember the exact steps it was something to do with disabling SMB2 on the file server and it fixed it straight up! Maybe you could have a similar issue? This may help Word 2010 Slow opening files on network share
SHimmer45 Posted June 17, 2014 Posted June 17, 2014 i didnt think disabling SMB versions was recommended by microsoft as the client and the server auto negotiate what version to use
abillybob Posted June 17, 2014 Posted June 17, 2014 i didnt think disabling SMB versions was recommended by microsoft as the client and the server auto negotiate what version to use Worked perfectly fine here never had an issue since!?
Duke5A Posted June 17, 2014 Posted June 17, 2014 Server is an ML350 running Server 2008R2. It is about 4 years old now. It has a Raid 5 system - all drives have checked OK with every sort of diagnostics I could find. I will consider removing a drive at a time if the problem persists. The drives with the bugged firmware will pass tests right up until they fail outright. Reading your latest posts I'm pretty confident this is your issue. Run the ACU, generate a diagnostics report, and attach it to this thread. The report will contain the drive models and firmware revision they're on. This will tell me for sure and I can help you sort out of the mess that is the HP website. If they're the same model drives I may even have the firmware laying around. I think HP went to a model where they will only let you download firmware updates if you have a valid support contract now.
randidiot Posted June 17, 2014 Posted June 17, 2014 Stick to Bitdefender or kaspersky for server security, uninstall every program from the server that isn't a role and run the server COMPLETELY blank, disable internal security, and slowly bring features online to you narrow the issue down, replace cable from server and check that your swtiches are running STP, a user may be accidentally pluging a Eth back onto itself, causing the switch to flood, next time the server does it, logon and do \\servername\share if your able to access the share via the server name then the share is working, DO NOT USE \\LOCALHOST this bypasses alot of things, if your able to access the share locally then the server is working CORRECTY, and you must track back from the ethernet controller to the swtiches.
pete Posted June 18, 2014 Posted June 18, 2014 There's also been a recent advisory from HP regarding SATA disks with online spares connected to certain Smart Array controllers: HP Support document - HP Support Center Any HP ProLiant server listed under Platforms Affected configured with any of the following HP Smart Array Controllers running Firmware Version 6.40 (or earlier): HP Smart Array P212 Controller HP Smart Array P410 Controller HP Smart Array P410i Controller HP Smart Array P411 Controller HP Smart Array P711m Controller HP Smart Array P712m Controller HP Smart Array P812 Controller Note: This issue occurs only on HP Smart Array Controllers with one or more HP SATA HDDs or SSDs configured as Online Spare Drives. Smart Array Controllers configured with SAS HDDs or SSDs are NOT affected.
witch Posted June 18, 2014 Author Posted June 18, 2014 Stick to Bitdefender or kaspersky for server security, uninstall every program from the server that isn't a role and run the server COMPLETELY blank, disable internal security, and slowly bring features online to you narrow the issue down, replace cable from server and check that your swtiches are running STP, a user may be accidentally pluging a Eth back onto itself, causing the switch to flood, next time the server does it, logon and do \\servername\share if your able to access the share via the server name then the share is working, DO NOT USE \\LOCALHOST this bypasses alot of things, if your able to access the share locally then the server is working CORRECTY, and you must track back from the ethernet controller to the swtiches. Thanks for that. Unfortunately I am only here during the day when the server is in use so it is all a bit difficult. The issue is very random and so far I have only been here for one event - but I will certainly try the \\servername\share thing if I see it again. New issue today - I have been called up by two members of staff - so far - as their internet doesn't work. It turned out their proxy had disappeared. I reset the profile, deleted local profiles and in one case got them to log on to another machine with no luck. I can put the proxy in but it obviously isn't being picked up by the machines. (one wireless laptop, one wired PC) Have checked for loopbacks
witch Posted June 18, 2014 Author Posted June 18, 2014 ACU Diagnostics Report.zip Here is the diagnostic report
Duke5A Posted June 18, 2014 Posted June 18, 2014 (edited) [ATTACH]25137[/ATTACH] Here is the diagnostic report Alright, I went through it and here is the load out: Controller: P410i 5.12 Drive 1I:1:1: MM0500FBFVQ 500 GB HPD1 - Original Drive 1I:1:2: MM0500FBFVQ 500 GB HPD1 - Original Drive 1I:1:3: MM0500FBFVQ 500 GB HPD8 - Original Drive 1I:1:4: MM0500FAMYT 500 GB HPD5 - Original Drive 1I:1:5: MM0500FBFVQ 500 GB HPD1 - Original The drives are a bit out of date and the controller is really out of date; according to ACU the drives are all wearing the same firmware they came loaded with from the factory. The current firmware links are below. Controller: P200i v9.30 5 May 2011 http://h20566.www2.hp.com/portal/site/hpsc/template.PAGE/public/psi/swdHome/?sp4ts.oid=3902575&spf_p.tpst=swdMain&spf_p.prp_swdMain=wsrp-navigationalState%3DswEnvOID%253D4064%257CswLang%253D%257Caction%253DlistDriver&javax.portlet.begCacheTok=com.vignette.cachetoken&javax.portlet.endCacheTok=com.vignette.cachetoken MM0500FBFVQ Drive: HPD8 © 18 Feb 2014 Drivers, Software & Firmware for HP SAS Hard Drives - HP Support Center MM0500FAMYT Drive: HPD6 © 18 Feb 2014 Drivers, Software & Firmware for HP SAS Hard Drives - HP Support Center The controller firmware needs a valid support contract in order to be downloaded. If you don't have one then this can still be found on Google or I might even be able to download it - I think I may have an HP box with that controller still covered under a Care Pack. BTW - the firmware is designed to be loaded using the HP Firmware Maintenance CD. You can download any version and place the updated firmware onto it. Let me know if you need any help. Edited June 18, 2014 by Duke5A
witch Posted June 18, 2014 Author Posted June 18, 2014 Thanks I dont have a firmware maintenance CD that I can find, and I do not have a support contract any more either, sadly. But the diagnostic showed no errors that you could see?
DMcCoy Posted June 18, 2014 Posted June 18, 2014 Have you tried a chkdsk /f on the volumes? Windows can fall over sometimes with ntfs corruption but not report it, although it's quite rare, only seen it once or twice. FYI also running a few file servers with eset without issue, although uninstalling AV is always a good start when diagnosing server lockups.
witch Posted June 18, 2014 Author Posted June 18, 2014 I haven't.But I couldn't do that when the server was working, can I?
DMcCoy Posted June 18, 2014 Posted June 18, 2014 No, you will need to reboot for C: and unmount the other drives, it can take quite a long time.
witch Posted June 19, 2014 Author Posted June 19, 2014 ..and therein lies the problem. I am going to have to arrange to be here after hours, I can see
Duke5A Posted June 24, 2014 Posted June 24, 2014 Thanks I dont have a firmware maintenance CD that I can find, and I do not have a support contract any more either, sadly. But the diagnostic showed no errors that you could see? Sorry, I've been away. The update CD is an ISO you download from HP. The individual firmware downloads I linked to are added to the CD and the utility will automatically load them. You do have to take the server down for about a half hour though.
Driftingashore Posted December 9, 2014 Posted December 9, 2014 Digging up an old post, but I stumbled across this after having a similar sounding issue so after getting to the bottom of it I thought I'd share what our problem was in case it helps anybody else in the future. Basically, the server's RAID card had caching disabled, resulting in poor write speeds. Over the years, our network got faster and files grew bigger, resulting in clients being able to saturate the disk under heavy usage and cause anybody reading/writing to a share to freeze. Drove me crazy trying to troubleshoot an intermittent problem leaving no logs anywhere, and no recent changes pointing to an answer. Of course, everything was working - it was just working too slowly to keep up until the load eased (however long that would take) - and then it would be back to normal. Enabled the disk cache, and boom, life is worth living again. Phew.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now