Jump to content

Server 2012 File Server - suddenly stops serving requests (but otherwise looks fine)


Recommended Posts

Posted

So are some of you saying if you had 2012 or 2012 R2 with Windows 7 clients the problem isn't there, but when running 2012 or 2012 R2 with Windows 8 or 8.1 clients the problem's there?

 

Again this would point to the SMB versions for sure. And presumably if you're running 2008 R2, but using Windows 8 or 8.1 with RSAT the problem isn't there?

Posted (edited)

For us the problem only appears to occur when Windows 8.1 machines access our 2012 R2 file servers. We set group policy to stop our Windows 8.1 clients from loading roaming profiles and this seemed to fix the problem - 700 ish Windows 7 machines were still loading roaming profiles quite happily.

 

Until last night I could recreate the problem simply and reliably by getting a couple of Windows 8.1 machines to load a roaming profile from the server - as soon as they tried logging in the file server would stop responding to requests.

 

This morning I couldn't reproduce the problem at all, so we allowed 150 ish Windows 8.1 computer to load profiles. All was fine until 12:30 when we allowed another 60 or so Windows 8.1 computers to load profiles and within 5 mins the server had died again.

Edited by DJL
Posted
We only have Windows 7 clients and it only happens a couple of times a month. If restarting the server service gets it going again, then I can live with it until a proper cure is found.
Posted (edited)

Definitely looks like that's true for us. We're just trying to work out if there are any differences in updates etc between the 150 that seemed ok this morning and the 60 we enabled at 12:30

 

The other thing I've noticed is occasionally the powershell command Get-SmbConnection shows that Windows 8.1 is talking to 2012 R2 using SMB3 instead of SMB3.02 which is odd. I'm wondering if there is an issue with SMB dialect negotiation

Edited by DJL
Posted
So it looks like Windows 8 and Windows 8.1 clients are provisionally the problem :)

 

Be interesting to get feedback from others.

 

I'm running a Win7 client and I had as many problems as anyone at our place. We don't have Win8 at all. Sorry :-(

Posted
I'm running a Win7 client and I had as many problems as anyone at our place. We don't have Win8 at all. Sorry :-(

 

Windows Server 2012 R2? Are all your Windows 7 machines running all the latest updates from WSUS?

Posted
The one other thing to try is to have a 2012 or 2012 R2 server without File Services switched on, but could be a DC running AD, DNS, DHCP and Print server roles.
Posted (edited)

I've had this problem exactly once about 6 weeks ago, and assuming it was the exact same problem, it happened simultaneously with Windows 7 SP1 and Windows 8.1 clients, all running the latest updates for the time. My file server is running Server 2012 (not R2), and also had the most recent patches at the time.

 

No other roles are installed on the server, it just runs File And Storage Services. For what it's worth, I do have the Data Deduplication role service installed, not that it should make any difference given that local access to files was working fine.

Edited by AngryTechnician
Posted
The latest from our end (remember I'm not privy to all the technical details):-

 

1 - we start the first day of a new half term with plans to restart the server service every once in a while

2 - our external support has escalated this to MS and following some memory dumps MS are now paying attention

3 - somebody (perhaps MS) remotes into our server and "removes" some of the software - I would love to know exactly what this means too!

4 - since then there have been no further instances of the problem

5 - the latest suspect is VEEAM

 

So, we are pootling along and as a user I am not experiencing any of the issues we've had for the past few weeks.

 

I don't think this is the end of the matter, and if I can find out what "software was removed" I'll repost.

 

Further update:-

 

The "removed software" was, in fact, services that were stopped. Apparently we have made progress as we no longer have the loss of our file server happening several times a day but we are running without various services.

 

We are bringing the services back one at a time to see what effect each has.

 

More when I know what happens.

  • Thanks 1
Posted (edited)

ETA: After re-reading the whole thread, I suspect mine is a separate issue from what's being discussed here. I'll try posting my own question later. Leaving what I wrote before in case I'm wrong...

I'm not positive the problem I'm having is exactly the same as what's described here. Here's the reason I'm not sure this is related to what you guys are discussing: All of our clients don't fail at once. It's one client at a time, not all at once. And often as not, after a little waiting, the client WILL eventually reconnect to the file share. Not always, but usually. The problem is, once one person is having the problem, it often occurs on other client machines and the only sure way to fix it all at once is just to reboot the server. Details below:

 

Single Server:

 

VMWare ESXi 5.5 build 1474528 with Windows Server 2012 R2 (Domain Controller) running DHCP, DNS, AD, File and Printer sharing

 

Client computers: Mostly Windows XP, some Windows 7.

 

For us, the problem manifests as a stall on a client pc (Most often Windows 7, very rarely Windows XP) when we go to save files in Office 2007 to the mapped network drive that points to the server. It doesn't happen every time, and the length of time a computer has had the file open doesn't seem to affect it either. That makes me think it isn't an opportunistic locking problem. It doesn't appear to be limited to office files, but that's where we see it most frequently as that's the most frequent type of file in use. When we try to access the mapped network drive during the stalled save, it stalls opening the window for a while, but eventually opens in most cases. Often as not, you can then find a way to pull office out of its stall and save properly it properly. It creates a .tmp file in the directory that was being saved to sometimes as well, which isn't unusual for Office. During these stalls, a ping command to both the IP and the name of the server works fine. NSLookup works fine. Address resolution seems to work fine for everything on the network. Here's the weird thing: Eventviewer isn't showing any consistent errors across these instances besides the app hang errors and credentials being submitted and verified by the server. The server isn't showing any consistent errors either.

 

The events that are being logged on a fairly regular basis are 1001 (hang) and 4648 (log on, resolves successfully) I see a simultaneous event on the server, ID 4776 credential validation, and then 4624 logon successful. Once in a while, I also see a warning - ID2012 System Log warning indicating a network error during transmitting/receiving data.

 

I've tried swapping out our network devices (switches & gateways) and that didn't solve the issue.

 

I've tried disabled smb2/3 on the server and the client machines on account of reports that it could cause issues just like this in Office during saves to mapped network drives. No dice so far. This happens at least once a day, and it's driving me bonkers. The ONLY real fix that gives me a few hours of peace is to reboot the server. Thankfully, it reboots inside of 3 minutes, but it destroys workflow in the office on anyone using shared files. (That's everyone in the office.)

 

Now that I know what service to try restarting (I'd previously tried restarting a bunch of services related to SMB and file sharing, but hadn't tried restarting the server service.) I'll give that a shot next time it happens and see if I can avoid a reboot. I'm going on a month and a week of troubleshooting this, and I've pulled out just about all my hair. Very much looking forward to that list of services they disabled. I can't get funding from my company to get outside help or open a case with Microsoft. So you guys are my best, last hope. ;)

 

Interesting links worth browsing related to all the searches I've done trying to figure out what the hell is causing this, some of which come from this thread.

 

Whitepaper on Opportunistic Locking and possible file corruption problems:

Opportunistic Locking and Read Caching on Microsoft Windows Networks

 

SMB commands for enabling/disabling smb 1, 2/3 on various operating systems:

How to enable and disable SMBv1, SMBv2, and SMBv3 in Windows Vista, Windows Server 2008, Windows 7, Windows Server 2008 R2, Windows 8, and Windows Server 2012

 

Current hotfixes for 2012 & 2012R2 related to file sharing issues:

List of currently available hotfixes for the File Services technologies in Windows Server 2012 and in Windows Server 2012 R2

Edited by bergmbe
Posted

@bergmbe

 

Your description of the user experience is exactly what we are seeing here; hung connections, not all machines at the same time, .tmp files being left behind etc.

 

As I promised an update here it is - there is nothing more to report. We remain functional, but still have some services switched off, and since last week we have not had any repeats of the mass disruption seen prior to the half-term holiday.

Posted (edited)
I've had this problem twice since I migrated from 2003 to 2012 R2 on my file server, October 2013. It is a pure file server, no additional roles. Win7 clients and a single Win8.1 client (mine, local profile). One user reports that they are having trouble opening or saving a file (both instances was a different excel 2010 spreadsheet), then slowly more and more people lose connectivity. Both times I have tried to access shared drives and my Windows 8.1 machine has locked up and had to hard reboot. I haven't experienced the issue since before the half-term and I did a complete server estate SUU and windows update deployment. Edited by cogrady84
Posted

@cogrady84

For what it's worth, disabling SMB2/3 seems to have stopped me from having to reboot workstations. It still stalls out when they go to access a shared drive after having a stalled save, but it seems to eventually reconciles itself after anywhere from 20 seconds to ~3 minutes. It's not a solution, but it allowed me to keep my system up in most cases. This only applies to Windows XP and Windows 7 home/pro

 

Basically if someone has a problem saving, I do this and it seems to avoid a workstation reboot:

Let Excel/Word/whatever chug along and try try try in the background. It leaves the 'saving' prompt up with a cancel button, which I do not hit.

Attempt to open the shared drive.

Wait for the shared drive to properly display contents - usually 40 seconds to 2 minutes.

Once it properly displays contents, navigate back to the program trying to save, hit cancel, wait for it become responsive again, and the try saving again.

 

Obviously this is by no means a solution, but believe it or not it's the only way I've been able to avoid a reboot on a workstation. Which is (unfortunately) necessary when they have 15 documents open that need to be saved. Here's the kicker... I get through that entire process, and it can take up to 5 minutes... and the only events I see in the logs on either the server or workstation relate to a credential negotiation (Event 4776) that returns a successful result (Event 4624). Once in a while, there's a 2012 network error event on the server, and of course it always registers apphang events related to excel/word/etc and explorer if that froze up for too long trying to display contents of a network share.

Posted

@bergmbe How you describe your issues almost mirror ours to the tee. Doesn't affect all clients all the time, but when it does it typically manifests itself as a stall out when opening Explorer (usually on Computer, so stalls refreshing the list of mapped drives) or when attempting to open/save office documents.

 

As with the others though, name resolution, pings, RDP - everything else appears to work fine during the periods of this beligerent behaviour.

 

Our "problematic" server is also a DC, hosting RID and PDC FSMO roles. I performed an "nltest /SC_VERIFY" early last week against the box and it couldn't find it's own domain name - a touch worrying - but other DCs could, so I figured the server hadn't correctly promoted itself so moved the two FSMO roles off to other DCs in preperation for a demote/re-promote to see if that cured it.

 

Now I didn't get around to the demote/repromote but interestingly the DC that got those FSMO roles now fails the "nltest /SC_VERIFY" test with the can't find domain error, and the problematic server now passes it - and I haven't seen the problem in near enough a week.

 

I know that may not help as people with this issue seem to be running a mix of 2012/R2 in both File server with DC, without DC, Hyper-V modes - but thought I would add it to the pot as anecdotal in case it helps somebody figure out what's going on.

Posted
So can everyone log into their Domain controllers and run netstat -an and look for 389 connections (more than one per device). Does it look like a lot of connections? When the File share server goes done (and run before it does as a comparison) run netstat -an | Select-String -pattern ":389" from powershell. Before connected after not connected? I am chasing this issue down for our environment. If you see what I am seeing we might be onto something.
Posted (edited)

I'm only seeing a few entries at a time using :389, and no discernible difference in the before, during and after file share issues. What I am seeing, and this may be normal as I haven't used netstat often enough to know if this is normal, is seeing a TON of entries overall. Pages and pages of them, most for TCP 127.0.0.1:randomport# showing state: esablished, or UDP 0.0.0.0:randomport#, or UDP [::]:randomport# both of the latter two showing *.* as their state. Perhaps you could tell me what those mean? Because that's an awful long list.

 

PS: If you have time to explain, assume I'm a total admin noob.

Edited by bergmbe
Posted
So those are listening ports, that is good. I have verified that the finals your server is listening on 445 which is SMB. Domain controller should be listening on 389. It might be but I have more issues besides the file server server. The occurrences I'm still happening I have a case open with Microsoft. We're doing some packet captures and memory dumps I will keep you posted.
Posted
Any way to filter the results so it doesn't list listening ports, but only active ones? I did a netstat -? in powershell and looked up a few things, but didn't see an easy way to filter for that in the results.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...