Jump to content

Recommended Posts

Posted

Hello,

 

I don't contribute nearly as much as I should and selfishly I am here to ask for HEEEEEELP! I really hope someone can help me get out of this horrible mess.

 

So, our file server is Server 2008 R2 running on vSphere 5.5. Since the kids and staff came back, we are now suffering from intermittent performance lag to the file server data drive, usually followed by the D drive completely dropping out and then after a minute or so, it comes back and resumes normal service for ten mins or so before it happens again. It doesn't happen on any of our other VMs so it doesn't make sense for it to be the virtual environment. We have three ESX hosts connecting to a SAS (all Dell).

 

I tried moved the file server to a different ESX host and the problems remained. The only change to that environment was that we did a firmware upgrade to the SAS over the holiday. Today, we have also firmware upgraded the disks within the unit, but that's made no difference. But I don't see why any of that should affect it. The status of all of that hardware is optimal.

 

Right now I am wondering if it's just a more traditional low-level OS issue, but my brain is fried and if anyone could give me a fresh perspective, I would be eternally grateful.

 

The best way to describe the problem is when RDCing on to the file server, you can see that the D drive icon changes for the standard hard drive icon and the blue used space bar disappears, then the progress bar on the Explorer window crawls along for a minute or so, then the hard drive icon returns to normal and you get normal service again.

 

We have two WMDK files on the virtual environment which are spanned on the file server to make one logical D drive of about 2.8TB.

 

I promise I will contribute more to the forums from now! Please help!

 

Mike

Posted

start with the basics of troubleshooting and then increase the troubleshooting difficulty:-

what does event viewer say on the 2008 box when the drives "drop" ?

how many luns on the SAN? any event errors?

is the iscsi traffic VLAN'd off from the LAN on vmware?

is the iscsi traffic using TCP4or6?

if 4 is jumbo frames enabled?

is there anything else on ISCSI vlan that could effect?

is there any other protocol on the nic that could effect?

how much free space is there?

do you have enough free space to recreate the drive?

do you have a testbed you can recreate the problem ?

can you rollback the firmware to prove this is the problem?

 

i could go on forever but get my drift.

Posted

Hi Andyis,

 

No, please do go on forever - any ideas are most appreciated!

 

Half of the problem is that I am not very experienced in this particular area, but I'll get there.

 

The other problem is that the event viewer isn't giving much away. We are however getting errors under event ID 7011:

 

A timeout (30000 milliseconds) was reached while waiting for a transaction response from the VMTools service.

 

We have reinstalled VM Tools but this hasn't helped. The tools we had running have been happy for around three years without issue and I can't think what may have changed.

 

We have two LUNS and as far as I can tell in vSphere, there aren't any errors.

 

I believe we are VLANd off but I'd have to double-check - quite VLAN heavy here and I know there is a VLAN I belive called vMotion

 

I believe we're only using TCPIP4, haven't touched 6 at all - hmmm, I will Google where to find jumbo frames

 

D drive has 172GB free space, which is more than we had a couple of days ago as we're trying to reduce it in case that helps.

 

Unfortunately, we do not have enough free space to recreate the drive - we're that drastic that I'm actually wondering about getting a new server in to set it up as a workaround. The SAS itself is mostly all allocated (I should point out that that hasn't changed for quite a while, and we do not have any snapshots).

 

Maybe we should try the firmware rollback...

 

Thank you for your help - any other ideas most welcome!

Posted
When you say SAS, do you mean a locally attached SAS array? Or do you mean a SAN, in a dedicated storage controller? This is shouting dying disks or RAID controller to me, as that's the symptoms I had with one of my hosts. Eventually the RAID card died - replaced it like for like and fine since.
Posted

Just reading about jumbo frames here, which I admit I didn't know much about before today...

 

What is jumbo frames? - Definition from WhatIs.com

 

Hmmm, I suppose packet loss could occur over iSCSI between the VM/ESX and SAS if there is an incompatibility somewhere re jumbo frames? But once again, nothing had changed in this regard, although when we were reinstalling VM Tools and separately following a server reboot, the IP 4 settings dropped out of the NIC and I had to console into it to put them back in. Possibly related? Not sure why the settings dropped.

 

I looked in the Event log for the SAS and just found a few of these errors:

 

Date/Time: 31/08/17 15:50:34Sequence number: 11242Event type: 1710Event priority: CriticalDescription: RAID Controller Module wide port has gone to failed stateEvent specific codes: 0/0/0Event category: ErrorComponent type: RAID Controller ModuleComponent location: Enclosure 0, Slot 1Logged by: RAID Controller Module in slot 1Raw data:4d 45 4c 48 03 00 00 00 ea 2b 00 00 00 00 00 00 10 17 18 01 3a 22 a8 59 14 00 00 00 00 01 01 00 00 00 00 00 01 00 00 00 22 00 00 00 22 00 00 00 08 00 00 00 00 00 00 00 02 00 00 00 01 00 00 00 0a 00 00 00 01 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 01 00 00 00 00 00 01 02 10 00 00 00 04 00 00 08 00 00 00 00 04 00 26 08 02 00 00 00

 

But at this precise moment it says that the storage array status is "Optimal"

Posted

Date/Time: 31/08/17 15:50:34Sequence number: 11242Event type: 1710Event priority: CriticalDescription: RAID Controller Module wide port has gone to failed stateEvent specific codes: 0/0/0Event category: ErrorComponent type: RAID Controller ModuleComponent location: Enclosure 0, Slot 1Logged by: RAID Controller Module in slot 1Raw data:4d 45 4c 48 03 00 00 00 ea 2b 00 00 00 00 00 00 10 17 18 01 3a 22 a8 59 14 00 00 00 00 01 01 00 00 00 00 00 01 00 00 00 22 00 00 00 22 00 00 00 08 00 00 00 00 00 00 00 02 00 00 00 01 00 00 00 0a 00 00 00 01 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 01 00 00 00 00 00 01 02 10 00 00 00 04 00 00 08 00 00 00 00 04 00 26 08 02 00 00 00

 

Are there any more of these logs that co-incide with your drive dissappearing?

 

if so thats your problem.

Posted

Hi 3s-gtech,

 

Thank you for your thoughts! Although I currently have an optimal state, do you think the error in the event log I found above could back up your theory?

 

Yes, we have a Dell MD3200 SAS unit attached to the three ESX hosts.[TABLE=width: 248]

[TR]

[TD=width: 248]Dell MD3200[/TD]

[/TR]

[/TABLE]

Posted

Thank you all very much for your help here. I have found a design diagram showing how our three ESX hosts are connected with SAS cables to SAS box.

sas.jpg

 

The RAID wide port errors don't appear to have occurred in the last week, but it did happen about 6 times during August.

 

I'm not sure if we use MPIO but do you think I should deploy the hotfix just inc case?

Posted
Hmmm thinking about it, we last had the RAID error literally 10 minutes before we firmware upgraded the SAS. Googling suggests that the firmware update can solve the error, which makes me think that that may be a red herring. Hmmmm.... I kind of want something to catastrophically fail as at least then we'd know what the problem is!
Posted

In the event log we seem to have consistently had the following error throughout the day, which seems to correspond with the drive "dropping out":

 

event ID 7011:A timeout (30000 milliseconds) was reached while waiting for a transaction response from the VMTools service.

 

I actually logged a call with VMWare this morning but we're on an 8-hour response as we have basic support, so not sure how long I'll be waiting.

 

Not where I want to be on a Friday...

 

Posted
Oh okay, an external SAS array. What was the last RAID error? Any predicted failures etc? As another example, I have another server with a pair of SSDs (in RAID 1) in it. No errors on the RAID controller UI, but any VMs put onto that array just don't work. The SSDs have died in some way, but no logs about it at all. Any way of narrowing down which drives may be at fault (how many do you have total)?
Posted
Might be a stupid suggestion, but considering you have redundant controllers, have you tried leaving only one connected, then do the same for the other, to see if its an issue with just one of the controllers?
Posted

Hi all,

 

Thank you for your suggestions! Yes, maybe having a go with the redundant controllers will be worth a look.

 

Yes - we are currently under warranty with Dell. We have spoken to them for advice about the firmware upgrades that we have done. I'm not sure how helpful they will be considering we don't have any other alerts, but I am going to crack on to them on Monday I think and see what they can do.

 

It would be interesting to try and narrow it down to a particular drive. My head hurts right now but I will think about how we might do that. Hopefully something will fail properly over the weekend.

 

Right now I'm considering spending some money on a new physical server and migrating the file server to it as a workaround for now.

 

Interesting how most of the kids and some staff have gone home, and the server is running more smoothly now. I poked it with a stick by running WinDirStat on it (just to create some intense activity) and then the drop-out happened again. Increased IO definitely seems to affect it.

 

Thank you all very much for your help on this! I will keep you posted.

Posted
I've just unplugged one of the redundant link between the ESX running the file server at the moment and the SAS - currently everything seems to be running OK. I suppose we'd either see more regular and more catastrophic drop outs, or better performance. Let's see how it goes! 2 minutes to 4pm - not sure what else I can do today!
Posted

if it was me i would also

 

1) write a quick batch file to alert/email me if exist d:\folder was not found and loop it.

1a) run constant ping to the iscsi SAN (to see if it drops / increases )

2) if there is any space left on a LUN , attached a new Disk (to see if this drops out too when the first disk drops)

3) get the bottom of the SSD disks issues you also have.

Posted
The only change to that environment was that we did a firmware upgrade to the SAS over the holiday.

 

Do the RAID controllers have thier own dedicated battery backup packs? If not, they will default to write-back caching - write-through caching could have been set before the firmware upgrade, but doing the upgrade might have set them back to their defaults. If you have a UPS powering your servers you should be covered for backup power anyway.

Posted

Morning all!

 

I hope you had good weekends. One good job I did was to resume backups of the filer server on DPM, which will help my next plan... to buy in a new dedicated server and migrate the file server to it. I will just recover from DPM as a form of data migration as it should use the same equivalent locations and NTFS permissions on the new server. It will in turn free up some much needed space on our SAS as we are chocked for space at the moment - can't take snapshots or fire up any new VMs, which is not good. While I'm working on that we will work on other options in the meantime.

 

Dhicks, interesting what you're saying about the caching mode. I will see if I can find that setting to see what it's on. That said, we have UPS and I believe there is a battery on the controllers. It got me wondering if the firmware upgrade could have done something like corrupt the VMDK file for the D drive? That said, we did shut down as many of the servers as we could prior to doing the firmware upgrade.

 

Re rolling back, I thought my colleague had created a rollback plan but sadly, he hadn't done that. So, we don't know what version we were on prior to the upgrade (although it would have been about three years out of date). A lesson learned there for future change management and getting the plan signed off before proceeding! I was so busy dealing with other things that I was just happy that he was dealing with it - I should have paid more attention. That said, we can't say for sure if the firmware upgrade was the cause.

 

So for now, full steam ahead on the new server. We use redirected App Data for users on the file server so I can't imagine that that's helping end user performance at the moment (perhaps why PCs start freezing when the D drive drops out on the file server) - I might see if I can re-localise those or move them to another server the minimise disruption.

 

Thanks for your help everyone and any brainwaves that you might still have are most appreciated. Sometimes it's difficult to think straight when you have people banging on the door, as I'm sure you can imagine.

 

Mike

Posted

I've just read this from the manual of the MD3200:

 

Write-Through Cache

The RAID controller automatically switches towrite-through if cache mirroring is disabled or if the battery is missing or has afault condition.

 

I'm not 100% sure how to work out our cache mode, but I can see that mirroring is enabled for both controllers. These are the settings I can find:

 

cacheFlushModifier=10;cacheWithoutBatteryEnabled=false;mirrorEnabled=true;readCacheEnabled=true;writeCacheEnabled=true;mediaScanEnabled=true;consistencyCheckEnabled=true;cacheReadPrefetch=true;modificationPriority=high;preReadRedundancyCheck=false;

 

Does writeCacheEnabled mean that we are on write-through? Anyhow, the manual indicates that the controller would default to write-through in the case of a missing battery of fault, which makes me think that it would manage a potential problem to minimise potential problems caused by write-back cache. So maybe that's not it in this case?

Posted
Anyhow, the manual indicates that the controller would default to write-through in the case of a missing battery of fault

 

Indeed - so with write-through caching, every disk write to the RAID controller would wait until it's actually been written to the disk(s), the value written is simply cached for the next read. It makes for write speeds that depend on the performance of your disks, which can be rather slow, instead of write speeds at the speed your RAID controller can cache them, which would be pretty much as fast as the controller's RAM.

Posted

Thanks for the clarification dhicks! I think I'll leave the cache alone unless anyone has any strong thoughts that it could be part of the cause. I do just wonder if we've just got a corrupt VMDK file for that server's D drive. We've had no other issues with any of our other 12 virtual servers.

 

The file server was as good as gold over the weekend but with everyone in today, it gets progressively worse as the day goes on. Just had to reboot it, which now causes a weird scenario of the default gateway setting going missing from the NIC. I have to completely strip all of the NIC IP4 settings and put everything back to get a connection again. Seems a bit buggy but I've never seen it before.

 

I have to say, I think I'm pretty good at troubleshooting most things but this is either a very strange problem, or my lack of experience in SAS boxes and VMWare is tripping me up!

 

New HP server on order which should be here Wednesday, then hopefully by COB Friday the file server will be good to go - I'm really hoping so as I'm supposed to be on leave next Monday and Tuesday! I'm already looking forward to that cold beer on Friday night...

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...