nick3young Posted April 20, 2015 Posted April 20, 2015 Hi all, We are currently in crisis mode with SIMS. It is sat on a virtual Windows Server 2008 R2 64bit server with 16GB of RAM assigned and a quad core 2.6Ghz CPUs. For the past few weeks, the SQL service for SIMS has been maxing out the processor running at 100%. We have obviously tried a reboot which made no effect. I also tried adding an additional quad core CPU but this made no effect either - SIMS bascially just carried on maxing out the original 4 cores (I suspect this is a limitation of SQL Server 2008 R2.). Has anyone ever come across this? Does anyone have any ideas to start troubleshooting it? The CPU usage does drop down to around 20% - 40% in the evening but during the day it is straight back to full capacity again. This is task manager showing 4 of the cores running at full: Also, interestingly in vCenter, the server is showing incredibly high latency. Until the end of March, everything looked normal but since then, we are regularly seeing spikes of well over 1000 milliseconds (see graph below). We don't remember doing anything around that time but it could be an auto-update perhaps...? But if this is the case then I'm sure other schools will have been effected by it. Thanks Nick
nick3young Posted April 20, 2015 Author Posted April 20, 2015 Haha..... and I assume you're not joking? I've never heard of reducing RAM to improve performance! What is the reason?
3s-gtech Posted April 20, 2015 Posted April 20, 2015 That won't be it. What else is on that host? Anything that may be hammering the disks? Do you have AV on the SIMS server, and is it set to not use on-access scanning?
zag Posted April 20, 2015 Posted April 20, 2015 (edited) Haha..... and I assume you're not joking? I've never heard of reducing RAM to improve performance! What is the reason? Not a joke at all SQL server eats as much ram as you give it. It is a common scenario of people providing too much ram on virtual machines with SQL server and getting high CPU usage. We started with with 16 as well, and reduced to 10gb and its been great ever since. Saying that, it shouldnt cause 100% cpu, sounds like something else to me. Try running a SQL query analyzer on it for a hour or so. You should see any slow queries show up. Edited April 20, 2015 by zag
nick3young Posted April 20, 2015 Author Posted April 20, 2015 Thanks for the replies both. Including the SIMS server, we have 4 virtual machines on the same host (an AD server, a user storage server and an application server e.g. running Clickview, Impero etc). The host has 16 cores and 64GB of RAM which is connected to a 4TB EMC SAN. I don't feel like we are overworking the host but I'm prepared to be proven wrong. I do have to admit, latency is ridiculously high across all the VMs on this host. Could a single troublesome VM have a knock on effect to the latency in the guest OSs on the same host? In other words, could one troublesome application be causing these latency issues across four servers? Could this be SIMS, ha?! I'm clueless because other than looking at SQL Server sat at 100%, I don't know what it is doing. Zag, your SQL query analyser sounds like a good idea. I have managed to arrange for a SIMS developer to remote in this afternoon so I might wait to see what he has to say before I start tinkering too much. I'm now just worried I could be chasing after the wrong cause though (with latency being high on all servers)....
zag Posted April 20, 2015 Posted April 20, 2015 Yes high IO will cause high CPU as well in a virtual environment. We have Sims on Hyper-V with local SSD storage and a dedicated server. Sims is very resource hungry over the last few releases.
kmount Posted April 20, 2015 Posted April 20, 2015 What spec/config is the SAN? How are the host's connected to it?
nick3young Posted April 20, 2015 Author Posted April 20, 2015 The SAN is an EMC VNXe3100 connected via iSCSI.
kmount Posted April 20, 2015 Posted April 20, 2015 (edited) What disks are in the DPE/DAS' on it and what RAID level? Most importantly, if you look in Unisphere what kind volume performance are you seeing? If you're not sure where to look let me know and I'll pull up one of our VNXe's to check. On the vCenter console can you see what kind I/O your VMs (combined & individually) are generating? PS. Is it a single or dual service processor VNXe? Are your datastores spread over both SPs if it's the later. Are you noticing performance issues on other VMs or only this one because of the 100% CPU? Edited April 20, 2015 by kmount
nick3young Posted April 20, 2015 Author Posted April 20, 2015 It has been a while since it has been set up and strangely, I can no longer connect directly into the SAN's web interface to pull up the information (it is timing out in the browser)... I wonder if this is related? But going from memory, it has 12 600GB SAS drives configures into 4 datastores. I believe each datastore is configured as RAID-5 - I could be wrong but the capacity of each datastore is what it would be configured as RAID-5. I can access a lot of information relating to disk I/O in vCenter... what would be most useful to look at? I can get 'highest latency', 'disk throughput usage', 'disk throughput contention', 'general usage', 'write rate', 'read rate' amongst lots of other counters.... Not sure if it is a single or dual service processor to be honest. Do you know how I could check? Yes, it's difficult to tell entirely, but we are noticing performance issues on other VMs. We are noticing that backups are taking longer to go through too. The graphs I showed in my initial post showing latency... this is also high for the other VMs on that host (some of those VMs are in different datastores though). Sorry, I feel that some of these answers have been a little 'wishy washy' but I'm right at the extreme of my knowledge with SANs here! Cheers.
kmount Posted April 20, 2015 Posted April 20, 2015 Hi there, No worries, I can understand it's frustrating with hassle going on. First thing's first, look at the back of the SAN and you'll know pretty quickly if it's dual or single controller, it'll either have two 'sides' each with their own power/network etc or it'll have just one side and big space where another one could be! I suspect those 600GB's are 15k, (they certainly are in our VNXe's) so performance on that for a small/medium workload should be OK though RAID5 may play in but if it worked fine before and had been fine until recently lets not start SAN bashing just yet. Not being able to get into Unisphere is a worry, it could be something as trivial as a memory leak or such that could be ebbing away at your performance, or it could be something completely different! (first step here is to get onto EMC Support and get them to diagnose/resolve that problem). Have a look at https://pubs.vmware.com/vsphere-51/index.jsp#com.vmware.vsphere.monitoring.doc/GUID-44252CB8-5561-488A-A8CE-CF05C8F584BA.html & https://pubs.vmware.com/vsphere-51/index.jsp#com.vmware.vsphere.monitoring.doc/GUID-92C91273-F466-4B51-89CC-C7064E6171CE.html -- Appreciate it's a pretty old link but should be useful to at least see if the SIMS box (vs the others) is generating loads of I/O (i.e. whether it is causing/contributing to the issue) or whether it's just a victim. Going back to your VM spec, you say you've got 18 physical processor cores, can you confirm what other vCPU allocations you have on the same host? (i.e. if the SIMS box is 8vCPU how many others are allocated and importantly, in what quantities) Might also be worth looking at the performance snapshots to look at the CPU RDY and CPU WAIT on the host + VMs to see whether the host or VM is having problems getting a slice of the pie but we can come back to that. Getting into the Unisphere would be good to see what kind of workload it is reporting.
nick3young Posted April 20, 2015 Author Posted April 20, 2015 Yep, it is dual controller then. I suspected it was but I didn't want to fill my response full of guess work! The drives are 15k - sorry I could have told you that bit. I will open a call with EMC to see if I can get them to troubleshoot the Unisphere connectivity issue then. I can't understand that one. I've always been able to connect into it fine - it is only today where I have first discovered that I can't. We have 16 physical cores and 16 virtually assigned across the 4 VMs on that host. We did have 20 assigned but I have recently reduced that to see if it was causing the issue. I had read that you can assign more vCPUs than you have physcially though... isn't the ratio as a general rule something like 4 to 1? I'll check the links out (thanks for those), open this call with EMC and report back. Cheers.
kmount Posted April 20, 2015 Posted April 20, 2015 You can indeed assign more than you have physical cores, that's part of the beauty of virtualization. - The caveat however is in understanding the below: If you assign a VM 8vCPU and it wants to do _anything_ it will have to wait until the host can find it 8 pCPU cores (or threads) upon which to execute the task. This is pretty bad if the VM just wants to say "hi!" Now, having one big VM and several little VMs is OK too because the scheduler on the host will be able to pop those 1-2 vCPU machines in and out to make space for your 8vCPU one but if you happen to have 16 CPU cores/threads to play with and you have 3 x 8vCPU VMs you'll find that only two of them can ever get service at a time with the third being effectively held still. It's not quite as black/white as that though as we're talking operations that occur in seconds. ESXi also likes a bit of spare core for itself to do the magic. - A long but useful outline is available - http://www.vmware.com/files/pdf/techpaper/VMware-vSphere-CPU-Sched-Perf.pdf So, you're 4:1 is fine, to be honest you'd likely get more but you 'up' your chances by being careful with vCPU allocation due to the scheduling, it's rare we find a VM that needs more than 4vCPU to be honest but of course it does happen. (we typically split roles out if we need more grunt than 4vCPU) I think your plan of action is solid, get the SAN sorted so we can see what it's saying about life. Something else I forgot to mention is to make sure you're not running any of your VMs in snapshots as that will increase your resource consumption (storage i/o side) as it may need to pick bits from layers.
nick3young Posted April 20, 2015 Author Posted April 20, 2015 Thanks for your ongoing help with this kmount. And a brilliant description of how VMWare handle's vCPUs to pCPUs - it makes perfect sense. It is almost like a queue system between the VMs for available pCPUs isn't it? I guess if you have 1:1 they never have to queue. It is interesting you mention snapshots because we use Veeam to perform our backups which works in part by taking a snapshot before removing it. One of our performance related issues is that VMWare\Veeam is taking an AGE to remove the snapshot from a couple of our VMs after a full backup. Last week, the snapshot removal ran to the end of Monday... the full backup started Friday night! It went through a bit quicker this weekend but I suspect this is all linked in with the same cause - whatever that is. I thought maybe that one of them could be stuck in limbo after your suggestion, but I've double checked and none of our VMs are running in snapshots. Nice thought though.
zag Posted April 20, 2015 Posted April 20, 2015 An easy way to prove it would be to move the server to another host with local storage. Just a thought...
nick3young Posted April 20, 2015 Author Posted April 20, 2015 I'm happy moving the server to another host Zag but I'm not sure how to do the 'local storage' bit. Do you literally mean, move the server on to a physical box and out of the virtualised environment? Kmount, I'm still waiting for EMC to troubleshoot the SAN, but in the meantime I installed VEEAM One reporting software because I'd heard some good reports about it. It appears to tell me a little more about my environment than vCenter does. One for example is this: It is reporting that two of our datastores have latency issues - the other two are fine. But I guess this doesn't answer the key question - is it a hardware issue on the SAN which is causing performance issues or is a troublesome application (e.g. SIMS) which is causing it. I've had SIMS in this afternoon and they believe it is an 'environmental issue'. Hopefully EMC can get to the bottom of it....
kmount Posted April 20, 2015 Posted April 20, 2015 Hi there, Sorry for the delay in replying. vCenter (at the request of Veeam) taking yonks to delete snapshots is very very worrying, that is very indicative of an underlying issue. One thing I was curious about, does it only take a long time to remove the snap from the SIMS VM or from other VMs too? (busy VMs can make snapshot removal take longer because the snapshot can be potentially rather large). The VNXe doesn't support VAAI either so storage operations need to run up and down the wire but I'm not convinced that is relevant right now. Those read latencies are pretty high, I'd be quite concerned about them.. Now, lets get cracking to figure out what's happening... I can see DS1 & 3 are showing latency where DS2 & 4 aren't. I'm going to go out on a limb here but I suspect we're going to find that DS1&3 are on one of the controllers(SP) and DS2&4 is on the other one. If I'm right in that thinking I suspect we're going to find that it could be as simple as a CPU issue on the SP but once we get into Unisphere we'll know more. --> A reboot of the SP hosting DS1&3 if I'm right that they are together with 2&4 on the other might be a sensible move; if the multipathing is set up properly you shouldn't feel the reboot. (we don't) Can we drill down to get the iops for each of the VMs? (or even datastores) -- interestingly right now you're seeing high read latency but reasonably OK write which is quite unusual because it's normal to have some read cache in place to try and 'offset' a chunk of that. I'm confident we'll get to the bottom of this though! Cheers, Kim
nick3young Posted April 20, 2015 Author Posted April 20, 2015 Ha... please don't apologise. I already feel guilty the amount of your time you have dedicated to this - you're a trooper. Are you a network manager for a school? The snapshot removal which took an age last week wasn't the SIMS server, but it is another server which is showing latency issues. Good thinking with the DS1 & 3 being on the same controller. I was trying to think what could connect the two and this seems like an obvious assumption. I guess it's typical that I currently cannot get into Unisphere to gather more info on this. EMC have replied though and asked me to SSH in.... which, like HTTP, didn't work. So they have now asked me to serial connect directly in to it. So that's the morning's challenge..... find the serial cable! I will also have a look to see if I can drill down further into the monitoring software. I'll report back! Thanks again.
teejay Posted April 20, 2015 Posted April 20, 2015 You need the latest re-index patch off Capita, usually solves this, did for us anyway and I now have it set as a regular maintenance task.
nick3young Posted April 20, 2015 Author Posted April 20, 2015 Did this patch drop your CPU usage down then Teejay? I'm beginning to think that the 'SIMS CPU' and 'disk latency' are two separate issues which are perhaps exaggerating each other. I'm surprised SIMS haven't mentioned it though because they have remoted in this afternoon.
teejay Posted April 20, 2015 Posted April 20, 2015 Yes, we were 100% cpu with latency problems, the re-index patch fixed it for us.
nick3young Posted April 20, 2015 Author Posted April 20, 2015 Interesting! Thanks Teejay... I'll have a look tomorrow.
kmount Posted April 21, 2015 Posted April 21, 2015 Hi Nick, No need to feel guilty chap, always happy to help the community. I was a network manager for a number of years, now I look after technical services for a services company prodominantly around this kind of technology. Teejay's shout about the SIMS reindex patch is good, there's always going to be cause and effect here, the latency may be an effect rather than a cause. Good luck with the serial cable, getting that hooked up will hopefully get Unisphere sorted out again and then you can confirm if it's the shared SP (lets hope it is, narrows down our search) and the reindex patch is definitely worth running.
zag Posted April 21, 2015 Posted April 21, 2015 I'm happy moving the server to another host Zag but I'm not sure how to do the 'local storage' bit. Do you literally mean, move the server on to a physical box and out of the virtualised environment? Yes its just a physical host with a large hard disk. This is basic virtulization and very easy to do. SANs are typically used in large data centers with hundreds or thousands of VM's.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now