Jump to content

Recommended Posts

Posted
Unrelated almost, so i apologise for that, but looking at that report you created with Veeam One, i thought it would be good for us to have. i installed it to play with but I don't understand what to put in the veeam one server popup box?
Posted

If it wasn't for the latency/slow snapshot issue, i wouldn't be looking at the setup or hardware, as if it worked before, and nothing has changed...

 

Just to throw some other ideas out there:

 

Snapshots can be a killer especially if you have kept multiple snapshots. It could aid the snapshot removal process, if you take and keep one, so that the subsequent snapshot has a newer starting point. I'm not expert in this area, so i may be wrong or have described it incorrectly, but i usually think of it as a bunch of diffs so after a snapshot, everything else that changes gets written to a file. Making a recent one could help, if it's to do with volume of data. If there are other storage issues, that could be something else.

 

I would look into if Antivirus has changed in anyway, worth looking at what processes are running on the SIMS server, i've had Symantec kill an SQL server, with it's background checking, even though it's not checking the files. The process was ccSvcHst.exe , even though it's not taking a large portion it affects performance a lot. Similarly, winrar will kill the server if i choose to use it when sql is under use.

 

If just SQL is going crazy - have there been any changes in SIMS? The recent update and / or migration to newer SQL version can cause issues.

Even something that you can't see like lots of users adding homepage widgets could suddenly up the load, and you'd never known it has happened.

Posted

No matter what anyone says, latency is always going to be significantly worse on a NAS/SAN via Ethernet (NFS or iSCSI) when compared to SAS JBOD or SAS Local Storage.

 

Have a look at this. The latency on my array of solid state drives is significantly worse than the latency of an array of local storage 10k SAS drives. (local storage means on the raid card of the server with the disks plugged directly into the VM host server)

 

latency.JPG

 

Going back to your system, you're latency are very very poor. It does indicate a problem with the SAN / iSCSI interface.

 

Can you drill down to the latency of individual VM's? They should all be roughly the same latency on the same interface, if they are, then that means the latency is within the interface rather than the individual VM.

  • Thanks 1
Posted

Thank you all for the wealth of replies here... lots of food for thought. I'll try and go through and answer any questions asked here (in the order they were asked!):

 

1. Zag, I thought so. I did actually setup a separate physical box with Server 2012 R2 and Hyper-V running a pool of 20 VMs locally. This is for our remote desktop services and it works a charm for tasks such as remote SIMS access (when SIMS is working!!). I'm guessing it wouldn't be easy to convert our SIMS server from VMWare and fling it across to this Hyper-V box though?!?

 

2. TomFreeman, I think you are perhaps installing the Veeam One client. You need to install the server first and this will become the address it is asking for when you install the client. We have the pro version so it might be slightly different to the free version that I assume you are trying?

 

3. Vikpaw, I get your point about hardware not changing... but perhaps it has. I though something could have perhaps failed in the SAN... I've heard that they can be quite temperamental! I've since learned that the SAN is 'ok' (I think/hope), but I'll come to that in a moment. Interesting you mention AV because we have upgraded it recently. But it was only a very small version upgrade of Symantec. I know I would be a fool to rule this out though so I'll see if I can do some digging. There hasn't been any obvious changes in SIMS though to my knowledge. We do run 'SIMS InTouch' which is painfully sluggish and seems to delay the home page loading.... but we have run that for about a year now. I am looking in to getting rid of it if I'm honest.

 

4. I get your point Abutters about latency.... but it shouldn't be 1500 odd milliseconds like you recognise. It's running over CAT6... not string, ha! Fibre is way more efficient I assume... but way more expensive too no? I can drill down to the latency on individual VMs, I'll have a look tomorrow to see if they all match up. It's a good idea, thanks.

 

5. Matt40k, yep, the ISCSI is on it's own dedicated network. There are CAT6 cables flying all over the place! 4 out of the back of the SAN... a pair of cables into a pair of different switches. There are also cables coming out of the SAN into the main network switch (I assume for management etc). It's a similar story for the pair of hosts that we have... but yes, it's all dedicated.

 

6. Kmount, EMC remoted in today and worked their magic. We can now get on to Unisphere again. We haven't drawn a lot of information from it... two things really. One being that the software is incredibly out of date (you'll probably not be surprised to learn!). So I need to get onto that... they said it takes about 3hrs - how long?! The 2nd thing is that there is nothing obviously faulty on the SAN. Do you have any further tips on where to look in Unisphere?

 

Thanks again all.

Nick

Posted

Evening,

 

Normally I budget for an hour/two on an upgrade, the counter has a fixation with 80 minutes whenever I kick one off. (We have a few 3300's and one or two 3100s left) - I'd definitely get the upgrade done at your convenience just to keep on top of it. If all set up properly it should be fairly painless (in the future, it might have a hissyfit now)

 

As to digging through Unisphere lets get a few bits to start with:

 

- Storage Tab, VMware Storage (if empty, look in iSCSI Storage), and you should see your datastores listed, click on "Details" and in the top right hand corner it will say which "storage server" it is running on. (note that down for each datastore)

- Settings Tab, iSCSI Server Settings, and you'll see a list of the iSCSI Servers configured along with the all important "Storage Processor" column - that combined with the above should be enough for us to deduce where each datastore is running

- System Tab, System Performance, lets get a look at those graphs please, it should help us see how busy each SP is

 

Cheers!

 

Kim

Posted

No Teejay... was too wrapped up in getting connected to Unisphere with EMC remoted in yesterday. I am definitely going to do it though and I'll post the results back here. I assume you have to do it when everyone is off? How long did it take to run?

 

Kmount, good news! I hadn't noticed or didn't even think to check last night... but after EMC rebooted the SAN management interface last night, our latency issues have gone! How weird? We can see from the graph, at 5pm when the rebooted... latency has dropped to well below 20ms from 1500ms. They didn't reboot the SAN... just the management element of it? So how could that have had such a drastic but positive effect?

 

SIMS CPU is still at 100% but like we sensed, this is separate issue which hopefully Teejay suggestion will resolve.

 

As for your suggestions, these are the results. Hopefully, it may shed some more light on the issue:

 

Datastore00: SPA

Datastore01: SPB

Datastore02: SPA

Datastore03: SPB

 

iSCSIServer_SPA: SPA (eth2, eth3)

iSCSIServer_SPB: SPB (eth2, eth3)

 

It doesn't look like the graphs have been capturing data until the reboot yesterday. Here are the graphs though for what they are worth....

 

Untitled.png

 

So you were right with your hunch.... 01 and 03 are on the same datastore.

 

Can I just say at this stage, thanks again everyone for all your help and suggestions. I feel hugely relived this morning now that the latency issues appear to be gone!

Posted

Yep! Thanks everyone.

 

And to add to the good news, re-index patch has dropped our CPU usage down from 100% to an average of around 25%. Success.... thanks TeeJay.

  • 1 month later...
Posted

I just had this problem again after the autumn upgrade and ran the 2 capita patches 14265 and 15589

 

Seemed to fix it!

 

I also setup the maintenance schedule as described earlier.

 

CPU has gone down from about 90% to 10%

Posted

Please see below a notice placed on MyAccount earlier today.

 

The SIMS Teacher app high CPU/connectivity issue has now been resolved

25/06/2015

What's the issue?

We can confirm that an issue experienced by a small number of schools related to the SIMS Service Manager components of the Teacher App, causing a high level of CPU usage has now been resolved.

What caused the issue?

Microsoft have now released an update to their Azure Service Relay Bus, which we have implemented for the SIMS Teacher App. The update has been automatically released for Teacher app schools.

In addition, Microsoft have published additional guidance that a further two ports should be opened if accessing the Teacher App from behind a firewall (ports 5671 and 5672). Supporting documentation for the Teacher App is also being updated with the latest guidance from Microsoft.

We apologise for any inconvenience caused by this issue and thank you for your patience.

  • Thanks 3

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...