Jump to content

Recommended Posts

Posted

Hi all,

 

I am hoping that someone will be able to help me as I'm slowly going crazy over this issue.

 

We have an issue whereby at the same time every hour XX:45 the network locks up for up to a minute. During this time, all of the machines hang and stop responding and then suddenly come back to life.

 

This was happening before the summer but maybe once a day at a random time but it seems to have become regular on an hourly basis since our return in September.

 

Over the summer all of the server hardware and most of the switching hardware was replaced.

 

We use observium to monitor all of our switches (HP Procurve) and the bandwidth graphs on observium look exactly the same when this happens, there is no peak at all.

 

Likewise rpvst is on and there is nothing in any of the logs to suggest ports are being blocked due to spanning tree. We also have loop-detect enabled and nothing on any of the switches to suggest this is an issue either.

 

I'm slowly running out of ideas. We use Sophos on all clients and servers and they all pull down updates at different times, so that isn't the issue either.

 

I'm thinking wireshark next might shed some light (although I'm not 100% sure how to use it), but I am slowly running ideas.

 

Any suggestions would be greatly appreciated!

 

Kind regards, Ed

Posted
Windows updates? Any scheduled tasks? Any scripts? Anything set on impero?Not got spice works or something similar doing a network scan? Av doing a scan ? Have to tried a process of elimination? Turning all machines in school down and seeing if it happens in one room at a time?
Posted

The servers are fully patched on Server 2012 R2 and are running across three VMware ESXi 6.0 Hosts which are also fully patched.

 

The workstations are almost fully patched Windows 7 Enterprise SP1 (since 1st August 2016 when the image were made) and although we run WSUS no new updates have been approved for installation yet.

 

We do not use spiceworks and I have checked that it is definitely not installed on either of our servers that we used to run it from.

 

In terms of Class Management software we use NetSupport and are on the latest version across the site.

 

Each IT room is on its own subnet as well as servers, printers, cashless catering, projectors, wireless, etc all also having their own individual subnets and the freeze can be seen across all devices.

 

We have Sophos antivirus which is set to update at 60 minute intervals but looking at event logs across a range of devices, this seems to happen at different times across the various devices.

Posted
Each IT room is on its own subnet as well as servers, printers, cashless catering, projectors, wireless, etc all also having their own individual subnets and the freeze can be seen across all devices.

 

Probably way too many VLANs then by the sounds of things. Why a different subnet per room?

Posted
Probably way too many VLANs then by the sounds of things. Why a different subnet per room?

 

We only have 6 IT rooms and the rest of the desktops (approx. 600) are in one large subnet. The IT rooms were put onto their own subnets for ease of management more than anything.

 

This has been configured in this way for almost 2 years so if the issue were due to too many VLANs surely the issue would have always been present?

Posted

It's not an issue just becomes overkill I think @Michael means.

If this is seen across all devices then its either your core or server infrastructure.

Are any backups configured to run at the time?

What are the exact symptoms ? Does the PC completely freeze of just lose connection to network resources.

Posted
Is there anything in event viewer on the PCs? If you remove network connection prior to the time of freeze does it still do it?
Posted
Failing that you will need to run a packet sniffer like wireshark as you mentioned. Just turn on the monitor and let Google be your best friend when analysing results
Posted

Veeam is configured to backup out of hours and I did disabled all of the jobs last night to see if that helped today which it did not.

 

Our core is a new Hewlett Packard Enterprise \ Aruba 5412Rzl2 which is not showing a lot in a "show log" despite the logging being set to high. Mainly log in and log out from me.

 

The virtual server infrastructure seems happy enough - Looking at the network monitoring for each host via vcentre they don't seem busy at all - as per the attached screenshot.

 

The PC completely locks up - so during that time we are unable to launch command prompt to try pinging the core \ any servers \ etc.

 

NetworkMonitor.png

Posted
Is there anything in event viewer on the PCs? If you remove network connection prior to the time of freeze does it still do it?

 

Being logged in locally the issue does not occur. There is nothing on the local machine in event log to suggest an issue - in fact there is nothing in the event log approx. 10 mins either side of the issue happening.

Posted
Based on being logged in locally not having an issue I would look at file servers... Anything being redirected like the desktop? We had an issue on a few admin machines connected to an old network drive on a really old server which had a raid problem and it cause their explorer.exe to freeze up.
Posted
I'm going for server too . Gp result a computer and ensure everything adds up. E.g network drives . Or turn some policies off one by one to narrow it down
  • 2 weeks later...
Posted
Just a thought, are you running any flavour of 2012 Sever? It sounds familiar and is well discussed on Edugeek. The core issue appears to be SMB related (top of my head).
Posted

Do a data capture on server and analysis my guess is you have a rogue broken card or.machine on network flooding network or you have chimney or less enabled on nice card of server search for disable chimney on google

There are lots of setting to try disabling

Posted
I had something similar YEARS ago. Turned out to be a virus that was writing scheduled tasks then deleting them. Pretty sure this won't be the case for you, but thought it was worth a mention just in case.
Posted
Is the Aruba 5412Rzl2 fully up to date with firmware. The reason I ask is we are also having very similar issues to your network. And the only major change over the summer at my school was a new Aruba 5412Rzl2. Very interesting you also have the same switch at your core.
Posted

@eddyc If it helps the issues I was seeing reached a braking point today with the network freezing every 5 minutes.

Updated the core switch 5412Rzl to the latest September released firmware and the issues looked to have stoped. Will find out Monday when we have a full number of users back on the network if it was a 100% resolution.

Posted

The constant reference to user induced "Loops" is in many cases, pure speculation especially it seems in the education sector and is often suggested after all "obvious" causes have been exhausted.

 

Yes, it does happen, sometimes deliberately by users but more often I find it to be due to network admins not having documented LAGs/Trunks or multichannel connections correctly.

These often get fixed by accident as secondary links get unplugged in a blind panic!

Many so called loops get killed by unplugging redundant links that have lost LACP config or by the blind panic induced deployment of stp/rstp when really all you have done is add 10-15% CPU overhead on an already over taxed low cost switch!

 

In this case, it seems more likely that what you are seeing is more likely to be excessive multicast traffic.

If your switches are not correctly setup to handle this, a single PC sitting at one end of your LAN can bring everything down regardless of your VLAN settings.

 

Correctly set up multicast packets will only appear on ports that are participating in the same multicast group but this also means they can appear and often do on uplinks and inter connects.

 

Do check the real estate for high levels of multicast traffic, the timing issue is clearly a key factor and that's why I'm going to suggest you look at your desktop PCs first.

 

Anything that uses an Intel or some Broadcom NICs (including some switch ports) is potentially to blame, this includes many HP desktops, Laptops and switches (but this also includes many other vendors).

 

The regular (hourly) occurrence is most likely due to a machine or groups of machines entering a sleep mode.

This is what catches people out as the machines are often unused when they are potentially the most active!

 

For years these NIC vendors have had issues with multicast flooding on both IP4 & 6 systems.

 

To fix it you need the latest and correct firmware and drivers installed everywhere, otherwise as a machine goes into energy save mode it can start to flood a segment with multicast packets.

As it wakes up the flood stops!

 

We have many sites some with dozens of VLAN segments and whilst VLANS will separate and isolate subnet broadcasts it's not going to stop multicast issues across the physical layer unless you have it properly configured.

 

Another common gotcha is the regeneration of machines with dodgy drivers through bad imaging practises.

Check your OS images and SCCM driver packs in case your images are harbouring out of date drivers that are affected by these multicast issues.

 

Here is a link to just one of the hundreds of articles published over the last few years, relating to similar issues.

http://packetpushers.net/good-nics-bad-things-blast-ipv6-multicast-listener-discovery-queries/ regarding theses issues.

Posted
The quick test for loops is flashing lights on switches instead of the twinkle that's normal. It spreads throughout a network and causes switches to crash's and become hubs or reboot and cause lack of connectivity
Posted
@eddyc If it helps the issues I was seeing reached a braking point today with the network freezing every 5 minutes.

Updated the core switch 5412Rzl to the latest September released firmware and the issues looked to have stoped. Will find out Monday when we have a full number of users back on the network if it was a 100% resolution.

 

Sorry for the delayed replies, it has been somewhat manic here the past week or so!

 

It sounds like the exact same issue then - Can I just check that you are now running KB.16.02.0013 built on 12-Sep-2016 and posted on 14-Sep-2016?

 

The freezing doesn't seem to be anywhere near as bad as it has been over the past week or so - a few seconds most when it has happened.

 

If that firmware seems to have helped you, I'll pencil in the upgrade here too.

 

Thanks Ed

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...