Jump to content

Recommended Posts

Posted (edited)

Morning all,

 

Well this one is truly mashing my swede.

 

Not long been at a new place and having network dropouts. Just a minute or two maybe once or twice a day during school hours, but definitely get a couple overnight too ( set a -t ping going.)

 

Cisco switches throughout (except a couple of lil Netgears in the office,) only major change I've made here was moving to Jamf over Christmas. Have a Mac mini here as cache which seems to be working ok. Could be a red herring, could maybe not be.

 

No real pattern to it it seems, and whenever I seem to find one suddenly there's not one. However logging onto the core switch (which is four switches stacked, two in the server room and two across in another building) things get very curious.

 

Firstly, logs have this repeating loads:

 

certerrors.JPG

 

Next, this blip appears just the one time I noticed:

 

multiplelldpsmall.JPG

 

A full-on crash (amidst the SSL cert howling on it) looks like this (lots and lots of pics of this one:)

 

crashpart1.JPG

 

crashpart2.JPG

 

crashpart3.JPG

 

crashpart4.JPG

 

crashpart5.JPG

 

crashpart6.JPG

 

crashpart7.JPG

 

crashpart8.JPG

 

crashpart9.JPG

 

crashpart10.JPG

 

crashpart11.JPG

 

crashpart12.JPG

 

From dashing to the server room it does look for all the world like switch #2 in the stack is flapping somehow, but I'm not seeing at the moment what those logs are trying to show me.

 

If anyone who is a whizz with Cisco gear is looking and it's really obvious please do let me know! If I crack it in the meantime I'll post up what fixes it. Really has me stumped for now though.

 

Thanks all.

Edited by JRA
Posted (edited)

I haven't looked at all the logs as for whatever reason I am struggling to open the images, but it looks like something is going wrong with the stacking.

 

What model switches do you have and how are they stacked? You said they are split across 2 buildings, so I am guessing you're stacking using VSS? If so I guess this is going over fibre optic?

 

Are you able to run "Show Switch Detail" and "Show Version" in the CLI please and post the result here (or via PM) please?

 

Thanks

Edited by FN-GM
  • Thanks 1
Posted
I haven't looked at all the logs as for whatever reason I am struggling to open the images, but it looks like something is going wrong with the stacking.

 

What model switches do you have and how are they stacked? You said they are split across 2 buildings, so I am guessing you're stacking using VSS? If so I guess this is going over fibre optic?

 

Are you able to run "Show Switch Detail" and "Show Version" in the CLI please and post the result here (or via PM) please?

 

Thanks

 

Thanks there! :)

 

Ok hopefully these pictures come out. Cisco SX550X switches, switches 1 and 2 in the server room, switches 3 and 4 in another building. All over fibre but not sure what VSS is tbh.

 

Snips of all those bits here:

 

clibits.JPG

 

switch1.JPG

 

switch2.JPG

 

switch3.JPG

 

switch4.JPG

 

Interesting #2 switch says it's on "standby".

 

Cheers again. :)

Posted (edited)

These are quite different from what I am used to. They are from the small business line. The commands are different and the technology by the looks of it.

 

Are you able to run “Show Switch” please?

Edited by FN-GM
  • Thanks 1
Posted
A bit of a guess here but there could be an issue with your fibre connections. Maybe the fibre cable between the 2 buildings. Are you able to see if the ports are dropping out?
  • Thanks 1
Posted

Running that gives me "incomplete command" - actually did have to turn on telnet earlier as it looks like nobody's ever putty-ed/CLI-ed into these before now.

 

Did wonder that briefly but I'm not sure on how I'd determine that. Anything useful in the logs? (if you can even see those there ofc.) Sure could be a possibility. I suppose a good test would be to -t ping switch #2 in the stack for a day and see if that ping ever drops throughout the dropouts. I myself (bear with me here) connect to another switch which is on an LACP trunk to switches #1 and #2 so if I drop packets to switch #2 during a dropout it's a clue switch #2 is upset, and if I don't it's likely connectivity between switches #1 and #2 <--> switches #3 and #4.

 

Worth a stab?

Posted
Any chance one of those stacking ports isn't configured right and the spanning tree is kicking in because it thinks it's a network loop?

I mean it could be I guess, any easy ways to tell? I haven't changed anything stacking-wise at all. Which I know is not the same thing as saying nothing has changed, more meaning it wasn't me if so I swear!

 

Pinging a ubuntu box I had sat around which is now in switch #2, also pinging a switch on the far end of the fibre run in the other building, logging both to text files. Hoping that'll show me a clue in if they BOTH drop out at the same time then switch #2 might be the issue, and if that ubuntu box ping carries on at any time the other drops out then more likely connectivity between switch #2 and the other building might be it.

 

Thanks everyone for joining me on my annoying network adventure! :)

Posted

And as luck would have it, just had a dropout. Lost the ping to the other building AND the ping to the device going into switch #2.

 

So I'm thinking some issue with that switch as I can still get onto the stack by the IP address, and switch 1 shows as the only member of the stack.

 

Of course, only lasts a minute then it's all back on.

 

Urgh. :/

Posted

I wondering if it might be switch 1 that is the problem if all the others drop off? When switch 1 loses sight of the other two... do the other two remain up and functional?

 

I might also troubleshot by unplugging the redundant paths between the switches... that way you can test what order the switches are dropping out of the cluster and through systematically working through the options for the physical topology isolate which switch (or cable/port) is causing the drops.

 

To be honest, if this was an old procurve network I'd be suggesting you check for spanning tree changes, because a minute is about the length of time would take one of those to reconverge spanning tree after some topology change.... and if it is spanning tree, you can see/infer from the show spanning-tree commmand which port/switch triggered the event.

  • Thanks 1
Posted (edited)

I'll see if I can try that somehow, that might be worth finding out indeed.

 

Ok with my Sherlock Holmes hat on I've sniffed out the following schedule to the dropouts:

 

16th Jan:

14:00

18:00

20:00

22:00

 

17th Jan:

 

00:00

02:00

04:00

06:00

10:00

14:00

16:00

20:00

22:00

 

18th Jan:

00:00

02:00

06:00

08:00

 

To me that heavily implies it's something traffic-ey overloading the poor things. Currently exploring options in the phone system spamming out updates, maybe the iPads now (recently moved to Jamf.) After that I dunno, cosmic rays I guess.

 

Thanks everyone for carrying on with me. :)

Edited by JRA
Posted

Juicy juicy clue this morning from Jamf; dropouts seem to coincide with the iPads "checking in".

 

I knew it was Apple's fault. ;)

 

Let's see if (hopefully) we can throttle that...

  • Thanks 1
Posted

The regularity of the events is astounding. Points to a software/firmware issue in my mind. Just checking the are running latest/stable firmware?

 

Or (long shot) power problems - if you've got UPS's but the switches are not protected, it might be worth cross referencing with voltage drops/spikes etc.

  • Thanks 1
Posted
The regularity of the events is astounding. Points to a software/firmware issue in my mind. Just checking the are running latest/stable firmware?

 

Or (long shot) power problems - if you've got UPS's but the switches are not protected, it might be worth cross referencing with voltage drops/spikes etc.

 

Could be for sure. I'll check over those if my Jamf theory doesn't pan out, though I'm pretty sure it's quacking like that particular duck at this point. Cheers. :)

Posted

Sounds slightly similar to an issue I had not so long ago

 

Do you have anyway of monitoring wireless network traffic? A measure I put in place that helped was to limit our Chromebooks bandwidth to Google Update, as whenever a ChromeOS update as released, all 700 devices would try to update, and as standard, updates were given top priority, so my network crashed.

 

I would also look into restricting Apple Bonjour as much as possible.

  • Thanks 1
Posted

So I'm fairly certain at this stage it's something iPad-ey. Dropped in for a few quiet Sunday hours, scheduled the student wifi off over the weekend and have had NO dropouts.

 

Sounds slightly similar to an issue I had not so long ago

 

Do you have anyway of monitoring wireless network traffic? A measure I put in place that helped was to limit our Chromebooks bandwidth to Google Update, as whenever a ChromeOS update as released, all 700 devices would try to update, and as standard, updates were given top priority, so my network crashed.

 

I would also look into restricting Apple Bonjour as much as possible.

 

Thanks for this one. Had a good rummage through your thread, also have some options to explore in our (ancient) Ruckus wifi along those lines. Happy to say though the problem is at least pinned down!

 

Cheers everyone who helped. :)

  • Thanks 1
Posted

Ok, now I think I might be even closer to it. It's likely the iPads causing the issue to manifest rather than them being the actual culprit. Which is a shame, I enjoyed it being Apple's fault.

 

Looking at the spanning tree config on the core stack, I see this:

 

0ncore.JPG

 

Now, my knowledge of the guts of spanning tree isn't amazing, but I should imagine (and please do steer me right if I'm not right) that I should INSTEAD be seeing the root bridge ID as being the bridge ID here (as in the "boss" route for spanning tree should be the links twixt the four switches in the stack which make up my core switch here.) And also, that priority for the "boss" bridge should be 0. Or lower than default at any rate.

 

So, what the heck is that odd bridge ID? Well it's this:

 

edge.JPG

 

Looks like one of the edge switches in an out-building we have has the LAG connection (going to switches #1 and #2 of the core stack) with the root bridge ID! And by my reading of it, port TE2/0/14 is a major culprit in the flap logs, which plugs into my favourite suspect in this, being switch #2 of the core stack (the other being in port 14 of core switch #1.)

 

A couple more edge switches to rub it in:

 

edge2.JPG

 

edge3.JPG

 

That looks to be worth re-working. As I mention not super with Cisco gear so if anyone passing can check my thinking that'd be great.

Posted

Normally we'd set the core to have the lowest bridge ID, I usually set to 8192.

 

That said even with them all at the same priority, the root will be elected based on MAC address / Switch age.

 

Setting the cores bridge ID lower would probably be a good start.

 

With all that said we've had a lot of issues over the years with this series of switches and instability, can't believe Cisco ever put their name to them.

  • Thanks 1
Posted

Cheers for that, from what I'm digging up setting that priority lower looks indeed to be the way (I take it you mean set the priority lower, not the bridge ID.) No idea how I knock off the root port yet but will be one to try as part of all this.

 

With all that said we've had a lot of issues over the years with this series of switches and instability, can't believe Cisco ever put their name to them.

Haha deep joy, my network here comprises these Cisco small business switches.

Posted
Cheers for that, from what I'm digging up setting that priority lower looks indeed to be the way (I take it you mean set the priority lower, not the bridge ID.) No idea how I knock off the root port yet but will be one to try as part of all this.

 

Yep, sorry priority lower

  • Thanks 1
Posted (edited)

Okay so we're still getting them. Student wifi off overnight and no dropouts, student wifi on in the morning, dropouts (mostly) every even-numbered hour throughout the day.

What I'm wondering is if the routing here is contributing to it. The network transitioned to VLANs from a flat network about 2 years back, but the next hop is on the same VLAN (VLAN1, default) as the rest of the switches. I have made a diagram (beautifully) in MS paint:

 

"Normal" network with the next hop on a firewall <> core switch VLAN (example IPs: )

 

nexthopvanilla.png

 

This network:

 

nexthopus.png

 

So the thing I'm thinking is, the Jamf check-in from the ipads is looking like a broadcast storm on VLAN1 to the core switch, which then wets itself and breaks the stack. The stack then quickly rebuilds and we're all fine until the next one. On the rare occasions there are NO dropouts with the student wifi on, on the even-numbered hours, we may just be squeaking enough traffic through to not upset it.

 

To that end I'm tempted to take a bravey pill and turn off spanning tree completely just on the core stack and let it chew through all that traffic just for the 2 min it needs to, see if that does it.

 

Failing that the core stack is due a firmware update later so I'll run that too, changing two variables in one night like a good scientist.

 

Anyone passing have any thoughts on that one? Thanks whoever's still reading!

Edited by JRA
Posted

Does the Smoothwall detect a spike in traffic? Also, have you updated your Smoothwall at all? I had an issue using the SFP port and moved to a standard network port which resolved a similar issue.

 

I also wonder whether the iPads are talking to each other across your network, and this communication is causing spanning tree to kick in. Meraki has a NAT mode (all devices connecting to the SSID are effectively isolated from everything else on the network) which I put the Chromebooks on and this defintely helped.

  • Thanks 1
Posted
Does the Smoothwall detect a spike in traffic?

 

Had it in mind to give that another look. Did check that early on in the process but didn't see anything, though wasn't sure if it'd dropped before anything registered.

 

Also, have you updated your Smoothwall at all?

 

Not super-recently, end of Sept, still on Leeds 62.

 

I had an issue using the SFP port and moved to a standard network port which resolved a similar issue.

 

Oooohh, yeah we are in an SFP. I might keep that idea in my pocket. :)

 

I also wonder whether the iPads are talking to each other across your network, and this communication is causing spanning tree to kick in. Meraki has a NAT mode (all devices connecting to the SSID are effectively isolated from everything else on the network) which I put the Chromebooks on and this defintely helped.

Did spot that from your earlier messages and isolating clients from other clients on the same APs now in Ruckus, though haven't turned on the option available to isolate from all hosts on subnet (option there too to whitelist GW, might also try and re-jig the NIC on the Mac server/s to live in that VLAN also for the cache-ing should I enable that one.)

 

Thanks much for all that.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...