Jump to content

Recommended Posts

Posted

Afternoon all,

Thought I would open this up for some suggestions as I feel like I've looked/tried everything.

 

Over the last week or two we've been having random network outages, our network cabs lock up and all port lights on solid... so first thought is a broadcast storm. This isn't every day and doesn't seem to have a set pattern.

It's affecting our entire network so must be localised to our primary core switch (we have a second site with it's own core which is in turn connected to the main core switch).

When this happens, we lose connectivity to our Dell VRTX running both of our hypervisors, almost like someone's unplugged the uplink to the VRTX.

 

So to the switch logs I go... absolutely nothing. Tells me it can't reach the ntp server (not surprised as it's running on the VRTX) and that's it. Spanning-tree doesn't show any ports being blocked either.

I can get on the hypervisors and all VMs via KVM without an issue, they tell me there's a network connection problem yet the VRTX Chassis and the 10Gb VRTX Switch logs don't show any issues.

 

Hardware wise we're running a Dell PowerEdge VRTX with two blade servers and the core switch is an Aruba 5406Rzl2 (J9850A).

 

Just wondering if anyone has experienced anything similar or if anyone has any suggestions.

Posted

To clarify then, the whole lock-up situation only occurs when your 2nd site's core is connected to your 1st? or are these two completely separate issues that are unlrelated?

 

To me it sounds like your routing, or Spanning Tree not working as you expect because you have a conflict between your two L3 switches and their subsequent networks. They might even think they're part of the same spanning tree for example and be recalculating constantly and fighting over who is the CST root. Do you have IP addressing conflicts/overlaps between the two sites also which would mean for instance if site 2x has the same subnet(s) your traffic might be heading that way when looking for your server...

Posted
To clarify then, the whole lock-up situation only occurs when your 2nd site's core is connected to your 1st? or are these two completely separate issues that are unlrelated?

 

To me it sounds like your routing, or Spanning Tree not working as you expect because you have a conflict between your two L3 switches and their subsequent networks. They might even think they're part of the same spanning tree for example and be recalculating constantly and fighting over who is the CST root. Do you have IP addressing conflicts/overlaps between the two sites also which would mean for instance if site 2x has the same subnet(s) your traffic might be heading that way when looking for your server...

 

No, sorry haven't made it clear on the setup. The connection between the cores isn't an issue, they've been setup like this and running without issue for years on both these switches and the previous ones.

Core 1 is the root, on Core 2 the root port is the fibre connection to Core 1 and spanning tree will deal with moving traffic to a backup site link if the fibre goes down.

 

If core 2 goes down, Core 1 stays up and Site 1 is unaffected.

If core 1 goes down, the whole network falls over. I guess you could technically discount core 2 as a core switch...

 

As this issue is network-wide, I'm fairly confident that the issue as is at the Core 1 end

Posted (edited)

Last topology change time?

 

If the last topology change is more recently than you'd expect, you can then hunt down which port the election was triggered from, and this helps narrow the problem down to a subset of the network.

 

 

If the topology change time is unrelated to this issue, then here's a horribly manual way of troubleshooting, but try half-splitting your network when it next happens:

 

Unplug/disable half the inter-switch uplink ports and wait a few minutes to see if it calms down, if not unplug half again, and repeat until the problem goes away - again hopefully you have then narrowed down the problem to a more manageable subset of the network.

Edited by psydii
Posted

I did have similar back in the mists of time and we eventually tracked it down to a network card causing a broadcast storm, we had been looking for loopbacks as that is what it looked like.

 

It seems odd that it is only happening every now and then though, so could someone be plugging in a device? Would it be worth unpatching any port not in use?

Posted
Just adding to TechMonkey's reply. This did happen with a certain network card (with certain firmware) would cause a broadcast storm when the device when into sleep mode. Next time it happens I would plug in a laptop with Wireshark and sniff those packets. It's soon become apparent what's causing it
  • Thanks 1
Posted
Worth checking that your 5406 is on the latest firmware. A couple of revisions ago there was an issue where a switch would reboot unexpectedly under a lower than normal CPU load.
Posted
Last topology change time?

 

If the last topology change is more recently than you'd expect, you can then hunt down which port the election was triggered from, and this helps narrow the problem down to a subset of the network.

 

 

If the topology change time is unrelated to this issue, then here's a horribly manual way of troubleshooting, but try half-splitting your network when it next happens:

 

Unplug/disable half the inter-switch uplink ports and wait a few minutes to see if it calms down, if not unplug half again, and repeat until the problem goes away - again hopefully you have then narrowed down the problem to a more manageable subset of the network.

At the moment, last topology change is when I rebooted the switch at 3pm yesterday. I haven't noticed any changes coinciding with the issue but I'll make sure I double check.

Posted
I did have similar back in the mists of time and we eventually tracked it down to a network card causing a broadcast storm, we had been looking for loopbacks as that is what it looked like.

 

It seems odd that it is only happening every now and then though, so could someone be plugging in a device? Would it be worth unpatching any port not in use?

 

Annoyingly this first started at 3am and couple of times before jumping to 5am, 9:30am, 10:30am, 11am, 12pm, 3pm. I did have a quick CCTV check to make sure there was no one in the building at 3am though...

Interestingly, this first occurred in the early hours of the day after one of our blade server motherboards went caput. Certainly thought (and hoped) the two were related with some form of loop caused by the VTRX switch but then it happened yesterday (Blade was back up and running Monday afternoon).

 

 

Just adding to TechMonkey's reply. This did happen with a certain network card (with certain firmware) would cause a broadcast storm when the device when into sleep mode. Next time it happens I would plug in a laptop with Wireshark and sniff those packets. It's soon become apparent what's causing it

Wireshark now ready to go and ports set to mirror (not sure why I didn't do this first to be honest). Now to play the waiting game!

Posted
Just adding to TechMonkey's reply. This did happen with a certain network card (with certain firmware) would cause a broadcast storm when the device when into sleep mode.

 

We had something similar-sounding which (I think) we've tracked down to a faulty 10G port / transceiver in a server.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...