Jump to content

Recommended Posts

Posted (edited)

With our core network being relatively stable over the years, it's very rare that we have to diagnose potentially network related issues. We have had a relatively large number of reports that either the network is very slow or things aren't connecting. This can range from two PCs, connected to two different switches in different buildings are auto-negotiating 10MB or 100MB, not 1GB (or just say there detecting and then saying no network connected), reports that file transfers are incredibly slow, despite the file servers and storage devices looking perfectly healthy, issues with wireless clients dropping off or getting very slow download speeds and probably more that I can't remember off the top of my head.

 

We are a HP ProCurve (5406zl core and 2530 edges and the odd Aruba CX6000) house with a Ruckus ZD1200 controller and Ruckus R500 APs. Short of going into each switch individually and checking the event log (which isn't the easiest to read), is there an easy way to diagnose where network issues are coming from?

 

Here's a snippet from our latest logs from switches

 

Core switch:

W 02/06/24 13:40:27 02581 ip: AM1: IPv4: Duplicate IPv4 address 10.22.100.1 is detected on port B7 in VLAN 100 with a MAC address of 
W 02/06/24 13:40:56 00331 FFI: AM1: port E23-High collision or drop rate. See help.
I 02/06/24 13:43:09 00166 update: AM1: Console inactivity timer is reset as the file transfer is in progress.
W 02/06/24 13:48:45 00332 FFI: AM1: port A5-Excessive Broadcasts. See help.
I 02/06/24 13:53:09 00166 update: AM1: Console inactivity timer is reset as the file transfer is in progress.
W 02/06/24 13:53:32 00332 FFI: AM1: port A5-Excessive Broadcasts. See help.
W 02/06/24 13:55:01 00331 FFI: AM1: port F3-High collision or drop rate. See help.
W 02/06/24 13:57:24 00331 FFI: AM1: port E20-High collision or drop rate. See help.
W 02/06/24 13:57:24 00331 FFI: AM1: port E21-High collision or drop rate. See help.
W 02/06/24 13:57:24 00331 FFI: AM1: port E22-High collision or drop rate. See help.
W 02/06/24 13:57:24 00331 FFI: AM1: port E21-High collision or drop rate. See help. help.
W 02/06/24 13:57:24 00331 FFI: AM1: port E22-High collision or drop rate. See help.help.
W 02/06/24 13:57:24 00331 FFI: AM1: port E23-High collision or drop rate. See help.help.
W 02/06/24 13:58:37 02581 ip: AM1: IPv4: Duplicate IPv4 address 
W 02/06/24 13:58:40 00332 FFI: AM1: port A5-Excessive Broadcasts. See help.
I 02/06/24 14:03:10 00166 update: AM1: Console inactivity timer is reset as the file transfer is in progress.
W 02/06/24 14:08:55 00332 FFI: AM1: port A5-Excessive Broadcasts. See help.
I 02/06/24 14:13:10 00166 update: AM1: Console inactivity timer is reset as the file transfer is in progress.
W 02/06/24 14:14:03 00332 FFI: AM1: port A5-Excessive Broadcasts. See help.

Our A-E ports are our fibre links to our edge switches, though I don't know if these broadcast messages are related to the use of Impero, though it could be related to spanning tree. What does look concerning is relating to the duplicate IP address for the MAC address for our default gateway for VLAN100, on ports for two different switches in two different buildings.

 

Logs for the switch connected to the switch on fibre port A5:

W 02/06/24 12:57:25 00331 FFI: port 9-High collision or drop rate. See help.
W 02/06/24 12:57:25 00331 FFI: port 10-High collision or drop rate. See help.
W 02/06/24 12:57:25 00331 FFI: port 15-High collision or drop rate. See help.
W 02/06/24 12:59:08 00332 FFI: port 26-Excessive Broadcasts. See help.
I 02/06/24 13:06:48 00166 update: Console inactivity timer is reset as the file transfer is in progress.
W 02/06/24 13:09:23 00332 FFI: port 26-Excessive Broadcasts. See help.
W 02/06/24 13:13:49 00331 FFI: port 5-High collision or drop rate. See help.
I 02/06/24 13:13:54 00077 ports: port 8 is now off-line
I 02/06/24 13:13:57 00435 ports: port 8 is Blocked by STP
I 02/06/24 13:13:57 00076 ports: port 8 is now on-line
I 02/06/24 13:14:03 00077 ports: port 8 is now off-line
I 02/06/24 13:14:08 00435 ports: port 8 is Blocked by STP
I 02/06/24 13:14:08 00076 ports: port 8 is now on-line
I 02/06/24 13:14:18 00077 ports: port 8 is now off-line
I 02/06/24 13:14:23 00435 ports: port 8 is Blocked by STP
I 02/06/24 13:14:23 00076 ports: port 8 is now on-line
W 02/06/24 13:14:30 00332 FFI: port 26-Excessive Broadcasts. See help.
I 02/06/24 13:16:48 00166 update: Console inactivity timer is reset as the file transfer is in progress.
W 02/06/24 13:19:38 00332 FFI: port 26-Excessive Broadcasts. See help.
W 02/06/24 13:24:25 00332 FFI: port 26-Excessive Broadcasts. See help.
W 02/06/24 13:24:25 05737 FFI: port 26-Excessive Multicasts. See help.
W 02/06/24 13:24:45 00331 FFI: port 6-High collision or drop rate. See help.
W 02/06/24 13:24:45 00331 FFI: port 7-High collision or drop rate. See help.
W 02/06/24 13:24:45 00331 FFI: port 8-High collision or drop rate. See help.
W 02/06/24 13:24:45 00331 FFI: port 9-High collision or drop rate. See help.
W 02/06/24 13:24:45 00331 FFI: port 10-High collision or drop rate. See help.
W 02/06/24 13:24:45 00331 FFI: port 15-High collision or drop rate. See help.
I 02/06/24 13:25:40 00077 ports: port 7 is now off-line

The STP entries suggests there is a loop on those ports - but I've checked and double checked - there isn't. Port 5 goes into a patch panel, then into printer at the other end. Port 6 goes into the next patch panel port, into the same double socket next to the printer port and into a HP USB C dock for the teacher laptop.

 

Logs for the switch connected to fibre port B7:

W 02/06/24 12:28:26 00332 FFI: port 50-Excessive Broadcasts. See help.
W 02/06/24 12:33:34 00332 FFI: port 50-Excessive Broadcasts. See help.
W 02/06/24 12:33:34 05737 FFI: port 50-Excessive Multicasts. See help.
W 02/06/24 12:38:41 05737 FFI: port 50-Excessive Multicasts. See help.
I 02/06/24 12:45:22 00077 ports: port 47 is now off-line
I 02/06/24 12:45:31 00435 ports: port 47 is Blocked by STP
I 02/06/24 12:45:31 00076 ports: port 47 is now on-line
I 02/06/24 12:46:27 00077 ports: port 47 is now off-line
I 02/06/24 12:46:30 00435 ports: port 47 is Blocked by STP
I 02/06/24 12:46:30 00076 ports: port 47 is now on-line
W 02/06/24 12:48:56 00332 FFI: port 50-Excessive Broadcasts. See help.
W 02/06/24 12:49:37 00331 FFI: port 20-High collision or drop rate. See help.
W 02/06/24 12:54:04 00332 FFI: port 50-Excessive Broadcasts. See help.
W 02/06/24 12:54:04 05737 FFI: port 50-Excessive Multicasts. See help.
W 02/06/24 12:55:05 00331 FFI: port 2-High collision or drop rate. See help.
W 02/06/24 12:55:05 00331 FFI: port 14-High collision or drop rate. See help.
W 02/06/24 12:55:05 00331 FFI: port 30-High collision or drop rate. See help.
W 02/06/24 12:55:05 00331 FFI: port 32-High collision or drop rate. See help.
W 02/06/24 12:55:05 00331 FFI: port 36-High collision or drop rate. See help.
W 02/06/24 12:55:05 00331 FFI: port 40-High collision or drop rate. See help.
W 02/06/24 12:55:05 00331 FFI: port 42-High collision or drop rate. See help.
W 02/06/24 12:55:05 00331 FFI: port 45-High collision or drop rate. See help.
W 02/06/24 12:55:05 00331 FFI: port 47-High collision or drop rate. See help.
W 02/06/24 12:55:05 00331 FFI: port 48-High collision or drop rate. See help.
W 02/06/24 12:59:11 00332 FFI: port 50-Excessive Broadcasts. See help.
W 02/06/24 13:09:26 00332 FFI: port 50-Excessive Broadcasts. See help.
W 02/06/24 13:11:29 00331 FFI: port 20-High collision or drop rate. See help.
W 02/06/24 13:14:34 00332 FFI: port 50-Excessive Broadcasts. See help.
W 02/06/24 13:19:21 00332 FFI: port 50-Excessive Broadcasts. See help.
W 02/06/24 13:24:28 00332 FFI: port 50-Excessive Broadcasts. See help.
W 02/06/24 13:24:28 05737 FFI: port 50-Excessive Multicasts. See help.
W 02/06/24 13:27:53 00331 FFI: port 2-High collision or drop rate. See help.
W 02/06/24 13:27:53 00331 FFI: port 14-High collision or drop rate. See help.
W 02/06/24 13:27:53 00331 FFI: port 30-High collision or drop rate. See help.
W 02/06/24 13:27:53 00331 FFI: port 32-High collision or drop rate. See help.
W 02/06/24 13:27:53 00331 FFI: port 36-High collision or drop rate. See help.
W 02/06/24 13:27:53 00331 FFI: port 40-High collision or drop rate. See help.
W 02/06/24 13:27:53 00331 FFI: port 42-High collision or drop rate. See help.
W 02/06/24 13:27:53 00331 FFI: port 45-High collision or drop rate. See help.
W 02/06/24 13:27:53 00331 FFI: port 47-High collision or drop rate. See help.
W 02/06/24 13:27:53 00331 FFI: port 48-High collision or drop rate. See help.
W 02/06/24 13:29:36 00332 FFI: port 50-Excessive Broadcasts. See help.

 

If anyone could provide tips on how to help diagnose this, that would be great.

Edited by CHiLL
Posted
Is it possible the laptop is causing the loop, if it's connected to both the physical network via the USB-C hub and WiFi? Shouldn't be an issue, but doesn't hurt to disable one or other.
Posted

Blocked by STP is always logged I think when a port comes on line, then it releases it once it is happy there is no loop , as long as its a few seconds is correct behaviour

 

the Duplicate IP addresses i'd try and track down first what the MAC is from the OUI and what it should be

  • Thanks 1
Posted

I have just come out of the other side of this so good luck. My network was like yours, up and down all over the place. Mine turned out to be a Yealink phone going crazy.

 

Made any other changes lately? Changed any cables? mini switches?

  • Thanks 1
Posted (edited)
Is it possible the laptop is causing the loop, if it's connected to both the physical network via the USB-C hub and WiFi? Shouldn't be an issue, but doesn't hurt to disable one or other.

I'll do some more digging.

 

Blocked by STP is always logged I think when a port comes on line, then it releases it once it is happy there is no loop , as long as its a few seconds is correct behaviour

 

the Duplicate IP addresses i'd try and track down first what the MAC is from the OUI and what it should be

Thanks for clarifying the STP, I wasn't aware that was the behaviour and you've saved me chasing my tail! The MAC address displayed is that of the Core Switch, which is somewhat concerning.

Edited by CHiLL
Posted

Possibly not related in your case, however had this when I installed Ruckus a few summers back, using the R750's. Somehow during a firmware update (we think) a couple of AP's randomly enabled network bridging. Very similar STP errors as seen on your switch above. It would gradually get worse, as the network drop outs would cause the other ruckus AP's to go into bridge mode and create more loops and spread the chaos, reset all ap's would fix it for a while then it would start again. It is really strange that they are being rate limited on ethernet, I would only think that network interference on that actual cable would cause that. However, your printer on port 5 may have a very basic network card and is getting swamped causing it to freak out, seen that on HP office printers when we had issues.

 

Best of luck in fault finding.

  • Thanks 1
Posted
Is it possible the laptop is causing the loop, if it's connected to both the physical network via the USB-C hub and WiFi? Shouldn't be an issue, but doesn't hurt to disable one or other.

I don't think it's specifically either the printer (Xerox C405 with 1Gb) or the laptop, as swapping the ports around still produced the error and I'm seeing the same logs on other switches. Thanks for the tip on Ruckus, though sadly we're out of support with them and aren't eligible for updates (or support if I force an update and brick it).

Posted (edited)

Do you use VRRP at all on the core switches. I've seen such duplicate IP errors on HP Procurve switches and the MAC address is the other switch.

 

Doubtful it is both wifi and ethernet on the same device. For one they'd be using different MAC addresses so a layer 2 loop should not be a thing. If you had a duplicate IP address then either device night not work if it is a client machine. If it is a gatway then yep can cause bigger issues.

 

The excessive broadcasts is worthy of further investigation.

Edited by Davit2005
  • Thanks 1
Posted
Do you use VRRP at all on the core switches. I've seen such duplicate IP errors on HP Procurve switches and the MAC address is the other switch.

Our core config shows that VRRP is disabled. Since I don't know what VRRP is or what effects changing that setting would do, I'll not be doing it during term time at least.

 

The excessive broadcasts is worthy of further investigation.

If all our workstations are on the same VLAN, would that cause an issue regarding Impero broadcasts? I presume Impero broadcasts work just like network broadcasts in that way, right? I'm not actually sure. We have spanning tree enabled on each switch, so that would trap broadcast traffic to only the switch the clients are connected to?

Posted

In that case I'd be chasing down the duplicate IP. If it is sharing the same IP address as your gateway it's not good.

 

VRRP is a way to get redundant gateways (without logical stacking, etc). The IP address is shared between 2 different routers but they have a different address at least on one of them as the actual vlan address, the VRRP IP (gateway) address can move across if one router fails. If you have not got it enabled then fair enough :-)

  • Thanks 1
Posted
In that case I'd be chasing down the duplicate IP. If it is sharing the same IP address as your gateway it's not good.

 

VRRP is a way to get redundant gateways (without logical stacking, etc). The IP address is shared between 2 different routers but they have a different address at least on one of them as the actual vlan address, the VRRP IP (gateway) address can move across if one router fails. If you have not got it enabled then fair enough :-)

I'm trying to look at that now, but being pulled from pillar to post, you know how it is. I have noticed that the core switch clock is an hour ahead, as the log times were out compared to the other switches. While it might not have a direct effect, it's worth fixing. SNTP is set to our DCs, which are the correct time, Time Zone is set to "60" and DST on the is set to "Western-Europe". I wonder if setting Time Zone to 0 would bring it to the correct time and if so, would it adjust when we change to BST?

Posted

Irrespective of DST/BST/GMT the time zone will need amending to 0. By amending this it will sync to your designated NTP server and correct the time accordingly. In other words, whatever the time is on your DC will exactly be the same for your core.

 

The command is simply

time timezone 0

 

And set it to poll your DC for the time every 10 minutes

 

sntp 600

  • Thanks 1
Posted

Turns out we did indeed have a loop on our network and removing it stopped all the multicast warnings that were appearing on every switch.

 

I remember previously using a free version of HP IMC network monitor (like 10-15 years ago), but it's not been free for a long time. Spiceworks once also did an on-site hosted platorm that you just added your SNMP targets and it gave you an overview and alerts. I have tried PRTG and Zabbix, but I found them extremely complicated to set up and eventually gave up as we had other things on the go at the time. Are there any other decent free solutions?

 

Irrespective of DST/BST/GMT the time zone will need amending to 0. By amending this it will sync to your designated NTP server and correct the time accordingly. In other words, whatever the time is on your DC will exactly be the same for your core.

 

The command is simply

time timezone 0

 

And set it to poll your DC for the time every 10 minutes

 

sntp 600

Thanks, I'll have a look at that next week, when nobody is in.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...