Jump to content

Wireshark Capture ARP Broadcasts - Do I have a loop?


Recommended Posts

Posted

Hi All,

 

We've been experiencing excessive broadcasts in my company for a while now which sometimes cause brief outages. I started running Wireshark to capture broadcasts during these storms by mirroring the uplink port of one of the switches. The output of the capture is as follows:-

 

289837 2013-11-04 16:43:46.503029000 Vmware_a2:00:c0 Broadcast ARP Who has 10.163.255.5? Tell 10.163.255.89

289838 2013-11-04 16:43:46.503036000 Vmware_a2:00:c0 Broadcast ARP Who has 10.163.255.5? Tell 10.163.255.89

289839 2013-11-04 16:43:46.503044000 Vmware_a2:00:c0 Broadcast ARP Who has 10.163.255.5? Tell 10.163.255.89

289840 2013-11-04 16:43:46.503053000 Vmware_a2:00:c0 Broadcast ARP Who has 10.163.255.5? Tell 10.163.255.89

289841 2013-11-04 16:43:46.503060000 Vmware_a2:00:c0 Broadcast ARP Who has 10.163.255.5? Tell 10.163.255.89

289842 2013-11-04 16:43:46.503066000 Vmware_a2:00:c0 Broadcast ARP Who has 10.163.255.5? Tell 10.163.255.89

289843 2013-11-04 16:43:46.503071000 Vmware_a2:00:c0 Broadcast ARP Who has 10.163.255.5? Tell 10.163.255.89

289844 2013-11-04 16:43:46.503078000 Vmware_a2:00:c0 Broadcast ARP Who has 10.163.255.5? Tell 10.163.255.89

289845 2013-11-04 16:43:46.503083000 Vmware_a2:00:c0 Broadcast ARP Who has 10.163.255.5? Tell 10.163.255.89

289846 2013-11-04 16:43:46.503089000 Vmware_a2:00:c0 Broadcast ARP Who has 10.163.255.5? Tell 10.163.255.89

289847 2013-11-04 16:43:46.503094000 Vmware_a2:00:c0 Broadcast ARP Who has 10.163.255.5? Tell 10.163.255.89

289848 2013-11-04 16:43:46.503100000 Vmware_a2:00:c0 Broadcast ARP Who has 10.163.255.5? Tell 10.163.255.89

 

This literally goes on for hundreds or thousands of packets within the same second. Does this mean I have a loop somewhere in the network causing duplicate ARP Broadcast packets? The device (10.163.255.89) is a server for video conferencing units and 10.163.255.5 is a video conferencing unit. When we get these broadcast storms, the captures seems only to pick up these ARPs from devices on the video conferencing VLAN which is an end-to-end VLAN, therefore it is geographically spread across most of the network with QoS priority so when this happens, it throttles all of the uplinks, sometimes causing outages. As far as I know MSTP is configured on all switches, but I'm relatively new to this network and there are over 200 switches.

 

Thanks in advance for any advice on this.

Posted

I'm looking into a similar issue with a VMWare host.

Ping any VM hosted on the VMWare cluster and the first view packets are lost, after this the VMWare host responds and the ping is sustained at the expected rate.

We are not yet certain if the issue is hardware based on the host (an HP3000 Blade) or the VMWare 4.1 virtual switch.

 

It's causing the client a lot of grief the original supplier and their VMWare experts have so far drawn a blank, but out testing has brought us to the Host/VMWare config which we can see has not been updated in 3.5 Years!

 

How is your Host/Virtual switch configured? and what physical resources do those IP addresses refer to?

 

In our case it appears to be an ARP issue at the VMWare level.

Posted
Do you always get the tell to the same device or is it different devices too?

 

Hi Marshall, it's not always the same device, but they are always related devices on the same subnet (10.163.255.0/24), ie they are always video conferencing units or servers for video conferencing.

Posted
I'm looking into a similar issue with a VMWare host.

Ping any VM hosted on the VMWare cluster and the first view packets are lost, after this the VMWare host responds and the ping is sustained at the expected rate.

We are not yet certain if the issue is hardware based on the host (an HP3000 Blade) or the VMWare 4.1 virtual switch.

 

It's causing the client a lot of grief the original supplier and their VMWare experts have so far drawn a blank, but out testing has brought us to the Host/VMWare config which we can see has not been updated in 3.5 Years!

 

How is your Host/Virtual switch configured? and what physical resources do those IP addresses refer to?

 

In our case it appears to be an ARP issue at the VMWare level.

 

You might be onto something. We don't manage the configuring of our UCS Switch so I can't see the configuration unfortunately - my vision from a networking perspective stops at the HP switches it is connected to. The .89 address is the virtual server for video conferencing (TMS) and the .5 address is a video conferencing unit.

Posted
Not a VMWare expert, but I did find this post on a previous known issue with ESX sending out excessive RARP broadcasts... worth checking you have the updates that fix it?

 

Possible reasons for RARP storms from an ESX host | VirtuallyHyper

 

I think this is a strong possibility. There are other captures I have that are almost purely Gratuitous ARPs from the same virtual TMS Server. I will have another read through before I decide how I'm going to implement it. I will let you know the result. Thank you all for your help.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...