Jump to content

New Aruba core switch / old Netgear core issue ...


Recommended Posts

Posted (edited)

Hi

 

Over summer we are trying to migrate our network running on old Netgear switches to Aruba / HP 5412 with 10gb fibre links and Aruba HP core switches.

 

Before summer term, we installed the Aruba core next to the old Netgear core linked together using a 10gb SFP connection, servers and most of the netgear edge switch links was still on the old netgear core stack, we tested one HP edge switch connected to Aruba core in our IT office live and worked beautifully for months.

 

Now over summer we are migrating the rest over to the new core, we are having very broadcast storm like issues but how it happens is really throwing me off. (We are not moving the old netgear edge over, we are replacing the switches, transceivers etc with Aruba stuff.

 

We started migrating switches over, and what happens (on certain switch stacks) when we move the fibre trunk over to the new Aruba core (from 1gb fibre to 10gb on Aruba setup) the network begins deteriorating with remote services, network drives failing, unplugging it Network returns to normal fairly quickly. Sanity checking the suspect edge switches reveals no looped devices.

 

Sometimes after several plug, unplugging tests the network will fail to the point it requires a reboot of the core switches

 

Then the problematic edge switch stops being an issue, thinking it’s all ok, we carry on migrating switches ...

 

Then randomly, after a couple of successful switches installed, another location we plug the fibre in throws the same problem again! (I wait an hour after every switch migration to ensure no broadcast storms) the problem pretty much appears the moment we plug it into new fibre...

 

Now I find it hard to get my head round that 3 separate locations so far are causing a network loop when migrated ?

 

I’m also trying to learn deciphering the logs and how to switch on Spanning tree properly on the core.

 

At the moment I’m trying to move all the stuff over to the new Aruba core switch, unplug the old netgear and see if the old equipment was causing trouble ...

 

Any other ideas ? Just don’t want it randomly going down just as results day coming round ! Thanks !

 

The only pattern I can think of so far is the edge stack runs ok on the old 1gb fibre, then it gradually causes a network wide issue soon as you move this into a 10gb fibre to the Aruba ...

Edited by MrWu
Posted

Immediate first thought is spanning tree, are both the old and new trying to be the root?

 

And thoughts around the l3 routing will follow next if the old netgear is getting smashed.

 

Logs on both cores is where I'd be looking.

  • Thanks 1
Posted

Spanning tree would be my first thought too.

 

Have you got multiple connections between the edge switches and the 5412. Check the logs of each switch, as you lose connectivity you may need to do this via console to the core.

 

This is typical behavior of spanning tree, broadcast storms can build up. Also make sure you have enabled at least MSTP on the core switch and the edge switches.

  • Thanks 1
Posted (edited)

Thanks, only one link between old core and new core

 

Will check.. but network has been solid and quiet last night and also this morning all smooth, will move the remaining servers and most of the edge links and isolate the old core as soon as I can today.

 

Will migrate remaining edge links one by one, wait and analyse

 

On the Arube 5412s and ARUBA 2540 edge switches, is it a simple matter of going into GUI and flicking Spanning tree to ON?

 

I'm think of just turning Spanning tree on the last edge location that caused a storm, plus the Core switch to do a bit of trouble shooting when I plug it back in?

 

PS currently looking at setup (I'm not familiar Aruba switching however) Spanning tree on Core is off and on the Netgear Core its ON

Edited by MrWu
Posted

Further investigation, and looking at 3 separate incidents of the broadcast storm over 3 switch locations, it looks like they all have IP phones in that segment (it’s not on VLAN, we only have around than 10 IP phones in the whole school ) I do think when an IP phone (Siemens) gets disconnected and moved over to new HP 10 gb segment it causes issues. The phones are not daisy chained to a PC, just directly plugged into the switch.

 

I have migrated edge locations for ICT rooms with no issues for example, and they had no ip phones.

Posted

We had a similar issue moving some of our core infrastructure from 1GB to 10GB.

Just before I started at my current place, the previous NM had an external company swap out old 3Com 1GB core switches with new Allied Telesis 10GB switches. Config was just manually copied from the old switches to the new and nothing was really changed. Over the next few weeks (and months) we experienced more and more packet loss, dropping connections, slow logins and just general terrible network performance.

After some investigating with Wireshark, it looked like our WiFi was to blame - sounds stupid but it was very obvious after looking at Wireshark!

After reconfiguring our entire WiFi system the issues went away. For whatever reason, these issues had only shown themselves after we upgraded to 10GB and had never been an issue before. I still can’t quite get my head around it but it was very strange!!

Might be worth trying wireshark and see if you can see any obvious issues.

  • Thanks 1
Posted

Thanks all again for suggestions

 

I have turned on spanning tree on the core and have nearly migrated all edge to 10gb bar one

 

The edge location throwing out CRC errors was sorted with a newer OM3 patch lead

 

There was one switch that I isolated that was causing broadcast storm like instabilities to the network, check equipment plugged into it one by one tomorrow

 

Overall, looking at the stats of the HP core, things are all green at the moment, no errors so will carry on monitoring ...

Posted
Turns out to be my Siemens IP phones (only certain models I believe) that causes a broadcast storm issue that switches can’t detect .. (as in it’s not a BPDU ?)
Posted
Turns out to be my Siemens IP phones (only certain models I believe) that causes a broadcast storm issue that switches can’t detect .. (as in it’s not a BPDU ?)
Out of interest which Siemens phones are you using.. have a similar problem but not as accute. We do also use Siemens phones and its Poe switches that have issues....
Posted (edited)
Out of interest which Siemens phones are you using.. have a similar problem but not as accute. We do also use Siemens phones and its Poe switches that have issues....

 

Well I'm stumped....

 

I have siemens optipoint 410

 

Turns out that the cabinets that had issues when switches was replaced had these phones (which were plugged into POE ports on the old Netgears) were mistakenly plugged into a non POE HP switch...)

 

Should not power up the phones right? Believe it or not, and I watched this happen, there is a flicker and occasional link on the switch port (I see it flash intermittently) there is no display on the IP phone itself but there's network activity somewhere and then it brings the segment of that network down, no spanning tree messages though.

 

So is the Siemens phone's internal PC/Phone NIC forming a bridge of some sort, passively?

 

At least now I have documented to my team and colour coded those cables !

Edited by MrWu
  • Thanks 1

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...