MrWu Posted August 8, 2019 Posted August 8, 2019 (edited) Hi Over summer we are trying to migrate our network running on old Netgear switches to Aruba / HP 5412 with 10gb fibre links and Aruba HP core switches. Before summer term, we installed the Aruba core next to the old Netgear core linked together using a 10gb SFP connection, servers and most of the netgear edge switch links was still on the old netgear core stack, we tested one HP edge switch connected to Aruba core in our IT office live and worked beautifully for months. Now over summer we are migrating the rest over to the new core, we are having very broadcast storm like issues but how it happens is really throwing me off. (We are not moving the old netgear edge over, we are replacing the switches, transceivers etc with Aruba stuff. We started migrating switches over, and what happens (on certain switch stacks) when we move the fibre trunk over to the new Aruba core (from 1gb fibre to 10gb on Aruba setup) the network begins deteriorating with remote services, network drives failing, unplugging it Network returns to normal fairly quickly. Sanity checking the suspect edge switches reveals no looped devices. Sometimes after several plug, unplugging tests the network will fail to the point it requires a reboot of the core switches Then the problematic edge switch stops being an issue, thinking it’s all ok, we carry on migrating switches ... Then randomly, after a couple of successful switches installed, another location we plug the fibre in throws the same problem again! (I wait an hour after every switch migration to ensure no broadcast storms) the problem pretty much appears the moment we plug it into new fibre... Now I find it hard to get my head round that 3 separate locations so far are causing a network loop when migrated ? I’m also trying to learn deciphering the logs and how to switch on Spanning tree properly on the core. At the moment I’m trying to move all the stuff over to the new Aruba core switch, unplug the old netgear and see if the old equipment was causing trouble ... Any other ideas ? Just don’t want it randomly going down just as results day coming round ! Thanks ! The only pattern I can think of so far is the edge stack runs ok on the old 1gb fibre, then it gradually causes a network wide issue soon as you move this into a 10gb fibre to the Aruba ... Edited August 8, 2019 by MrWu
kmount Posted August 9, 2019 Posted August 9, 2019 Immediate first thought is spanning tree, are both the old and new trying to be the root? And thoughts around the l3 routing will follow next if the old netgear is getting smashed. Logs on both cores is where I'd be looking. 1
Davit2005 Posted August 9, 2019 Posted August 9, 2019 Spanning tree would be my first thought too. Have you got multiple connections between the edge switches and the 5412. Check the logs of each switch, as you lose connectivity you may need to do this via console to the core. This is typical behavior of spanning tree, broadcast storms can build up. Also make sure you have enabled at least MSTP on the core switch and the edge switches. 1
MrWu Posted August 9, 2019 Author Posted August 9, 2019 (edited) Thanks, only one link between old core and new core Will check.. but network has been solid and quiet last night and also this morning all smooth, will move the remaining servers and most of the edge links and isolate the old core as soon as I can today. Will migrate remaining edge links one by one, wait and analyse On the Arube 5412s and ARUBA 2540 edge switches, is it a simple matter of going into GUI and flicking Spanning tree to ON? I'm think of just turning Spanning tree on the last edge location that caused a storm, plus the Core switch to do a bit of trouble shooting when I plug it back in? PS currently looking at setup (I'm not familiar Aruba switching however) Spanning tree on Core is off and on the Netgear Core its ON Edited August 9, 2019 by MrWu
MrWu Posted August 10, 2019 Author Posted August 10, 2019 Further investigation, and looking at 3 separate incidents of the broadcast storm over 3 switch locations, it looks like they all have IP phones in that segment (it’s not on VLAN, we only have around than 10 IP phones in the whole school ) I do think when an IP phone (Siemens) gets disconnected and moved over to new HP 10 gb segment it causes issues. The phones are not daisy chained to a PC, just directly plugged into the switch. I have migrated edge locations for ICT rooms with no issues for example, and they had no ip phones.
geohanson Posted August 11, 2019 Posted August 11, 2019 We had a similar issue moving some of our core infrastructure from 1GB to 10GB. Just before I started at my current place, the previous NM had an external company swap out old 3Com 1GB core switches with new Allied Telesis 10GB switches. Config was just manually copied from the old switches to the new and nothing was really changed. Over the next few weeks (and months) we experienced more and more packet loss, dropping connections, slow logins and just general terrible network performance. After some investigating with Wireshark, it looked like our WiFi was to blame - sounds stupid but it was very obvious after looking at Wireshark! After reconfiguring our entire WiFi system the issues went away. For whatever reason, these issues had only shown themselves after we upgraded to 10GB and had never been an issue before. I still can’t quite get my head around it but it was very strange!! Might be worth trying wireshark and see if you can see any obvious issues. 1
steveg Posted August 11, 2019 Posted August 11, 2019 Is the fiber you have capable of 10gig reliably? How old are the runs, and what spec cable is it? 1
MrWu Posted August 12, 2019 Author Posted August 12, 2019 Thanks all again for suggestions I have turned on spanning tree on the core and have nearly migrated all edge to 10gb bar one The edge location throwing out CRC errors was sorted with a newer OM3 patch lead There was one switch that I isolated that was causing broadcast storm like instabilities to the network, check equipment plugged into it one by one tomorrow Overall, looking at the stats of the HP core, things are all green at the moment, no errors so will carry on monitoring ...
MrWu Posted August 15, 2019 Author Posted August 15, 2019 Turns out to be my Siemens IP phones (only certain models I believe) that causes a broadcast storm issue that switches can’t detect .. (as in it’s not a BPDU ?)
robk Posted August 16, 2019 Posted August 16, 2019 Turns out to be my Siemens IP phones (only certain models I believe) that causes a broadcast storm issue that switches can’t detect .. (as in it’s not a BPDU ?)Out of interest which Siemens phones are you using.. have a similar problem but not as accute. We do also use Siemens phones and its Poe switches that have issues....
MrWu Posted August 16, 2019 Author Posted August 16, 2019 (edited) Out of interest which Siemens phones are you using.. have a similar problem but not as accute. We do also use Siemens phones and its Poe switches that have issues.... Well I'm stumped.... I have siemens optipoint 410 Turns out that the cabinets that had issues when switches was replaced had these phones (which were plugged into POE ports on the old Netgears) were mistakenly plugged into a non POE HP switch...) Should not power up the phones right? Believe it or not, and I watched this happen, there is a flicker and occasional link on the switch port (I see it flash intermittently) there is no display on the IP phone itself but there's network activity somewhere and then it brings the segment of that network down, no spanning tree messages though. So is the Siemens phone's internal PC/Phone NIC forming a bridge of some sort, passively? At least now I have documented to my team and colour coded those cables ! Edited August 16, 2019 by MrWu 1
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now