Jump to content

Recommended Posts

Posted
Following on from my comments in annoying thread, a message has gone out to colleagues, which I'm not entirely sure are fully accurate. However, a request has gone out to ensure a cable from a ethernet port isn't plugged into another. Is it really so easy to bring a network to it's knees? what defences protect against this?
Posted

Most switches (proper managed ones) have something like broadcast protection that send special packets out, if a port detects that packet it disables and cuts it off

 

But obviously have to config it not to be on uplinks :p

 

Steve

  • Thanks 2
Posted

Yes it really is. It is why having all patches live is not a great idea. I have had it caused by a dodgy NIC as well.

I'm not sure if it has changed but it used to be that spanning tree would limit the issue to a single switch, stopping it spreading. This allowed you to at least have half a chance to find the area the loopback was in.

  • Thanks 1
Posted
...I'm not sure if it has changed but it used to be that spanning tree would limit the issue to a single switch, stopping it spreading. This allowed you to at least have half a chance to find the area the loopback was in.

 

It's been a while now, but as I recall, spanning tree cause problems for (some) methods of station building/imaging.

Posted
Most switches (proper managed ones) have something like broadcast protection that send special packets out, if a port detects that packet it disables and cuts it off

 

But obviously have to config it not to be on uplinks :p

 

Steve

That makes sense. No idea how ours are configured - I know they are Procurve but no idea on model. We have a quote for replacement but it's clear there's lot of legacy, lots of lack of info and no one with good knowledge. It makes sense now that removing one of the cables prevent a switch to switch broadcast but how that happened remains a mystery.

Posted
...It makes sense now that removing one of the cables prevent a switch to switch broadcast but how that happened remains a mystery.

 

We suffered a broadcast storm when one of the support team, while tidying up after a particularly raucous and destructive class had left a network room, plugged a cable into an outlet (thinking he was helping) and created a network loop.

Posted

We nailed this down by having STP enabled, but making sure edge ports are configured with PortFast (so they come up as quickly as clients need) and BPDU guard so they shutdown if a BPDU packet is seen arriving. On the oldest switches of questionable spanning tree reliability edge ports have 'loop detect/protect' enabled instead which provides similar functionality. We haven't had a issue since*

 

 

*Well except when our Aruba IAPs very occasionally slip into mesh mode and suddenly all sorts of absolute insanity ensues, because I haven't figured out the correct settings for ports connecting to IAPs, and they bridge to other IAPs and thus create looped trunk ports. LOL.

Posted
All our switches have D-Link's LoopBack technology enabled on all ports. IN the even of any loopback anywhere in the network, it sends a syslog packet to our syslog server and emails me with switch , and the 2 ports looped. both ports will be disabled stopping the loop, it will check at x interval I set, if loop still exists will continue blocking the ports, if resolved, both ports are re-enabled. Easy way to nail down the culprits.
Posted

I’ve had this in the past. Luckily the week after I spent half a day re patching a cab - so it was easier to find from that end.

 

Our unifi switches manage this protection brilliantly tho. Not had an issue for years.

Posted

I've seen this once or twice. Student plugged network cable between 2 ports.

 

If you split up the network into vlans it segregates and localizes the network issues caused by loops. As well as employing loop protection, spanning tree etc.

Posted
This rumbles on - I've asked our current provider to let me know if STP is enabled on our network. They've replied they can't tell, as the previous provider won't provider router admin access (they retain the contract for Internet provision, but not for much longer if I get my way). Isn't STP enabled on the switch? Or would router access be required? The onboarding process indicates switch ip addresses were provided.
Posted

They should be able to tell you from the switches, my guess is they don't understand STP.

 

The Core switch(es) should be configured as the route bridge by setting the priorities to be the lowest & that's generally about it for most switches.

 

In the context of what we're discussing the router is irrelevant.

Posted
They should be able to tell you from the switches, my guess is they don't understand STP.

 

The Core switch(es) should be configured as the route bridge by setting the priorities to be the lowest & that's generally about it for most switches.

 

In the context of what we're discussing the router is irrelevant.

That's what I thought. I think first line support is the where the messaging is getting confused. I think there is expertise within the company but haven't got to it yet. We have a face-to-face next week with a technician so hopefully we can make progress. Bottom line is our network is in a pickle, e.g. one network name that is supposed to be 'guest' wifi only is showing as a wired connection. We have a number of room that have double connection faceplates. As far as I can determine all double connector ports, only one of the two work in each room. I have a suspicion it might be related to a redundant Avaya IP installation which is still in a cabinet but no longer used as we switch to teams voice with very few phone handsets. It all needs a thorough re-organisations and I'll be making that point at the face-to-face meeting.

  • 1 month later...
Posted

There were reports of the WiFi not working again yesterday, but cabled connections OK.

 

I happened to need to be in the office to meet a chap who is going to quote for a cabled structured cabling test. It appears to have got worse and deteriorated whilst I was there.

 

Initially I could get cable connection but not internet. Then no cable connection to the point Windows says 'are you plugged in'. It would appear we had about four ports working - that's only 10%. The Windows network check seemed to go with 'No valid IP address' could be gained. Both for cabled and WiFi. As it is half term, the impact is less severe but as our new IT support people are supposed to be our single point of contact, we raise with them. Because the internet line and router is run by our old IT provider they speak to them. As far as they are concerned there is no issue - and indeed plugging in to a spare port on the router works fine. This is severity 1, this is the second time it has happened, and my patience is wearing thin! I've been badging line manager for months, but his LM seems to be stalling. By chance, the CEO was in, and I made it clear what the issue is. I would appear he is not being fully briefed. He is now and isn't shy with putting the money forward to replace the ISP and router right away.

Posted
Admittedly not my speciality, but if you get no problems when you plug into the router it suggests no issues on the internet line or the router.
  • Thanks 1
Posted (edited)
Admittedly not my speciality, but if you get no problems when you plug into the router it suggests no issues on the internet line or the router.

 

Not my speciality either, but you are spot on and confirms the view I gave yesterday that it has to be downstream.

 

As a result of me escalating yesterday to get the issue on the highest priority, 3rd line support got a colleague today to email a picture of our, and I use the term loosely, 'cabinet'. They then point out of the two switches we have, only one has lights on! Replacement with a known good power cable proves the switch to be dead. My colleague then spends 3 hours on the phone re-cabling blindly following instructions. What really riles me is if they had someone on-site on Monday, the immediate issue could have been resolved in a fraction of the time. This is one of those situations where expertise and being physically present is a requirement. To say I am far from happy with the support currently would be an understatement.

Edited by Ditto
  • 2 weeks later...
Posted
Cable testing quote has come in at about £1.2k - thought to be a couple of days work. Probably 2 switches (another died last week!), 2 patch panels, 3 APs and just under 50 wall connections. Proper Fluke testing - I presume with certification, but checking. It's one of the more smaller quotes we've seen from our IT supplier - what do you think?
  • 2 weeks later...
Posted (edited)
We nailed this down by having STP enabled, but making sure edge ports are configured with PortFast (so they come up as quickly as clients need) and BPDU guard so they shutdown if a BPDU packet is seen arriving.

 

This is the way. BPDU guard is key. On a properly configured switch it’s not easy to create a broadcast storm.

 

 

But reading the post I am not sure that this is a broadcast storm.

 

I won’t repeat other suggestions, but one consideration. On the op you mentioned this issue happens when you connect one cable. Is there a chance that there is something on the network with a duplicate IP of your default gateway, layer 3 switch or router? It could be that plugging in that cable connects that device causing the issues.

Edited by FN-GM
  • Thanks 2
Posted

On the other thread it was reported

"A cable in the switch cabinet when has been identified as a problem bringing the internet to a crashing halt if left plugged in both ends - suggestions of a broadcast storm. However, the number it has on it can not be traced to anywhere on site and they are all labelled."
making it sound like a network loop. Also if nobody knows where the other end goes, just leave it unplugged?

 

However later on (back in this thread) other symptoms are reported, and I'd have put a small amount of money on @simpsonj having zero'd in on at least one part of the problem

One thing that struck me whilst reading this, have you checked your DHCP scopes aren't running out of IP addresses?
but as @FN-GM says, if there are two devices with the default gateway IP, then that would also explain shenanigans.

 

That all said, I would hope that the service provider has fully resolved the issue by now. I'm sure several of us would be interested to the actual root cause.

Posted

You can trace where a loop occurs by logging onto the switch and checking which 2 ports have massively higher packets than the others. That said it can be very difficult to log onto the switches in the first place when a loop has occurred. Prior to STP switches, we had a few loops usually caused by kids messing around (unplugging Ethernet to dodge RM tutor, then plugging them in wrong), since STP we haven't had a problem.

Yes I was shocked when I found out too, it was that easy to bring the entire network down, luckily students never figured it out or we'd have had it every single week.

  • Thanks 1
Posted
One thing that struck me whilst reading this, have you checked your DHCP scopes aren't running out of IP addresses?

I'm going to show have much I am not a Network person here! As I understand it, we have a router which may have minimal firewall functionality, but that than connects through to a switch/patch panel. From there, cabling goes to 2 other switches supporting up to 40 network point and supporting 2 or 3 APs. In terms of a DHCP scopes, I wouldn't know how to check that but we did have the following statement from our 'technical consultant' who visits once a quarter. 'There are delays in the devices receiving IP addresses from the Draytek router and sometimes they don't receive an IP address at all.

 

Since then, one of our switches died (wouldn't power up) and a colleague was walked through pulling/inserting cables on the 2nd switch/ patch - this was a 'technically blind' process, the colleague has no idea about what he was doing, just asking as a bizarre 'remote' service! I have made my dissatisfaction about that and that they should have come on site - they're only 30 minutes away. Now, as it happened I was in the office this week along with several others, using the boardroom. This room has never had a good reputation for WiFi, and recently just our CEO and a major donor were in and they couldn't get connected to the internet and I believe it was the day the switch popped it's clogs. I and everyone else in the room got perfectly good Wifi, strong and stable. They fact this was after the switch died and re-cabling, I doubt it is a coincidence. It's not all good though, a 2nd half of the building is without any WiFi whatsoever.

Posted
That all said, I would hope that the service provider has fully resolved the issue by now. I'm sure several of us would be interested to the actual root cause.

I would have wished for the same, but to be frank, they've been utterly useless with our onsite infrastructure. All they've done is provide hugely expensive quotes with new equipment, cable tests and excessive project management charges. Throughout, they've not provided a single credible shred of evidence of the underline causes of our problems other than the IP related comment in my last post. for balance, remote 1st line support is excellent - so if you've got a PC or MS software related issue, they're great. But for our tiny back-end core infrastructure, it feels a though we have been utterly abandoned and our only way forward is via the charities bank balance.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...