Jump to content

Recommended Posts

Posted

Hi everybody,

 

Wonder if anybody has experienced this kind of thing:

 

In the last few months we recently had a company upgrade part of our aging HUB controlled cat5e network with cat 6 switches. We now use gigabit to the desktop which is great. All seemed ok for a bit...

 

Then we had some segments starting to lose some ports! This required resetting the switch in that area. This resolved the issue temporarily. Since that time all sorts of areas have been affected. We thought maybe some power fluctuations could affect the switches and indeed we have data via PowerChute which goes some way to confirm this so we have protected the most troublesome areas with SmartUPS units. However, this to our surprise has had little or no effect!

 

We regularly on almost a daily basis now have 3 key areas which exhibit this issue and randomly others.

 

We have a large network over a very spread out campus with many fibre segments going back to two central fibre switches!

 

We have had the network installers back to check for storming switches and transceivers etc and they apparently did find an issue although reluctant to share this with us at the time of discovery. Even with all this the issue still exists and is inconveniencing many staff and making the situation difficult for us in IT Support as we have:

 

1. Checked the network for storming switches and hubs ourselves (A massive undertaking)

2. Checked for damaged cables

3. Checked for damaged network cards

4. Checked for damaged storming WAPS, min switches etc.

 

The only thing we have identified is that the new MAIN switch in question an AT 9924SP seems to fluctuate very quickly indeed and is the main fibre core switch that the company installed along with a number of new AT switches (copper) around the campus. The existing MAIN switch which is identical we had before did not exhibit this behaviour. Only since the other switch has been installed does the whole lot exhibit this manic behaviour. And another thing, we have to reset the remote segments first thing in the morning! Why is this?

 

We are considering purchasing a spare AT 9924SP switch (Not cheap) and changing it out to test the idea that the new core switch may be at fault so....

 

The upshot of all this is:

 

Our network seems less reliable since we have upgraded to full Gigabit, the company responsible for the upgrade has already attended, corrected an issue and felt nothing else was wrong even though they accepted that the switch seemed overly busy, we have checked the other switches around the campus, all computers and printers and other devices for errors and as much cabling as we can and found nothing obvious except the main (apparently chattering) main switch.

 

So the questions are:

 

1. Has anybody else experienced this kind of thing, if so, what was the resolution?

2. Has anybody else had experience with this switch AT 9924SP and modules ATSPSX, mini gbics? What do think of Allied Telesis switches?

3. Does Power have an effect on switched networks in your experience in relation to brown outs and spikes? Do you need to UPS protect the switches for this issue and in which case even though they are smartUPS and protect against spikes and surges and over charge under charge why do they appear to not have an effect?

 

Many thanks and apologies for the long post.

 

I appreciate any comments and answers.

:D

Posted

Unfortunately we have too many connections to do that and not enough ports. Many thanks for the idea though and many thanks for looking and replying...

 

Can you put the fibre connections from the new switch into the old one and turn the new one off to see if the issue vanishes?

 

Wes

Posted

Hi

 

Have you got a loop and is tree spanning turned on. Tree spanning would cut off ports if a loop was detected to protect your network.

 

How did you check for damaged network cards?

Posted

1) Without more info, it could be a few things - if spanning tree is turned on, what's configured as the root? (should have the lowest bridge ID). What does Wireshark tell you if you listen on the main switch? What do the error counters look like on the switches - take a note of any errors, clear the counters and look at them again in 30 mins.

2) No idea, just have a couple of fibre converters.

3) Yes (IME). All our switches are on UPS (area with unreliable power). It's still the best money I've spent in terms of £ > reliability improvement. If you're on UPS, chances are your issues aren't power related.

Posted
The installers have not set up the switch so that we may access it. I will have to refer to the company who did the work before I start attaching laptops to the managment port. In terms of a loop, we have checked as much as we can for this and it doesn't seem to affect everything, just limited it seems to our two main fibre switches. Actually I have to agree the behaviour is indicative of a loop. I will contact the company to see if they have any problem me connecting a laptop to the managment port...thanks very much for your thoughts and we'll look for the loop etc...
Posted
hi, I'd try another switch to be honest this would be the best way to rule out your main switch and see if a loop occurs. I'd grab a cheap Cisco switch or a HP or something to try it - you could also keep it in your store cupboard as a spare.
Posted

Have you tried something like wireshark?

 

Do a bit of a packet sniff and see what the traffic is and where it's coming from? If there are lots of spanning tree changes you'll see it in WireShark by filtering for STP, MSTP, or RSTP packets.

 

Did the providing company leave you with any documentation?

Posted

Yep, we agree....this is what we think too....

 

hi, I'd try another switch to be honest this would be the best way to rule out your main switch and see if a loop occurs. I'd grab a cheap Cisco switch or a HP or something to try it - you could also keep it in your store cupboard as a spare.
Posted

I agree it would be interesting to logon to the switch with a remote station and try wireshark on it. Unfortunately the company did not leave any docs at all so we'll have to download it all and try this.

 

Many thanks everyone for the ideas and thought so far....

 

Have you tried something like wireshark?

 

Do a bit of a packet sniff and see what the traffic is and where it's coming from? If there are lots of spanning tree changes you'll see it in WireShark by filtering for STP, MSTP, or RSTP packets.

 

Did the providing company leave you with any documentation?

Posted (edited)

If the company who installed the infrastructure didn't document anything or provide you with details of how to access it, vent extreme displeasure in their general direction with a view to someone coming onsite PDQ to a) document their work and do a proper handover b) diagnose the problem.

 

Lack of documentation would also suggest that they didn't document it for their own reference (unless you have a maintenance agreement with them?).

Edited by pete
Posted

We have already. The excuse is, if we need more switches, they come and install and that's all. They insist on us taking out an expensive network management package with them. I think you are right. I think they should have documented it and also handed it over in a timely fashion. I have vented my concerns to little avail. We have ordered some transceivers and I have managed to download the manuals. I have configured a netbook with wireshark on it and I will attempt to connect this up to the switch to see what's going on. Later on in the week when (and if) we get the transceivers we will replace both up and downlinks to see what happens with that too. Severing the links will be interesting in itself I think....

 

Thanks so much for the help so far, invaluable guys....I'll keep you informed...

 

If the company who installed the infrastructure didn't document anything or provide you with details of how to access it, vent extreme displeasure in their general direction with a view to someone coming onsite PDQ to a) document their work and do a proper handover b) diagnose the problem.

 

Lack of documentation would also suggest that they didn't document it for their own reference (unless you have a maintenance agreement with them?).

Posted

We think we will get another switch. If you have tried these the Async0 port does not connect! I have tested it with my Fluke tester and it reckons it is a short. Is that a faulty switch in your opinion.

 

hi, I'd try another switch to be honest this would be the best way to rule out your main switch and see if a loop occurs. I'd grab a cheap Cisco switch or a HP or something to try it - you could also keep it in your store cupboard as a spare.
Posted

I think it is probably a combo of a weak switch and bad wireing. Some switches fail rather ungracefully when confronted with less than perfect wireing, cat5e that may have been fine for 100mbit can be unstable at 1GB especially if it is old or installed by primates. The better switches seem to handle this alright simply resetting the port or not failing out at all (Cisco, HP) where as the cheaper gear (D-Link) drops the port and leaves it disabled until the switch is rebooted.

 

We have some issues with this at one of our schools where some epicly dodgey wireing by an external contractor which I beleive must be run wrapped around a power feed from the cabinet all the way to the ports will drop with our dirty DGS-1248T switches but are fine on higher quality switches.

 

I don't know much about AT apart from the fact that they did used to be well respected a couple of decades ago and that our MoE has a bulk purchase deal with them which personally makes me very suspicious of their quality.

Posted

I would be changing that switch, It's the main core switch I would be changing it to something high end like Cisco, Nortel or HP. Leave all the other switches alone though and just change the core.

 

We where using Nortel 5590's for our absolute core switches but now all the core switches are Cisco 3550's, these are also the switches which connect our other sites.

 

YOu might be able to get a couple off ebay cheap if you don't have much if a budget.

Posted

Having laboured over this switch I have deduced that the ASYNC0 which should be active and has a static IP of 192.168.242.242 does not WORK!!! I have copied this from the hardware manual: Out-of-Band Ethernet Management Port

The out-of-band 10/100/1000 Mbps Ethernet port is dedicated to management

traffic on the AT-9900s switch. Use it for initial configuration and on-going

management tasks. The default IP address for the port is 192.168.242.242 to

allow remote access. This port is reserved for management only; the switch

does not transmit frames between this port and switch ports.

 

I have tested this with my Fluke tester and it comes back with an error and NO link light! Do I have to activate this port or something? There is nothing I can see in the instructions!

 

Any ideas anybody? At this rate I am not going to be able to link a laptop to it to test it and configure it or use wireshark....am I missing something or is this port faulty?

Posted

do your issues only occur at certain times of the day? If so it's unlikely to be a loop.

 

as previous posters have mentioned you can use wireshark to check if any clients are producing a lot of traffic epecially broadcast frames otherwise unplug things until the problem goes away.

 

as for management you should be able to telnet or SSH to the IP address of the switch and log in that way. The out of band management probably hasn't been setup.

Posted

I agree with all you have said. However these checks I have completed by unplugging everything and plugging it all back in until the switching becomes excessive. Our network suppliers did this too and maintain that when they were plugging in the segments as soon as any of the segments started to be plugged in again the switching started to be excessive which IMHO points to the switch. Also, you are correct the port has not been set up it seems. Having said this on trying to set it up I cannot even get a link light let alone anything else. It seems the port is dead! Any ideas on that one.

 

I appreciate the comments and support with this by the way and thanks for taking the time and effort to respond....

Posted

I hope I'm not being too patronising ;-), but... given what you have available to you (e.g. no documentation, no management of the switch and the fact that you have a very busy core switch) I'd approach this in a methodical and logical manner.

 

If you've got a hub and spoke network topology, I'd do this:

 

1. Start by unplugging all of the fibre uplinks - monitor the switch activity, has the problem gone away? If yes, go to step 3, if not, go to step 2.

2. Unplug all devices from the core switch and reconnect them one by one (with a few mins gap inbetween each) this way if you've got a device that has a faulty NIC, or is sending out junk the switch will have a chance to start displaying 'poorly symptoms' again.

3. Reconnect each remote cab one by one with a suitable gap inbetween each, until the problem is identified as existing again - then you can be reasonable sure the problem is associated with the most recently connected remote cab - trouble shoot that!

 

Hope this helps.

Posted

HI,

 

No your not being patronising at all and what you have suggested is a standartd procedure which the company in question has tried but I have not....

 

I will be trying this one out myself soon I fear. We have now some new transceivers to try too....

 

I'll keep in touch and many thanks for the help...

 

I hope I'm not being too patronising ;-), but... given what you have available to you (e.g. no documentation, no management of the switch and the fact that you have a very busy core switch) I'd approach this in a methodical and logical manner.

 

If you've got a hub and spoke network topology, I'd do this:

 

1. Start by unplugging all of the fibre uplinks - monitor the switch activity, has the problem gone away? If yes, go to step 3, if not, go to step 2.

2. Unplug all devices from the core switch and reconnect them one by one (with a few mins gap inbetween each) this way if you've got a device that has a faulty NIC, or is sending out junk the switch will have a chance to start displaying 'poorly symptoms' again.

3. Reconnect each remote cab one by one with a suitable gap inbetween each, until the problem is identified as existing again - then you can be reasonable sure the problem is associated with the most recently connected remote cab - trouble shoot that!

 

Hope this helps.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...