Chaps,
I wonder if you can help?
Got a really odd problem that started about two weeks ago, just wonder if any fellow Procurve users have ever seen this... Really struggling to find the answer...
We have three Core switches (2x5406zl and a 5308xl) which are not fully populated, about three or four module slots in use in each, and none fully populated. Each connection (mix of fibre and copper) goes to another managed Procurve switch somewhere on the campus (mixed later 25 and 26-series - all the 2524s have been replaced!!) All switches are linear in arrangement, i.e. no redundancy loops for STP to manage. Almost all uplinks (fibre and copper) are gigabit.
We have a vLAN for each subnet, not ACL'd. There is a two-port trunk between the 5406's, and a single trunk from the 5308 into one of the 5406's... All vLANs come back to the Core Switches and the firewall is D/G for the core switches.
two weeks ago, for some unknown reason, our monitoring system started to get sporadic, random ping failures of the switches. This lasts for a minute or two then returns to normal. It is affecting switches across all the vLANs, normally 2-4 in any one instance but sometimes (though never more than) 7-10 (we monitor about 60 switches).
I have tested the ping from other subnets to confirm it's not a problem with the monitoring system or the subnet it is on (and tried turning off the monitoring system and doing pings of switches just to rule it out) and they fail also. You also cannot connect to the management interface via telnet/SSH on affected switches at that time.
However: the switches are not down. Traffic is passing through quite happily while , internet, local system access, IP phones the lot all keep working! So it can't be a blocked port...
The ports are not overly busy (visually inspected and port counters checked), the CPUs are doing no more than 20% at any time, no drops or bad packets anywhere.
Event logs do not reveal anything seemingly relevant, apart from these which we have seen on one of our 5406's but nowhere else:
00828 lldp: PVID mismatch on port 35(VID 2)with peer device port 51(VID 1)(396728)
Any insight would be gratefully received
Thanks
Richard