PLSMC Posted May 3, 2024 Posted May 3, 2024 I am at a split site school and at one of the sites we are encountering an issue whereby various devices, be it desktop and/or VoIP phone will not connect to the network. On the desktops, running the network troubleshooter, the error shows as unable to contact the DNS server. The desktop does have all network settings correct( IP, gateway, dns, etc) when doing the IPconfig all command. This happens to just a few devices randomly and might last a few minutes or hours or sometime days. No matter which network port i connect a device experiencing this issue it will not connect, however if i connect a working device into the port of a non working device it connects without issue. A related issue is that i do not seem to be able to connect from a device in vlan 1 to vlan 3. Once thing i have noticed is that it only seems to occur during term time. During school break all devices are able to connect and stay connected without any issue. I am also able to connect without issue to all vlans from the other site at all times. The network hardware is comprised of the following. HPE 1950 Core switch which is used to route traffic across vlans. The edge switches are TP-Link. We have 3 vlans which comprise of: computer devices, IP Phones and Paxton NET 2 doors. We have 1 physical server ( domain controller) and 1 VM for Paxton Net 2 which sits on vlan 3. I have replaced all patch leads and checked switch configs to ensure there are no loops but now i am running out of ideas. I imagine it is a switch config issue but i just cannot find it. The next step i will be taking is to disconnect all switches except the core to try and isolate where the issue might be and hopefully by adding a switch at a time the issue will present itself. Can anyone offer other suggestions on what might be the cause or where to start troubleshooting. Many thanks.
simpsonj Posted May 3, 2024 Posted May 3, 2024 Have you checked your DHCP scope to see if it's running out of addresses?
PLSMC Posted May 3, 2024 Author Posted May 3, 2024 Only 50% in use. - - - Updated - - - does the voip phone get power just no ip? Phone does power up.
pablo007 Posted May 3, 2024 Posted May 3, 2024 (edited) Its hard to see it as a switch config issue given that a working pc will work in any port and you get power to phones, are any of the pcs that fail plugged into the voip phones(piggy backing)? if you set the ips to static do they work? whcih side of the site is the dhcp server the working or non working? Edited May 3, 2024 by pablo007
psydii Posted May 3, 2024 Posted May 3, 2024 I'm also suspecting a DHCP issue, probably not enough IP addresses available in the pool. Butipconfig is showing everthing is ok... so it can't be. Lets start at the vlans.. Traffic can't pass between vlans without some special configuration. That is the purpose of the vlan concept. Typically you have an IP network (subnet or CIDR) allocated to each VLAN. There is usually a router (typically the core switch) that has an IP interface in each of the VLANs and the router handles passing packets between the subnets (check the core switch, does it have IP interfaces in each VLAN/Network, is routing enabled) Typically you don't want all traffic to be able to pass between subnets (do you really need students to be able to ping Paxton?). So on the router there may also be Access Control Lists which limit traffic flows (you might consider these to be like firewall rules, though they typically are not as flexible as rules on a dedicated firewall). Are there such rules in play here? Also there need to be a DHCP / UDP Helper / Forwarder /Proxy that takes DHCP request on one VLAN and passes them to the server specified in the router config. Is this present? Does it point to the expected DHCP Server? If the DC is on a different vlan? What it the output of ipconfig /all on the Domain Controller? (in vlan2???) What is the output of ipconfig /all on a typical PC (Vlan 1?) What is the output of ipconfig of a device on vlan 3 Also... as you suggest intermittent faults could be loop related. Does your core have spanning tree enabled? If so when does it show it last altered/recalculated the topology?
PLSMC Posted May 3, 2024 Author Posted May 3, 2024 Its hard to see it as a switch config issue given that a working pc will work in any port and you get power to phones, are any of the pcs that fail plugged into the voip phones(piggy backing)? if you set the ips to static do they work? whcih side of the site is the dhcp server the working or non working? I haven't checked with static IP and DHCP server is on site with issue. Affected desktops are plugged into their own network port.
PLSMC Posted May 3, 2024 Author Posted May 3, 2024 I'm also suspecting a DHCP issue, probably not enough IP addresses available in the pool. Butipconfig is showing everthing is ok... so it can't be. Lets start at the vlans.. Traffic can't pass between vlans without some special configuration. That is the purpose of the vlan concept. Typically you have an IP network (subnet or CIDR) allocated to each VLAN. There is usually a router (typically the core switch) that has an IP interface in each of the VLANs and the router handles passing packets between the subnets (check the core switch, does it have IP interfaces in each VLAN/Network, is routing enabled) Typically you don't want all traffic to be able to pass between subnets (do you really need students to be able to ping Paxton?). So on the router there may also be Access Control Lists which limit traffic flows (you might consider these to be like firewall rules, though they typically are not as flexible as rules on a dedicated firewall). Are there such rules in play here? Also there need to be a DHCP / UDP Helper / Forwarder /Proxy that takes DHCP request on one VLAN and passes them to the server specified in the router config. Is this present? Does it point to the expected DHCP Server? If the DC is on a different vlan? What it the output of ipconfig /all on the Domain Controller? (in vlan2???) What is the output of ipconfig /all on a typical PC (Vlan 1?) What is the output of ipconfig of a device on vlan 3 Also... as you suggest intermittent faults could be loop related. Does your core have spanning tree enabled? If so when does it show it last altered/recalculated the topology? I have attached images of the switch config relating to routing and IP helper just to confirm this is correct. I cannot see any ACL config enabled. The DC is on VLAN 1 which is the same vlan as the affected desktops but different to affected phones ( vlan 3) The output of ipconfig /all on devices on different vlans appear the same apart from the gateway entry which differs as each vlan has its own gateway IP which routes back to the core and obviously the device IP. STP is enabled MSTP. I need to look into where to find when it was last recalculated.
synaesthesia Posted May 3, 2024 Posted May 3, 2024 Am I missing something, or is the HPE 1950 a "managed smart switch" - therefore it's only capable of static routing and that screenshot isn't related to that. Are there any settings for RIP? They're barely suitable as baseline edge switches let alone a core. Do you not have a proper router/firewall in place for instance from your ISP?
PLSMC Posted May 3, 2024 Author Posted May 3, 2024 Am I missing something, or is the HPE 1950 a "managed smart switch" - therefore it's only capable of static routing and that screenshot isn't related to that. Are there any settings for RIP? They're barely suitable as baseline edge switches let alone a core. Do you not have a proper router/firewall in place for instance from your ISP? Unfortunately this is the switch and what i have to work with. I cannot find anything in the GUI relating to RIP configuration ( probably has to be done through CLI) but it does appear in the statistics as per attachment. We do have a proper router/firewall from our ISP ( LGfL) I should mention that the switch logs don't contain any error messages, however it does continuously report that the "maximum number of ARP entries on the device has been reached". Is this something i should look into. It only appears as an informational log entry. Maximum number of ARP entries on the device is reached.Maximum number of ARP entries on the device is reached.maximumMaximum number of ARP entries on the device is reached.
psydii Posted May 3, 2024 Posted May 3, 2024 "maximum number of ARP entries on the device has been reached" That's quite a surprise given the LAN ip ranges only appears to support 256*4 Hosts, I vaguely recall that the 1910 (earlier version of this class of switch) could support 4000, or was it 2000 hosts? Anyway that's way more than you likely have devices connected. A quick skim of the latest firmware release notes for this model of switch suggests that older versions (pre 2019) had lots of arp related bugs. If you aren't on the latest firmware release, try updating. The 19x0 series of switches are/were *amazing* pieces of kit. If you had a small lan and weren't trying throw multiple gigabits per second of sustained traffice across mulitple ports, they really were extremely good value. Particularly if you know the secret incantation to bring up the full Comware CLI. If you had a bigger lan, then they made perfectly fine edge switches., and (at least for a time) they came with a procurve-like warranty. Looking at the fragment of config you've been able to share, is it possible that some devices have the core switch as gateway and others have the lgfl router (180.1)? You would get extremely inconsistent experiece with devices trying to speak from VLAN 2/3 to the DC (in VLAN1) if, say, the vlan 3 device had the router as its default gateway, but the DNS Server (your DC) had 180.1 as its gateway. The DHCP relays seem to be configured correctly, assuming your DC/DNS/DHCP server has a single network interface and it is configured as 10.251.180.3 / 255.255.252.0 gw 10.251.180.210 I would need to run this past someone familiar with the lgfl router configs (Its friday evening, so I shan't @Mention him just now)... but this is how *i* *think* the end points should be configured Assuming LGFL is presenting 3 networks 10.251.180.0/23 10.251.182.0/24 10.251.183.0/24 There should be three DHCP scopes on your DHCP server: Vlan 1 devices Gateway 10.251.180.210 subnet mask 255.255.252.0 dns 10.251.180.3 vlan 2 devices gatweay 10.251.182.210 subnet mask 255.255.255.0 dns 10.251.180.3 Vlan 3 devices gateway 10.251.183.210 subnet mask 255.255.255.0 dns 10.251.180.3 This would/should allow all devices on all vlans to contact the DC/DNS/DHCP service on 180.3..... but things are going to get messy for traffic destined for the internet. That traffic will hit the core switch xxx.210 and then be routed to the lgfl gateway in vlan 1 (10.251.180.1), however returning/inbound traffic to that host will (probably) egress the lfgl router on the ip interface corresponding to the vlan of the host and bypass the core switch. So I'm not sure this is the best way to have things set up.. but I can't see how else hosts in vlans 2 and 3 would be able to communicate with the DC/DNS/DHCP server, unless it was multihomed (a trick I haven't done since about 2003, and I think you aren't supposed to multihome DCs) or instead the LGfL router was passing the traffic between your subnets. Hopefully though there is something in the above that helps you hone in on the problem and find a solution.
PLSMC Posted May 6, 2024 Author Posted May 6, 2024 Your description of our network is exactly how it is. The DHCP scope and vlan config etc, you got it spot on. The Core switch is running Software Version 7.1.070, Release 3507 which is the previous version so i will update as soon as possible. The strange thing is that the issue does not occur during the school holidays. It only seems to occur when the network is under load. I came in today (Bank Holiday) and all reported affected devices are connected to the network. It is becoming a real head scratcher. Many thanks to everyone for your input.
psydii Posted May 7, 2024 Posted May 7, 2024 … how is the lgfl router connected to the core? How are the vlans on these physical port(s) on the core switch configured? The fact this happens only when the network is busy points to arp or dhcp lease exhaustion. I’m wondering if some/any/all devices have the same physical interfaces configured across multiple vlans. How many devices do you typically see during a school day? How long are your dhcp leases? It still could be just a loop that the network copes with when idle, but struggles when ‘busy’.
PLSMC Posted May 7, 2024 Author Posted May 7, 2024 … how is the lgfl router connected to the core? How are the vlans on these physical port(s) on the core switch configured? The router is connected to Port 1 of the Core switch via the onsite firewall managed by LGFL. I have attached 2 images for how the vlans are configured for an endpoint device and for a trunk to another switch. The fact this happens only when the network is busy points to arp or dhcp lease exhaustion. I’m wondering if some/any/all devices have the same physical interfaces configured across multiple vlans. I am currently checking every port on every switch to ensure vlan assignment is correct. How many devices do you typically see during a school day? How long are your dhcp leases? Vlan 1 228 devices - Lease 8 hours - desktops, laptops, etc. Vlan 2 40 devices - Lease 1 day - IP Phones Vlan 3 145 devices - Lease 5 days - Paxton Door controllers It still could be just a loop that the network copes with when idle, but struggles when ‘busy’. I have replaced every patch lead in our comms cabinets to ensure that there are no physical loops.
Davit2005 Posted May 7, 2024 Posted May 7, 2024 I have replaced every patch lead in our comms cabinets to ensure that there are no physical loops.[ATTACH=CONFIG]71464[/ATTACH][ATTACH=CONFIG]71465[/ATTACH] What do the switch CPUs look like? Loops are most likely at the wall port end most of the time. I've only ever seen one (student had connected wall port to wall port) thankfully when it happened it only effected one vlan as we had multiple client vlans. When it happens normally all the activity lights on the switches flash really quick altogether this happened with us as all 3 switches in the stack were effected.
Chris_Cook Posted May 7, 2024 Posted May 7, 2024 You need to isolate whether its DHCP, DNS or routing. What's the Lease time on your DHCP Scopes? I'd set some devices with static IPs (and a suitable DHCP reservation) and see if you can still ping from/to them when its going wrong. If that works then its DHCP, otherwise its something else. I'd also look at the load on your DC. Is it possible that this is overloaded? It might be that the network is fine, but your DC isn't able to respond to DNS requests. You could use nslookup to test this, and compare responses from your server and LGFL or Googles DNS Servers. It might be worth testing nslookup on the DC and on a client PC on each VLan. Some baseline data from when the network is ok would be helpful for comparison. Wireshark's also a handy tool to see what's going on on the wire.
psydii Posted May 7, 2024 Posted May 7, 2024 (edited) ..so in my experience, each lgfl school ip subnet is presented via separate ports on their router, so only having one port connected between the LGfL router and your core makes me think either the lgfl config is not as I believe they are, or things are not as they should be.* I can't see from your screen shots which port on 10.251.180.0/23 your core switch is connected to the lgfl router, nor which vlans are presented on those port(s) on the core switch. Can you try to attach somethign that shows this? From the additional info you have presented, I will highlight/query the number of devices on the paxton vlan. This seems high given the number of end user devices you appear to have. (unless you have most doors in the building on the paxton system?) Do you have a wireless network that presents a separate IP range to wireless clients? If so are there a lot of devices on that? *So from experience, each lgfl subnet is on a separate port on the lgfl router. On the school core switch you should have 1 port and 1 vlan per LGfL subnet connected between the school core and the lgfl router. e.g. LGFL Port 1 (10.251.180.1) -> Core Switch Port 1 (VLAN 1 untagged, 10.251.180.210) LGfL Port 2 (10.251.182.1) -> Core Switch Port 2 (VLAN 2 untagged, 10.251.182.210) LGfL Port 3 (10.251.183.1) -> Core Switch Port 3 (VLAN 3 untagged, 10.251.183.210) (Core Switch port numbers are examples, since I can't see what you've got plugged where. Also the .210 ip addresses might be associated with the VLANs themselves rather than the specific interface, its been over ten years since I last looked at comware) Again I should stress that this is how I have done it, but I know LGfL have to deal with a lot of variations at schools, so they may have a different config at your site to enable it all to flow through one port, for example VLANs or routing, or maybe just all IP interfaces presented on that port?. I would have thought one port per subnet/vlan is the best way of handling this though. Can you show us the configuration of the port / vlan for on the core for the connection to LGfL? Also the potential elephant in the room is your opening sentance "I am at a split site school and at one of the sites....". How are the sites connected? Are these vlans/subnets present at both sites? Is/are their only one DC/DHCP/DNS server? Or is there one or more at both sites? Is there an equivalent 'core' switch at the other site? Edited May 7, 2024 by psydii
PLSMC Posted May 7, 2024 Author Posted May 7, 2024 What do the switch CPUs look like? Loops are most likely at the wall port end most of the time. I've only ever seen one (student had connected wall port to wall port) thankfully when it happened it only effected one vlan as we had multiple client vlans. When it happens normally all the activity lights on the switches flash really quick altogether this happened with us as all 3 switches in the stack were effected. CPU loads appear fine. See attachments for examples. We are a SEND school so whilst i will not rule out a student messing with the wall ports, it is very unlikely due to their condition. In my case it would be more likely to be a teacher.
pablo007 Posted May 7, 2024 Posted May 7, 2024 I haven't checked with static IP and DHCP server is on site with issue. Its time to do this now set some of the affected machines with statics and see how you get on
PLSMC Posted May 7, 2024 Author Posted May 7, 2024 ..so in my experience, each lgfl school ip subnet is presented via separate ports on their router, so only having one port connected between the LGfL router and your core makes me think either the lgfl config is not as I believe they are, or things are not as they should be.* We have our full IP range presented at a single port on the router and the Vlans are configured locally. LGFL have no visibility of the vlans. I can't see from your screen shots which port on 10.251.180.0/23 your core switch is connected to the lgfl router, nor which vlans are presented on those port(s) on the core switch. Can you try to attach somethign that shows this? From the additional info you have presented, I will highlight/query the number of devices on the paxton vlan. This seems high given the number of end user devices you appear to have. (unless you have most doors in the building on the paxton system?) We do indeed have the whole school from outer gates to all internal doors on the paxton system hence the high number of devices. Do you have a wireless network that presents a separate IP range to wireless clients? If so are there a lot of devices on that? We do but have a very limited number of devices and are lumped on vlan 1 *So from experience, each lgfl subnet is on a separate port on the lgfl router. On the school core switch you should have 1 port and 1 vlan per LGfL subnet connected between the school core and the lgfl router. e.g. LGFL Port 1 (10.251.180.1) -> Core Switch Port 1 (VLAN 1 untagged, 10.251.180.210) LGfL Port 2 (10.251.182.1) -> Core Switch Port 2 (VLAN 2 untagged, 10.251.182.210) LGfL Port 3 (10.251.183.1) -> Core Switch Port 3 (VLAN 3 untagged, 10.251.183.210) (Core Switch port numbers are examples, since I can't see what you've got plugged where. Also the .210 ip addresses might be associated with the VLANs themselves rather than the specific interface, its been over ten years since I last looked at comware) Again I should stress that this is how I have done it, but I know LGfL have to deal with a lot of variations at schools, so they may have a different config at your site to enable it all to flow through one port, for example VLANs or routing, or maybe just all IP interfaces presented on that port?. I would have thought one port per subnet/vlan is the best way of handling this though. All IP range presented on one port Can you show us the configuration of the port / vlan for on the core for the connection to LGfL? See attachment. Also the potential elephant in the room is your opening sentance "I am at a split site school and at one of the sites....". How are the sites connected? Are these vlans/subnets present at both sites? Is/are their only one DC/DHCP/DNS server? Or is there one or more at both sites? Is there an equivalent 'core' switch at the other site? There is a firewall rule for routing traffic between sites. See attachment. Each site has its own DC/DNS/DHCP server. There is an equivalent Core switch at the other site, only caveat is that other site is much smaller, only 70 devices in total.
PLSMC Posted May 7, 2024 Author Posted May 7, 2024 I will be going to the other site tomorrow to do the suggested static IP test.
PotNoodleTech Posted May 8, 2024 Posted May 8, 2024 1950 is a pretty old switch are you sure it's not the switch slowly going gaga? I'd also be looking beady eyed at the connection between the schools based on the fact it works fine in holidays but not in term time it could be the internet (i am assuming) link between the schools struggling with ping or bandwidth?
psydii Posted May 8, 2024 Posted May 8, 2024 Somewhere earlier in the thread I wrote "Assuming LGFL is presenting 3 networks" to which I expanded to mean on three separate ethernet ports. Subsequently you have corrected my assumption, stating the lgfl<->coreswitch is a single cable between a port on the lgfl side and your core switch. Do LGfL know about the subnets you have chopped up from the your IP allocation? i.e. does the lgfl router know that 10.251.180.210 is to be used as the gateway for 10.25.182.0/24 and 10.251.184.0/24? or was I also wrong about the default gateways set for each DHCP scope/VLAN? Also I note you appear to have jumbo frames enabled. Every time I've looked at this the consensus view is that this is not a good idea (except maybe between servers that only talk to each other e.g. vmotion, veeam backup, iscsi). I would be very tempted to try turning this off, and check that your client devices are not trying to send jumbo frames either! (Get-NetAdapter |Get-NetAdapterAdvancedProperty |Where-Object {$_."DisplayName" -eq "Jumbo Packet"}) But I am not sure how this could explain the symptoms you describe. That all being said, the ARP Cache error seems to me to be the key to unlocking this. As PotNoodleTech suggests, it might even be a fault in the switch itself.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now