Jump to content

Random issue - DHCP clients receive IP address but can't ping anything?


Recommended Posts

Posted

Hopes have been dashed :( We are still having issues this morning.

 

I had a laptop in science which would pick up the IP address but then would only let me ping the core switch (default gateway). I can ping each default gateway address that is stored on the switch but I can't ping beyond that?? Any ideas??

 

I had a few instances of this last week. I'm wondering if what I set up for the Aruba Airgroup is causing issues somewhere? I don't think we had this exact issue at the start of January, but it was a couple of weeks in when I enabled the airgroup settings....

Posted

So I've undone my settings that enabled the AirGroup feature on the WiFi, so will see what happens now. Its possible I suppose that the extra VLAN tagging required was causing some chaos somewhere. If anybody else can weigh in with some thoughts, I'd appreciate it!

 

Cheers

Posted

Any reason why you are pointing internal servers forwarders to Smoothwall and not directly to Google?

 

I always have set internal DNS servers forwarders to point out directly to google to be honest, firewalls and such point to the Internal DNS servers.

Posted (edited)

That's how it was configured when I started here and upon checking, seems to be a common way of setting it up:

 

http://www.edugeek.net/forums/internet-related-filtering-firewall/168964-smoothwall-ad-dns-not-resolving-hostnames.html

 

 

EDIT

 

I've just been in the room where we had a lot of issues period 2 today.... laptops. With my latest changes in place I picked 5 laptops at random, positioned them in random locations around the room and they all worked first time!

 

Fingers crossed

Edited by themightymrp
Posted (edited)

Lets completely eliminate DHCP and DNS from the equation for a minute. Next PC to have the issue (that you can use for a bit at least) give the pc a static IP in it's correct subnet but no DNS. Ping IP addresses of devices you know should respond (no firewalls) and then also try and ping the effected PCs fixed IP from those devices.

 

If giving a fixed ip and no DNS does seem to work, add your DNS servers back in manually and see where that gets you. Can you ping the names of some devices which should respond. Can those devices ping the client by name?

 

If the problem still persists then it looks like a vlan issue. Tagging or routing is still amiss somewhere. Or possibly a mac address table issue?

 

It might be time to crack open wireshark and take a close look at the traffic. You can try using it on the effected machine and on the machine which you are trying to ping. Otherwise you'll need to look into mirroring a port or a vlan to another port but that might add complication which may throw up other errors. Grab the portable version and shove it on a memory stick now so you run it on the next client to fail... https://www.wireshark.org/download.html Try ping the router, then a device on another network. Compare the packets being sent against each other and then look at the packets from the other end.

 

 

[edit] I must be a slow typer, some additional posts came in while I was typing this.

Edited by IrritableTech
  • Thanks 1
Posted

I'll do just that if any more issues occur. The more I think about it, the more I suspect that something I did to enable the AirGroup feature on the Aruba could have caused the problems. It involves tagging the VLAN of the local PC's onto the same ports as the Aruba AP's plug into. Aruba then filters the bonjour packets and routes them across VLAN's in a way not normally supported. If I've messed up on which VLAN's I should have tagged and where, I could have potentially messed up the IP ranges or address tables?

 

Not sure if that even makes sense but combine that with the multiple routes in the Smoothwall and 'routing issue' does spring to mind.

 

I'll let you know...

Posted

Well..... interesting morning. Definitely fewer cases today, only about 3 all morning. But the fix for them seems much easier. I just do an ipconfig /release and /renew and that's it. What I'm unsure of is if those machines have been switched on continuously since before I implemented the change. If they were already in a fail state then they could have just needed the refresh to pick up the network again.

 

I'll keep an eye out. I'll wireshark the next case if I get one - not that I'm great with reading the output though lol

  • Thanks 1
Posted

So it looks like I'll be using wireshark today. I'll try and do it on a machine which isn't desperately needing to be used which, typically, the faults tend to occur on :(

 

So I've had a spattering of cases this morning. On one, doing a release/renew worked. The others it didn't. For the others a simple reboot worked.

 

They all pick up an IP address from DHCP nice and quick but then I get one of 2 scenarios:

 

1) They can't ping anything including their default gateway

2) They can ping their default gateway but nothing else at all - not on their own VLAN or on another VLAN

 

It makes no difference if I put in a static IP with or without DNS server settings.

 

I'm so confused by this? There's no pattern to it, it only effects a handful of PC's each day. It doesn't take out chunks of PC's in ICT suites - it tends to be a single one, or maybe two. Once it finds the network again i.e. from a reboot, the machine works fine. Everything is quick and no issues with access.

 

Before I delve into wireshark and get lost looking at the information, has anybody any other ideas to look at? Could my 'core' switch be having difficulties with arp tables or performing the routing? As mentioned earlier, its a HP 5130 - so not a cheapo switch. But maybe not up to being a core switch?

Posted

You've rapid spanning tree configured on your core - what is configured on your edge switches?

 

We had an issue with standard spanning tree where our edge ports were sending out dhcp requests before the port went into forwarding mode so they never got a reply. Although not really the issue you are seeing, I'm still wondering about the port protection question I asked a while ago.

Posted (edited)

I'm going through my switches now, starting with the backbone 10gig switches (all HP 5130's). Some of them were not set to RSTP but instead to MSTP so I have changed them all to RSTP. There were no protections in place on these switches so I have now designated (just on these switches so far) which ports are Edge ports and I've enabled BPDU protection on them. I've then cleared the log files so I can see what they start filling up with.

 

I noticed that, particularly on the core switch, there were a LOT of entries for detected changes in topology. This shouldn't be the case as I don't really have an advanced enough network to have plenty of redundant links in place. From what I've read in the manual for these switches, not having the edge-port setting set correctly can trigger an influxof topology messages when devices are switched on/off.

 

We'll see how things go. I'm going to go through all the remaining access switches and ensure they are set to RSTP too.

 

Quick question about switch ARP tables. The other 5130's have an ARP table with about 40 odd entries in, give or take 1 or 2. They are all showing as dynamically learned and they are all VLAN 1 devices i.e. the network switches. Nothing else appears in them... should there be? Should I see the IP and MAC addresses of the PC's plugged into them? Or do they not populate as they are tagged on a VLAN? I ask because the core switch has 1024 entries with devices from all over the school! Not sure if that is right or wrong? I'm also not sure if that means the MAC address table is full or if it will only show me that many?

 

You know when you wish you hadn't bothered in the first place........

Edited by themightymrp
Posted

Well issues are ongoing :( I ran the "Best Practice Analyzer" on our DHCP and DNS servers and all came back with green ticks, so no errors in the setup at that end. It just randomly happens. Sometimes I can ping the default gateway, sometimes I can't. When I can ping it, I can ping other devices in the same subnet.... sometimes.

 

95% of the machines in school have had no issue at all, some are repeat offenders, others it has occurred on just once. But all in different regions of the school and therefore on different VLANs.

 

Open to suggestions....

  • 2 weeks later...
Posted

So, bit of an update on this thread as I was away over half term.

 

Issue is still causing havoc, seems to be more today than before! :( We actually just had a machine drop the network smack in the middle of somebody using their machine. I'm starting to suspect that the switch I am using as a core just can't cope.

 

I'm aware that it is possible to configure the Smoothwall to act as the VLAN routing device. Can anybody explain how to do it? I'm tempted to stay behind in the evening and put the settings in place to make that the core device, hopefully it has more ooompf and can handle it. I would need to create an interface per VLAN I assume and somehow tell it where to look for the DHCP server (ip-helper info). Can anyone help?

 

I would obviously have to muck around changing the default gateways for everything but at this stage I'm completely stumped. All I can think is that the switch is reaching a limit on something (arp ?) which is why most get on but then we have those issues. Maybe the arp addresses age out and the dodgy machines fill the gap???

 

Likewise, if anybody knows how to interpret a wireshark capture I'd appreciate knowing what to look for

Posted (edited)

Just been checking on the core switch from telnet and I may be barking up the wrong tree with the arp table idea.

 

It says 1484 entries are currently in use in the arp table with a possible limit of 16384 :-/ Blows my idea of it being a 1024 limit!

 

So maybe moving to Smoothwall is pointless as a core switch (if even possible). I'm back to needing help with wireshark I think

 

EDIT:

 

Can confirm the 16384 limit. Page 7 of the datasheet show this - we have the JG934A model

 

https://h50146.www5.hpe.com/products/networking/datasheet/HP_5130EI_Switch_Series_J.pdf

Edited by themightymrp
Posted (edited)

Random question: What Client OS? There were problems with Windows 10 showing very similar behaviour a few years ago. It was mostly related to devices that changed network (e.g. laptops transitioning from off-site to on-site), but we did se it on desktops from time to time.

 

Less random question: Have you considered updating the switch firmware? That switch model has had a huge number of updates since the version you are running. https://support.hpe.com/hpsc/doc/public/display?sp4ts.oid=7399472&docLocale=en_US&docId=emr_na-a00040082en_us

Edited by psydii
Posted

Client OS's are Windows 7 Pro in 99% of cases. I did at first wonder if the January hotfix rollup may have caused it but I'm pretty sure more people would be having the issues if so.

 

As for the second question - not so random. I haven't touched the firmware on this thing at all. Don't suppose you have a link to where the newest versions are?

Posted

Found the firmware downloads. You were right - there must be 20+ updates since the one that came with it!!!

 

Just need to figure out my best plan of upgrade as the latest version requires an updated BootROM too :-/ I assume that may have been included in one of the previous updates?

Posted

I had to go through this same process a few months ago. HPE gave me the best upgrade route which meant I had three to apply. All went well except for the last one - the switch didn’t reboot successfully. The next morning I reset the switch and that was fine, but our hyper-v cluster wasn’t happy because each server had been isolated from each other for a number of hours.

After that issue we created another route for the cluster traffic to add some redundancy.

 

Different switch, different issue, but if you’re in a similar situation it may be worth considering.

Posted

So... finding the time to do this upgrade is now a flipping issue! The only available evening this week where there isn't some kind of parents evening/department training etc going on is Friday!! :-( I want the upgrade in place by then to see if it works.

 

Might have to do it remotely and pin my hopes on it not failing. Does anybody know if the .ipe upgrade file will also upgrade the BootROM? I can't seem to find a definitive answer to that

Posted

OK, very interesting discovery!!

 

In the settings of the switch there are 2 sub-menu items under the Network tab. One is MAC and the other is ARP. Currently there are over 1400 devices listed in the MAC screen. Under ARP it was showing as 1024. I've had a number of device this morning that I've been really struggling to make work so I decided to test something. I deleted all of the IP to MAC addresses listed under ARP to do with the BYOD wi-fi domain, taking the total listed ARP addresses down to 934. By refreshing the settings on my troublesome PC's, they all found the network straight away!!!!!

 

It would appear that this buffer is getting full and causing the problem. No I need to figure out what configuration change I need to make to prevent this from happening. Is the ARP list essential??

Posted

you are having a problem and a related counter is at 1024? That sounds like a software/hardware limit!

 

And I think we might be able to solve that:

 

"By default, a device can learn a maximum of 1024 dynamic ARP entries.

If the value for the number argument is set to 0, the device is disabled from learning dynamic ARP entries"

 

"Set the maximum number of dynamic ARP entries for the device: arp max-learning-number number"

 

https://support.hpe.com/hpsc/doc/public/display?docId=emr_na-c04771711

  • Thanks 1
Posted (edited)

This looks like bad news:

"A device can learn a maximum of 1024 dynamic ARP entries."

https://support.hpe.com/hpsc/doc/public/display?docId=c04771726

 

I think your options are :

1) buy a bigger switch (5800 or 3600 series probably)

or

2) redesign your network so that each comms cabinet serves its own set of vlans/subnets, and your core switch routes to each of the comms cabinets, breaking the network down and ensuring the core and edges never have to worry about more than 1024 devices.

or

3) go back to a flat network.

 

4) as an interim measure you could set the arp ageing-timer to something short like 5 minutes to age out devices that have left the wifi, or gone to sleep. The problem is this may only mask the issue and cause further errors.

Edited by psydii
  • Thanks 1

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...