garethEds Posted November 5, 2019 Posted November 5, 2019 Evening all, So the in-place cluster upgrade that went very well has now started showing issues. It was all fine last MOnday when I left it. When I went to check today there were several hundred cluster errors. The CSV/Cluster network (as it's named) couldn't talk between the three servers - the three servers are 192.168.0.1, 10.168.0.2 and 192.168.0.3 - all have the same subnet and nothing has changed since the 2016 install (except the in-place upgrade). When pinging between the servers we are getting 'TTL Expired in Transit' errors. On one server the CSV/Cluster network isn't even showing as recognised - saying a network cable is unplugged (but it isn't). This is the first of the issues I am trying to solve. I've checked all the network drivers and all I can think of at this moment is that on the third server, there is a 4 port NIC with a driver from 2014. There are no drivers for 2019 that I can find. All servers are HP 380g8 (e and p models). I've put all of the virtual machines onto one server at the moment whilst I try and sort out the cluster. Everything else seems to be working - except the networking. Any advice anyone? Cheers Gareth
garethEds Posted November 6, 2019 Author Posted November 6, 2019 Right - the cluster network is fixed. Now the next issue..... UDP Port 3343 - seems the servers have closed port 3343. I've re-opened the ports on all HyperV servers in Windows Firewall - but none of the networks can talk to each other. When we validate the cluster this is the error we are getting. I've opened the ports. Gareth
ChrisMiles Posted November 6, 2019 Posted November 6, 2019 (edited) I get this problem too on our 3 server cluster with HP DL380 G9s. I think its due to bad drivers for the Melanox 10Gb nics in ours. What do you have in yours? Edited November 6, 2019 by elsiegee40 Edited for language by mods
garethEds Posted November 6, 2019 Author Posted November 6, 2019 I thought it could be the drivers: Server 1 / Server 2 - same spec 2 x 2 Port 1Gb 361T Driver: Intel / 10-05-2019 / 12.18.8.22 1 x 4 Port 1Gb 366i Driver: Intel / 10-05-2019 / 12.18.8.22 Server 3 HP NC 375T PCI Express Quad Port 1Gbit Driver: QLogic Corp / 01-10-2014 / 5.3.30.1001 HPE Ethernet 1Gb 4Port 331FLT Driver: HP / 25-10-2017 / 20.8.5.0 I'm going to look at drivers online to see if there are better ones. Gareth
ChrisMiles Posted November 6, 2019 Posted November 6, 2019 (edited) Hmm well none of those are Mellanox but the 331 is Broadcom which are also known to cause problems I believe. When you validate the cluster, which nodes does it say can't talk to each other? Edited November 6, 2019 by ChrisMiles
garethEds Posted November 6, 2019 Author Posted November 6, 2019 Hmm well none of those are Mellanox but the 331 is Broadcom which are also known to cause problems I believe. When you validate the cluster, which nodes does it say can't talk to each other? It says that none of them can talk to each other on any of the cluster networks. Gareth
garethEds Posted November 13, 2019 Author Posted November 13, 2019 Update: Have ordered a newer Intel based network card so I can remove the Broadcom one. Hopefully it will improve things. Gareth
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now