Jump to content

Recommended Posts

Posted

last night we had a total failure with something, i dont know what, but i was happily working away having a bit of a tidy up when all of a sudden the VMs failed and shutdown even though the hosts were fine and the SAN was fine. No power cut as i could still connect to my machine.

 

Anyway, ive got 2 hosts. HVH01 and HVH02. I have the VMs split across them both so 10 on each. If i'm on HVH02 in Failover cluster, i cannot connect to VMs that are on HVH01 but on HVH01 i can connect to VMs on HVH02.

 

Ive noticed on HVH02 under disk management ive got C:, Cluster Storage, QuorumDisk (D: ), System Reservered which are all ONLINE and on HVH01 ive got C:, Cluster Storage, System Reservered however cluster storage and what looks like the QuorumDisk are both OFFLINE. Should this be the case?

 

I cant put the Quorum or Cluster Storage online, it gives me an error... "The specified disk or volume is managed by the Microsoft Failover Clustering component. The disk must be in cluster maintenance mode and the cluster resource status must be online to perform this operation".

 

I've read that i should put the Quorum and Cluster Storage in Maintenance mode by going into the Failover Cluster Manager > Expand Storage > click on disks. Right click on either the Storage or Quorum > highlight more actions and turn on maintenance mode.

 

Im not sure what happens after that though and im not sure that maintenance mode does.

 

Should both Hosts have the Quorum and cluster storage online?

Posted

To add to this, i've just seen in FailOver Manager > Networks that the SAN network status has a red cross against it and says Partitioned.

 

The LAN network has a green arrow on it stating its Up.

 

Is there something majorly wrong with my HyperV setup?

001.PNG

Posted

IIRC The cluster services manage which disks are online on which server. in other words let it manage it dont mess with bringing disks online manually as you can corrupt the data on your drives.

 

That being said, looking at the Screenie in the second post it may well be a network issue of some form (have you updated card drivers or anything like that) does your HV boxes have the correct network configs etc?

Posted
Is this an iSCSI San with multi-pathing? We had an issue on just one switch and the whole thing failed. Agree with sister_annex, network looks a likely candidate. In the iSCSI config you can verify the paths and rebuild them from memory.
Posted

One host will show the disks online and the other offline--that's normal, and shouldn't be messed with. The on/offline can be mixed as well--so right now, one of my hosts has the CSV online and the Quorum offline, and the other vice versa.

 

The network thing looks suspicious though--I don't have my SAN listed in my networks, but I'm on multi-path fibre channel. I suspect an iSCSI SAN would show there, and if so, that could well be your problem.

Posted

Thanks guys for replying.

 

Ill leave the quorum and the cluster storage offline on HOST1.

 

As for the second issue, ill dig out the setup guide on how its been configured. I didnt do it, we got a company in to do it and ill report back.

Posted

I'm not sure if this helps?

 

HP servers – two identical specification, HP DL380PG8 2U rack-mounts, called HVH01 and HVH02. Each server has 128GB of RAM, 2 X 300GB drives in a RAID1 pair, single PSU and dual Intel E5-2640 CPUs.

• Each has an ILO card, connected via purple cable to management ports on the switch for out of band access, cabled via the HP2524 switch on the back of the rack.

• The base networking on the system is 4 X 1GB NIC ports. There is also an extra PCIe expansion card installed, with an additional 4 X 1GB NIC ports.

• Each set of 4 NIC ports will be split equally between production LAN and iSCSI networks, and with a connection to separate switches where possible.

• All Production NIC ports will be teamed via the Windows 2012 teaming panel into a combined trunk port between switch and server / VM Host, giving an aggregate 4GB of transit over 2 adapters.

• The servers have one power supply, with the option to add a second at a later date. It is strongly advised that the purchase of a second PSU for each server is completed as soon as is practical.

• Both servers are supplied with sliding rails and cable management arms to allow servers to be pulled from the racks whilst operational.

• Server 1 is configured with a management IP of 172.16.26.100, and an ILO address of 172.16.26.101

• Server 1 has iSCSI ports of 192.168.0.98-101

• Server 2 is configured with a management IP of 172.16.26.102, and an ILO address of 172.16.26.103

• Server 2 has iSCSI ports of 192.168.0.102-106

• The servers have all of the local hard drive space allocated to drive C, giving approximately 280GB of usable space. After Windows installation and the creation of the support drive and utility folders, there should be >200GB remaining.

• The password for the administrator account is detailed on the separate Password/Logon sheet at the end of this document.

 

SAN Storage – DotHill 2U disk array with 12 X 1TB SAS 7.2K hard drives

• The two ports on the top controller are split into red and green connectors. The two ports on the bottom controller are split the same as the top. Each controller also has a dedicated management port, connected via purple cable, cabled to the management switch.

• Red and green connections are made to the switch, with red connections to the red switch and green connections to the green switch.

• The disk array is configured with an effective level of RAID5 to provide the optimum blend of speed and resiliency for the School environment. Overall, there is 10,000 Gigabytes of usable storage on the array.

• Controller 1 is assigned an IP of 172.16.26.104

• Controller 2 is assigned an IP of 172.16.26.105

• iSCSI ports are assigned IPs in the range 192.168.0.106-109

 

As for pinging

On HVH01..

192.168.0.98 = Successful ping

192.168.0.99 = Successful Ping

192.168.0.100 = Successful Ping

192.168.0.101 = Successful Ping

192.168.0.102 = Successful Ping

192.168.0.103 = Failed Ping

192.168.0.104 = Successful Ping

192.168.0.105 = Failed ping

192.168.0.106 = Successful Ping

192.168.0.107 = Successful Ping

192.168.0.108 = Successful Ping

192.168.0.109 = Failed Ping

 

On HVH02..

192.168.0.98 = Successful Ping

192.168.0.99 = Successful Ping

192.168.0.100 = Successful Ping

192.168.0.101 = failed Ping

192.168.0.102 = Successful Ping

192.168.0.103 = Successful Ping

192.168.0.104 = Successful Ping

192.168.0.105 = Successful Ping

192.168.0.106 = Successful Ping

192.168.0.107 = Successful Ping

192.168.0.108 = Successful Ping

192.168.0.109 = failed ping

Posted

Just managed to fix this issue after reading a few articles on Spiceworks (https://community.spiceworks.com/topic/292189-windows-failover-clustering-networks-problem) and Experts Exchange (https://www.experts-exchange.com/questions/24893525/Cluster-network-is-Partitioned-network-connections-are-Unreachable.html).

 

I jumped on one of the servers HVH02, disabled the offending NIC (#7) and re-enabled. It reported that #8 was at fault, so disabled and re-enabled and it said #7 was the offending NIC. So once again, disabled and en-enable #7, refreshed Networks, and its all connected now :D.

Posted

I thought the problems were solved. I saw the little green arrow to saw it was up but it turns out it red X came back. I think the issue lies in the NICs but where im unsure.

 

Would you be able to send over the troubleshooting steps that you had.

 

Thanks

Posted
Just a thought, you don't have NIC teaming or trunking enabled on your switch do you?

 

Im not sure, it was setup by an external company who put in 2x Cisco SG300 switches. I dont even know the IP addresses of these switches to have a look at the config.

Posted

you should be able to see if there is NIC teaming set up on your servers

networkConnections.PNG

 

If you have that multiplexor driver setup you are using the windows NIC Teaming - I know there is other software available but this is what we use :)

 

I'd be getting in touch with your supplier and asking the question - they should have left you with config documentation at least!?

Posted
Im not sure, it was setup by an external company who put in 2x Cisco SG300 switches. I dont even know the IP addresses of these switches to have a look at the config.

 

Is the team setup as Switch Independant on Hyper-V??? This does not require any config on switch and links can go to 2 seperate switches from what I believe.

 

I think it is possible to only assign iSCSI Initiator to specific NICS then these can be vlanned off to isolate SAN from production.

Posted (edited)

Firstly can you confirm if the NICs you have highlighted above are in a team, or if each Nic has the IP address assigned directly to it?

 

Secondly can you upload an Ipconfig /all for each server pls?

 

Thirdly what is the Ip of the array that these NICs connect to?

 

Why so many IP's? If the iscsi NICs on the servers are in a team then it is in my view pointless having more than 1 ip on a teamed interface, personally I don't team iscsi interfaces.

Edited by bart21
Posted (edited)
you should be able to see if there is NIC teaming set up on your servers

[ATTACH=CONFIG]38648[/ATTACH]

 

If you have that multiplexor driver setup you are using the windows NIC Teaming - I know there is other software available but this is what we use :)

 

I'd be getting in touch with your supplier and asking the question - they should have left you with config documentation at least!?

 

See attached for the screenshots of the server NICs. THey do have teamed NICs.

 

Is the team setup as Switch Independant on Hyper-V??? This does not require any config on switch and links can go to 2 seperate switches from what I believe.

 

I think it is possible to only assign iSCSI Initiator to specific NICS then these can be vlanned off to isolate SAN from production.

 

Im not too sure, ive attached screenshots of the VirtualSwitchManager for for HyperV hosts.

 

Firstly can you confirm if the NICs you have highlighted above are in a team, or if each Nic has the IP address assigned directly to it?

 

Secondly can you upload an Ipconfig /all for each server pls?

 

Thirdly what is the Ip of the array that these NICs connect to?

 

Why so many IP's? If the iscsi NICs on the servers are in a team then it is in my view pointless having more than 1 ip on a teamed interface, personally I don't team iscsi interfaces.

 

Ive attached text documents of IPConfig from both hosts. HVH02 has 2 IPs attached to it 172.16.26.106 & 172.16.26.102. .26.106 is "the failover cluster and assigned an IP of 172.16.26.106/21."

 

The SAN RAID controllers are:

A = 172.16.26.104

B = 172.16.26.105

 

The SAN iscsi IP addresses are:

192.168.0.106

192.168.0.107

192.168.0.108

192.168.0.109

 

Ive found this in the documentation:

 

Each host has 8 physical Network connections. The connections on the motherboard are numbered from left to right, whilst looking from the rear, as 0-3, and the connections on the PCIE card are numbered from left to right, also looking from the rear, as 4-7.

 

Connections 0,1,4 and 5 are set to be members of a single team on the Hyper-V host, called HYPER-V-LAN. The team is set to trunk all traffic, and is not assigned to a vLAN.

 

This is used for production LAN traffic. Red and Green patch cables are connected for this team.

 

The team is set to use the following settings: Team mode – Static, Load Balance Mode – Hyper-V, Standby Adapter – None.

 

Connections 2,3,6 and 7 are set to be separate iSCSI connections configured with MPIO, and are not assigned to a vLAN. Yellow patch cables with red and green flashes are connected to this team.

 

In Hyper-V manager, one virtual switch is defined.

1. HYPER-V-LAN, default VLAN, bound to the default team interface of the Host team.

 

Each VM will thus have 1 network card assigned to it, bound to the production LAN. No connection is made to the cluster heartbeat / migration network or iSCSI networks

ipconfig_hvh02.txt

ipconfig_hvh01.txt

HVH02.png

HVH01.png

Edited by timbo343
Posted (edited)
I'm not sure if this helps?

 

 

I'm behind on this and this may be of no help, but I'd be questioning the documentation or setup. Both Server 2 and your SAN are documented as using the same IP address for iSCSI - 192.168.0.106

 

[edit] Reading it again it seems that you have eight physical NICs in each server, four teamed for your VMs, four using multipath for iSCSI and one for management? It doesn't add up - the documentation is awful, it looks like it's been copied and pasted from other documents rather than written with your install in mind from scratch.

 

The two servers have different references to what we assume are the same functions. vEthernet(BYPER-V-LAN) & LAN - Windows 2012 Team etc. It's a bit of a mess to be honest.

 

When did this company install? Have you contacted them?

Edited by IrritableTech
Posted
I'm behind on this and this may be of no help, but I'd be questioning the documentation or setup. Both Server 2 and your SAN are documented as using the same IP address for iSCSI - 192.168.0.106

 

[edit] Reading it again it seems that you have eight physical NICs in each server, four teamed for your VMs, four using multipath for iSCSI and one for management? It doesn't add up - the documentation is awful, it looks like it's been copied and pasted from other documents rather than written with your install in mind from scratch.

 

The two servers have different references to what we assume are the same functions. vEthernet(BYPER-V-LAN) & LAN - Windows 2012 Team etc. It's a bit of a mess to be honest.

 

When did this company install? Have you contacted them?

 

It would seem that way. I got lost with the documentation i must admit. It was installed summer 2014 and the install didnt go smoothly.

 

Im not going to name the company but they were local to us.

Posted
It would seem that way. I got lost with the documentation i must admit. It was installed summer 2014 and the install didnt go smoothly.

 

Im not going to name the company but they were local to us.

 

So it's been running two years, then suddenly...

 

To confirm, your VMs are running and your users are going about their day?

The issue is that you're seeing errors in your cluster setup and clearly things aren't optimal?

Your SAN network in Fail Over Cluster Manager is still partitioned? Personally I wouldn't have setup this network allowing cluster traffic as well as iSCSI - Do we know why they did that?

In MPIO Properties is you SAN listed as a Device? Is iSCSI enabled under Discover Multi-Paths and is this the same on both hosts? My compellent SAN did this when I set it up - I guess yours might be different.

In iSCSI Initiator are you seeing the same targets on both servers?

What does the management interface of the SAN tell you - is it complaining that a path has dropped, has it got permissions in for both servers, does it show the correct iSCSI targets for the servers?

 

Personally I'd start sketching the physical layer out on a bit of paper, noting IP addresses etc. and then start thinking about what that means in relation to the NIC teams and such. Grab a console cable and look at the config on your cisco switches, they may well be blank but something may have been configured. Then I'd probably think about the OS side of networking - NIC driver settings, Team settings, iSCSI Initiator, MPIO. Finally looking at application layer and your cluster / hyper v setup.

 

As a newbie to Hyper-V clusters (we built one this summer) I found it a little hard to get my head around all the various places you need to configure different elements. Hyper-V Manager - Virtual Switch Manager, iSCSI Initiator, Fail Over Cluster Manager - Networks, Storage, Nodes, MPIO Properties. etc. etc. etc. We're not teaming NICs though so that was one thing I didn't need to contend with. I'm rambling and not helping...

Posted
So it's been running two years, then suddenly...

 

To confirm, your VMs are running and your users are going about their day?

The issue is that you're seeing errors in your cluster setup and clearly things aren't optimal?

Your SAN network in Fail Over Cluster Manager is still partitioned? Personally I wouldn't have setup this network allowing cluster traffic as well as iSCSI - Do we know why they did that?

In MPIO Properties is you SAN listed as a Device? Is iSCSI enabled under Discover Multi-Paths and is this the same on both hosts? My compellent SAN did this when I set it up - I guess yours might be different.

In iSCSI Initiator are you seeing the same targets on both servers?

What does the management interface of the SAN tell you - is it complaining that a path has dropped, has it got permissions in for both servers, does it show the correct iSCSI targets for the servers?

 

Personally I'd start sketching the physical layer out on a bit of paper, noting IP addresses etc. and then start thinking about what that means in relation to the NIC teams and such. Grab a console cable and look at the config on your cisco switches, they may well be blank but something may have been configured. Then I'd probably think about the OS side of networking - NIC driver settings, Team settings, iSCSI Initiator, MPIO. Finally looking at application layer and your cluster / hyper v setup.

 

As a newbie to Hyper-V clusters (we built one this summer) I found it a little hard to get my head around all the various places you need to configure different elements. Hyper-V Manager - Virtual Switch Manager, iSCSI Initiator, Fail Over Cluster Manager - Networks, Storage, Nodes, MPIO Properties. etc. etc. etc. We're not teaming NICs though so that was one thing I didn't need to contend with. I'm rambling and not helping...

 

Yep, the VMs are running and the users are going about their day and not noticing any performance. It's been running ok-ish for 2 years, every now and then, HVH01 used to fall over but through trail and error ive managed to get it to be stable.

Yup, the issues im seeing i guess aren't optimal.

Im not sure why they set it up allow cluster traffic as well as iSCSI traffic.

Ive not heard of the MPIO and a quick google on where to find it - its not listed in either of the HVHosts server manager dashboards.

In iSCSI initiator, the target is the same on both hosts, however, HVH02 has data in the volume list but HVH01 doesnt have this.

The SAN isnt complaining about anything, it thinks everything is fine.

 

Once i get some time, ill link into the switches and see what they are telling me.

Posted
Yep, the VMs are running and the users are going about their day and not noticing any performance. It's been running ok-ish for 2 years, every now and then, HVH01 used to fall over but through trail and error ive managed to get it to be stable.

Yup, the issues im seeing i guess aren't optimal.

Im not sure why they set it up allow cluster traffic as well as iSCSI traffic.

Ive not heard of the MPIO and a quick google on where to find it - its not listed in either of the HVHosts server manager dashboards.

In iSCSI initiator, the target is the same on both hosts, however, HVH02 has data in the volume list but HVH01 doesnt have this.

The SAN isnt complaining about anything, it thinks everything is fine.

 

Once i get some time, ill link into the switches and see what they are telling me.

MPIO (Multipath I/O) is a feature that I would have installed to utilize your two paths to the same device. Once installed the management gui sits in Control Panel - Administrative Tools - MPIO. This all depends on if your SAN allows active/active connections though.

I'm confused about what you're seeing in your iSCSI Initiator. I don't see anything in my volume list on any of my servers - Under the discovery tab I have my SAN controllers IP addresses - LUNs just pop up in disk manager then I've added them as CSVs to failover cluster manager.

I'm not being very much help really, but it's very hard to diagnose without being a bit more hand on. Either way - it might be me back here in two years time saying I set it up like this but now things aren't working...!

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...