Jump to content

Recommended Posts

Posted

Hi guys,

 

Bit stumped, wondered if anyone has any ideas.

 

Yesterday at around 2:30pm, both our SAN switches lost connection to the SAN. Hyper-V tried to failover but because the connection to the SAN was lost, we were flooded with disk IO errors and most the servers seem to have corrupted, not to the point they are completed screwed, they did come back up... but enough that we are having an absolute world of issues this morning, I have had to restart all the VM's and they were running disk checks, clearing entries for corrupt files.

 

What is confusing me is that most the machines lost connection to the network as well as our cameras back to the NVR we have. Initially before I discovered the SAN was having a fit, I thought one of our access switches in the server room had gone awry, but the SAN switches are directly plugged into the SAN so how did they also lose connection (apologies if I sound dim here and am missing something, I am fairly new to delving into the technicalities of SAN storage)

 

Does this sound like a power issue? a switch issue? a cabling issue? am I potentially missing something? - head full of questions!

 

I do not want this to happen again, obviously, so am trying to work out what on earth happened!

 

Many thanks

Posted (edited)

Hi,

 

What model of SAN / Switches do you have, I've experienced strange issues like this in the past with HP P2000s ISCSI and FC versions, It could be the SAN having a fit.

 

I would say that would be more likely than the switches failing simultaneously.

 

Alex

Edited by Aprice
Posted
We have HP Proliant DL380's. The odd thing is that the cameras stopped working too, servers I can understand if the connection to the disks were lost and so the VM's were playing up, but the cameras don't rely on any VM's, they are just cabled to various switches and some back to the NVR, indicating there must have been a networking issue somewhere too! but both the two stacked cores and two stacked access switches in the server room have an uptime of longer than yesterday, so none of these went down! very confused.
Posted

Did all the cameras go down on the NVR, including the directly connected ones?

 

What SANs do you have? and how are they connected. ISCSI using ethernet or Fibre Channel?

Posted

On the face of it I'd say power issue (surge/spike/brownout), some equipment can handle it better than others and don't show any signs of it having happened, some it has drastic consequences for.

 

Do you have any equipment on UPS (and if not, now may be the time to invest - especially for servers/SAN)...?

Posted
We have HP Proliant DL380's. The odd thing is that the cameras stopped working too, servers I can understand if the connection to the disks were lost and so the VM's were playing up, but the cameras don't rely on any VM's, they are just cabled to various switches and some back to the NVR, indicating there must have been a networking issue somewhere too! but both the two stacked cores and two stacked access switches in the server room have an uptime of longer than yesterday, so none of these went down! very confused.

 

Does the SAN connect via iSCSI or FC, are they completely separate. If they are I'd look at pwr issues. Have you checked the logs on the SAN switch specifically the uptime.

Posted
On the face of it I'd say power issue (surge/spike/brownout), some equipment can handle it better than others and don't show any signs of it having happened, some it has drastic consequences for.

 

Do you have any equipment on UPS (and if not, now may be the time to invest - especially for servers/SAN)...?

 

That’s what we are thinking as a lot of the equipment that failed are unrelated to each other, and so it doesn’t make sense to be an issue between them. Interestingly, we do have a UPS and this didn’t show any signs of the battery having been depleted at all so if doesn’t appear to have had to kick in.

Posted
Did all the cameras go down on the NVR, including the directly connected ones?

 

What SANs do you have? and how are they connected. ISCSI using ethernet or Fibre Channel?

 

The ones plugged into the NVR stayed up, all the others went down, but what we have discovered is that the access switch that they are plugged into that has all the VLAN tagging, lost it’s configuration so we have had to reconfigure all the VLAN tagging for the cameras and they have now come back up.

 

The SAN is a HPE MSA 2050, the san switches are plugged into each other via Ethernet and into the SAN via fiber.

Posted
The ones plugged into the NVR stayed up, all the others went down, but what we have discovered is that the access switch that they are plugged into that has all the VLAN tagging, lost it’s configuration so we have had to reconfigure all the VLAN tagging for the cameras and they have now come back up.

 

The SAN is a HPE MSA 2050, the san switches are plugged into each other via Ethernet and into the SAN via fiber.

 

So was there a power cut, did the configuration get wiped, was it not saved :-) would not be the first, lol.

Posted
All the uptimes of all the equipment are still over 90 days. The switch connected to one of the SAN's lost some but not all of it's config (it's a Ubiquiti so auto saves so no chance of it not being saved) Essentially the SAN's have reported that they both lost connection to each other at exactly the same time, meaning the failover clustering tried to kick in and failed, which has caused issues with disk IO. I am incredibly stumped. Nothing actually powered off, it seems to be a network issues on the Ubiquiti switch, but the unifi panel is reporting no alerts for that switch and no indicator of an issue. All the sockets are surge protected also. I really have no idea what happened! I have been looking into it all day. :(
Posted (edited)
All the uptimes of all the equipment are still over 90 days. The switch connected to one of the SAN's lost some but not all of it's config (it's a Ubiquiti so auto saves so no chance of it not being saved) Essentially the SAN's have reported that they both lost connection to each other at exactly the same time, meaning the failover clustering tried to kick in and failed, which has caused issues with disk IO. I am incredibly stumped. Nothing actually powered off, it seems to be a network issues on the Ubiquiti switch, but the unifi panel is reporting no alerts for that switch and no indicator of an issue. All the sockets are surge protected also. I really have no idea what happened! I have been looking into it all day. :(

 

Any chance you can post a bit of a diagram, but please obscure any sensitive and identifiable information (public ip addresses, fqdn's, organisation names, etc.) But if you could include the switches Make and model and firmware.

 

I've setup iSCSI before and used separate switching and had a lot of redundancy i.e. separate network cards split between 2 different controllers on the SAN via 2 separate switches, this was on VMWare. At a previous place we also had a Fibre Channel SAN and that was split too for redundancy (multiple FC adapters, to redundant controllers via 2 separate switches) again on VMWare.

Edited by Davit2005
Posted
Complicated stuff!

 

I've only read a little bit about VMs on SAN vs NAS and thought the whole point of SAN was redundancy and resilience...?

 

@cdwyersandysecondary are your switches on your UPS?

 

Took me a while to understand a few years ago, in essence SANs present 'disks' to servers where as NAS presents files. The lines have been a bit blurred with ISCSI on NASs now.

 

SANS do tend to be more resilient with redundant management modules etc. It's not the first time I've heard or seen one have a wobble.

Posted
All the uptimes of all the equipment are still over 90 days. The switch connected to one of the SAN's lost some but not all of it's config (it's a Ubiquiti so auto saves so no chance of it not being saved) Essentially the SAN's have reported that they both lost connection to each other at exactly the same time, meaning the failover clustering tried to kick in and failed, which has caused issues with disk IO. I am incredibly stumped. Nothing actually powered off, it seems to be a network issues on the Ubiquiti switch, but the unifi panel is reporting no alerts for that switch and no indicator of an issue. All the sockets are surge protected also. I really have no idea what happened! I have been looking into it all day. :(

 

Don't suppose you can get on the SAN web interface? Might give some indication if all the storage interfaces went down at once.

Posted
I'd say that there was probably a broadcast storm or something has saturated the network and ground those switches to a halt. Do you have any kind of event logs on the switch with warnings for cpu utilisation or port utilisation?
Posted

Long ago (2015?) I had a P2040 lock up both controllers simultaneously. Turns out is was a 'known' bug that had been patched a few days before we hit it. HP woke up engineers on a Sunday night and sent them to site to help us put things back together. We had quite a lot of damaged VMs and worked with MS to fix those. Pretty solid response from HP given we were not on a 24/7 Care Pack.

 

The point to this anecdote is that it might be worth comparing your firmware levels with HP's current recommended, and consider updating.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...