Jump to content

Recommended Posts

Posted

I have begun to look after the IT at a Utc that is part of the trust I work for. We have a HP Virtualstorage SAN that has gone down twice due to the raid controller failing. This has happened twice in 9 months. I think there has to be a fault but convincing HP of that might be another matter.

 

The Utc has 2 core switches for resilience, 2 hyper-vs for resilience but only 1SAN. I know they are expensivr but what do other schools do, do you just have 1 SAN or do you go for total resilience and get 2?

Posted

How many physical servers are using the storage? Unless you're ~10+ physical servers then look at a local storage setup.

 

If you want redundancy you're going to want a hyper-converged system, ie Scale Computing, Starwind, HP Simplivity, or a "DIY" hyper-converged system running Storage Spaces Direct on Windows Server 2016 Datacentre or Starwind vSAN across your nodes.

Posted
Identify which kind of support contract you have in place with HP and leverage that. A HP Virtualstorage SAN would not have been cheap and you can darn well expect it to work! Twice in 9 years I would be disappointed with, but twice in 9 months in unforgivable.
Posted
1 SAN may suffice, but certainly 2 controllers for me

 

...which is absolutely fine until a fault in a firmware upgrade brings it all crashing about you....

 

"Shit happens"...and it doesn't matter how much you spend or what procedures you put in place....as many of the biggest IT infrastructure providers in the country will tell you....can guarantee your system won't go down.....and yes its possible that it can happen twice in 9 months....after all somebody wins the lottery every week....

 

I'd be checking and asking careful questions about what was done last time.....was it really the controller that had failed....did they really replace it? Was the original still in fact working.....if so perhaps it was a power related issue....or some unexpected disk issue that the controller did not handle well....did anyone actually look at the logs....are there errors which look similar this time? Had it simply lost or corrupted its internal configuration?

 

I'd be pretty nervous if (a) the supposedly failed controller - seemed to start up just fine on a clean system...(b) they simply replace the controller ...because they are following a menu script without doing any further work to identify the problem

 

...because if either or boh of the above are true....I be willing to wager a small bet (and I'm not normally a betting man) that your SAN is going to fail again.....

Posted
...which is absolutely fine until a fault in a firmware upgrade brings it all crashing about you....

 

Could you not opt to not do a firmware partner upgrade and see how the upgrade went prior to upgrading the working controller?

 

I'm so worried now I'm off to buy another SAN and 2 controllers :)

  • 2 weeks later...
Posted
When you say gone down? do you actually mean unable to provide service to users? We have one SAN but it has 2 x controllers 2 x PSU and spare disks waiting for failure at which point they would kick in. With regards to upgrades we always do them during breaks so failure would not interrupt service. If your SAN has failed completely i would defo be looking at how resiliency failed.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...