Jump to content

Recommended Posts

Posted

Evening all,

 

Following on from my post in the "annoying" thread, where Dell have quoted £7k for 1 years support on our EqualLogic P6100X(this is for top of the range "Plus" 4hr but even the normal 4hr was £5k!). It looks like I may be moving the SAN to non-mission critical backup duties quicker than I thought!

 

So currently we have 2 R710's and the P6100x with a dedicated Dell iscsi switch to link them all up, about 14TB of RAID, 21TB raw. 2-node cluster primarily running hyper-v based virtual servers

 

I want to replace everything, so new hosts, new storage, new switch etc.

 

Now all I need is HA for our VM servers and failover in case of failure of a host(technically same thing, but you get the idea), and on a 2 node cluster, compute is covered, but a san is a single point of failure for storage.

 

So normally, I'd be off to a vendor for new everything, but I am thinking there must be a better way to do this other than blowing tens of thousands on a SAN. I've had multiple SAN vendors through the door, with very large quotes that make me shudder! All well and good for oodles of nodes with redundant replicating SAN's, but our needs are simple!

 

I've been reading about storage spaces and Starwind(i've read the white paper on a 2 node cluster, but can't work out if it just replaces a SAN or does the Hyper-v functions as well, so hosts become the SAN) and other weird and wonderful linux products, but I do need some support(don't we all ;) )

 

So what options do I have in day and age for a system where my servers will always be up even if a node goes down(or I take it down for updates etc)? 20TB after RAID would be nice, and as fast as is possible :) No budget set, as we try and fit money to projects, rather than projects to money, if you get my meaning, it has to be right for the job not for the vendors monthly target(sorry for the cynicism!) I also like fit and forget, see below:

 

As a side note, I do love my EQ SAN, it has given me zero problems in nearly 5 years, bar one incident when it had been up for 365 days and it shut itself down which turned out to be a firmware bug, solved by upgrading the firmware!

 

Thanks

 

James

  • Thanks 1
Posted

Is it a school environment? I just sometimes question if warranty is strictly needed, does the school need 100% all of the time? What if the SAN is goosed and you need to restore backups anyway, would having warranty help with this? Just in a time where schools are in a difficult position financially, personally for where I work, I feel it really isn't needed on a 4 hour call out.

 

Take the system I'll be looking at implementing, 2/3 Refurb V-Hosts with fast local storage and then re-purposing an old but decent server to run Hyper-V replica, replicating core file servers and SIMS. Things like SCCM, printers, app server, door entry, classroom management and so on, I don't deem as mission critical, so being down for a few hours/couple of days whilst core systems are restored shouldn't majorly affect staff and students working. All this will also then be backed up to a DPM server which could be spun up as quick as the VHD copies or if the DPM server is up for, even on that.

 

I think that system will be fast, each host will have 6TB usable storage in RAID10, 96GB RAM, 16 Cores. It'll be quick, cheap, hopefully reliable and I've planned for disasters and redundancy.

Posted

It's a college environment, but we always run it as if its a enterprise and invest in our infrastructure, and we keep support up on mission critical systems. Saying that, I can count the amount of time I've had to call support(on servers and associated bits and bobs) on both hands in the 7 years I've run the system, and that's been the odd HD or PSU or the above mentioned 365 day error in the firmware. We already run DPM to our DR rack at one of our secondary centres, where we aim to be back up in a day if our main centre goes dark. Warranty would help in that they would replace the box, rather than us having to buy a new one, but yes the 4hr in that situation means we would be down to our day restore time whatever happened.

 

Dell only allow ProSupport for 7 years, so our hosts have to be replaced next year whatever happens, but this cost for the equalogic support may push that timetable up, as even the next day support is nearly 4k inc the VAT!

 

I am assuming if I go for a "hyper-converged" 2 node in the starwind style, I need the same amount of storage on both nodes, which means a hefty price in hard drives, especially if I want pure SSD..

Posted

We had a Fujitsu DX60 S4 with close to 20TB after RAID.

 

5 year warranty included, close to £9k. Included dual controllers.

 

I don't think that is bad!?

Posted
The Fujitsu and Nimble are both SAN's, I'm looking for alternatives to having a single point of failure(barring buying 2 of them!). ie no SAN infrastucture. Although thanks for the recommendations if in fact we do go the SAN route after all :)
Posted

To avoid a storage single point of failure you are looking at doubling down on the storage.

 

Even a raid array with dual controllers is effectively a single point of failure if the fault domain is power outage for example.

 

I think we need to understand what you are trying to achieve? Is your current SAN a single point of failure?

 

Do you have one or two server rooms, external generator and ups, redundant core / distribution / edge switches

Posted

I am basically looking to see if we need a SAN going forward, for whatever reason, lack of redundancy, price, etc etc. I am not really worried about a single point of failure, that is just one of the arguments being banded around for ditching them. I believe it is a lot of money to spend on something that we may not need in these days of SDS and hyper-converged. Yes SAN may be the best bet, and we will replace ours if needed, but I am after other options and opinions. Original question reworded, 2-node hyper-v cluster, 20TB after RAID, what would you do? and by you I mean everyone not just you @geezersoft :)

 

The price of renewing support on the current SAN has just pushed this quandary to the forefront, do we pay and keep going another year, or do we take that £5k and put it towards new infrastructure now. Fag packet maths, I can do 2 servers with enough storage for say 15-20k, or double the cost and have 2 hosts but with a SAN and switches. ie in this day and age, should we be shelling out 20k plus for a shiny SAN in such a small environment.

 

(By the by,We have a whole room 10K UPS with about half an hour of uptime, so barring a long power cut, power is sorted, 10 GBe dual switch core with enough spare to run on a single node if needed(albeit a manual wire pull and push to get going again))

Posted

If you're only having two nodes, you could even consider SAS based storage as you'd then be able to connect them up in a two node config, then you eliminate the need for the switching and extra costs of a SAN.

 

I'd say a two server setup with local storage could be a good idea, with the money you'd save on not buying a SAN, you could even put toward some SSDs for the server and you could use the old equipment as storage backup and Hyper-V replica seeing as it'll still be up to the job, but not 100% mission critical as hopefully you'll never need it.

 

You could even look in to storage spaces https://docs.microsoft.com/en-us/windows-server/storage/storage-spaces/storage-spaces-direct-overview

Posted
We had a Fujitsu DX60 S4 with close to 20TB after RAID.

 

5 year warranty included, close to £9k. Included dual controllers.

 

I don't think that is bad!?

 

@snagrat, can I ask who you got that from? We're looking to renew our SANs, that sounds like a bargain!

Posted
@NB9457 that's what I'm thinking, 2 matched servers with some kind of replication/3rd party offering in software, 10Gbe'ed together, load balancing in normal mode and 1 node able to take the load of both nodes if needed. SSD based would be good, with the savings. Then the old PS6100x will do on-site backup duty, replacing my aging PE2950 and MD1000 driveshelf
Posted
Right, been doing more of this reading up on things malarkey. My main concern is being able to still operate if a node goes down and the ability to move Vm's to another node to do maintenance. To get the storage size I want in a 2-node cluster is still going to cost an arm and a leg, be it local, DAS, SAN etc and be complicated. So it seems that there is an option for Share Nothing Live Migration, which sounds interesting, ie the ability to move running VM's between non-clustered hosts. The last R730XD I bought with 12TB of storage came in at around the £5k mark, dual Xeons, 96GB of RAM. That was at the end of last year, and I know pricing is severely wobbly at best these days, but I could get 2 of these with more RAM and more storage with 10GBe cards, use one port on each to my 10GBe core and the run the other between the servers for a migration network. This wouldn't give me HA, but it would give me the option of migrating all to one for maintenance windows(depending on storage capacity on a single unit) and we could live with the downtime on a 4hr prosupport contract for things breaking(although the VM's could be retrieved out of DPM storage to the one running server in the meantime)and be very very budget friendly. Does anyone have any real world anecdotal on SNLM and how it deals with transferring big VM's live, ie file servers with TB's of storage(no drives bigger than 1TB at the moment)? Thanks for the replies so far.
Posted

Time will be the biggest hurdle. I have read blog posts talking about transfer speeds as low as 2GB per minute, so for a 1TB drive you could be looking at a 10 hour+ migration time.

 

Depending on how many servers you need to move you could be looking at days to live migrate.

Posted

Our standard DR RTO is a working day.

 

I've read around and people are getting 7-8 Gbps on 10Gbe after tweaking and disabling broadcom NIC's and other twiddles, so that means circa a gig proper a second if I got my maths right, so 16 minutes per TB, I can live with that :) most of my VM's are a lot less

Posted
But do you mean DR or BC. DR means site loss, BC would be something like a host loss. If DR then you would need to effictively spin up a new site including hardware from a total loss.
Posted
I think so, and then my old SAN and cluster will become my DPM backup environment, probably only one of the hosts though, and probably forgo the switch and hard wire straight to the iSCSI optimized NIC card in the server(or 2 4 port cards for both controllers). then a 10Gbe link from the newly crowned DPM server to each of the new hosts to create a dedicated backup network. It will replace the MD1000 driveshelf and its aging 2950 dell head. Just need to sort some quotes and then go have a word with finance :) although as I've just spent high 5 figures on a Horizon VDI environment, I might have to grovel or wait till it was actually planned, which is easter next year...
Posted
We have a DR site at another centre where all our DPM backups are collected by a secondary DPM system, where we can spin up all the VM's in the case of the main site going dark, then we have round robin on our DNS for all our external IP's. That's the 1 working day RTO. Restore from backup and reprogram watchguard to point all the external IP's to the "new" servers, as at the moment, they all point to the main site, so round robin doesn't give us nothing found 5 times out of 10 :)

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...