Jump to content

Recommended Posts

Posted

Guys I'm starting to look at upgrading some of our servers next financial year and I'm considering the possibility of a single disk server with no RAID configuration.

 

As budgets get tighter and tighter I can probably afford to put in place two single disk Hyper-v host servers, one as a replica for fail over, for only slightly more than it would cost for a single server with multiple disks, RAID controller, etc. VM's would be locally stored on this server.

 

Having one of the servers as a replica would mean that 'should' the drive die in the main server I could switch over and maintain service continuity, then rebuild and reverse replication.

 

Am I being stupid?

Posted (edited)
That just strikes as being a really expensive, cross-server, slow, 'RAID' configuration. You're better off with one server with dual PSUs, RAID and a backup (because even having fail-over redundancy is no replacement for backups - same as RAID). Edited by 3s-gtech
Posted

This is what we do, RAID is old technology and not really relevant in the days of SSD drives and virtual image backups.

 

We run 2 file servers with SSD's in them. One is used as a cold backup, that is restored to once a half term.

 

We also take daily virtual backups of the primary file server.

 

Probably overkill, but I like to be safe :)

  • Thanks 2
Posted
That just strikes as being a really expensive, cross-server, slow, 'RAID' configuration. You're better off with one server with dual PSUs, RAID and a backup (because even have fail-over redundancy is no replacement for backups - same as RAID).

 

Why though?

 

We do have a backup, and the fail over server would be located in a far away part of the school, should a disaster take out the main server room.

 

I'm not disagreeing with you, just wondering.

Posted

Because a single disk failure is far more likely than a 'disaster'. Yes, things like fires, cable cuts, power surges etc can happen but in 11 years I've lost quite a few disks in RAID (and never lost any data) but I've never lost a server (did have a RAID controller die, that was swapped out quickly and business resumed). It just doesn't make sense to me - you're protected against major disasters, but minor and more frequent issues cause the redundant system to kick in. While you're sorting out the first server, and the disk in the second gets a problem, things could go downhill fast.

 

Plus - single high capacity disks are slow ;) Big SSDs not so much, but I thought this was a cost saving exercise?

  • Thanks 2
Posted

I agree with @3s-gtech on this one.

 

In 15 years of managing servers I have never fully lost a server. I've lost plenty of disks but never a full server. I tend to use RAID with a hot spare, if a disk goes the spare kicks in and rebuilds and I swap out the duff one, no drama.

As stated above you are protecting against a major disaster (which is good) but not against a more likely minor fault (loss of disk) which completely takes your server out. If funds are short I would be looking at protecting against the most likely issues first.

  • Thanks 2
Posted
As budgets get tighter and tighter I can probably afford to put in place two single disk Hyper-v host servers, one as a replica for fail over, for only slightly more than it would cost for a single server with multiple disks, RAID controller, etc. VM's would be locally stored on this server.

 

The answer is going to depend rather on what sort of load you think you'll have on your servers, but in my experience in most schools the disk is more likly to be a performance bottleneck rather than the amount of raw processing power available. As such, you might want to go for the option where you spend more on a single server with multiple disks than on two servers. If you're using Hyper-V you should be able to use the built-in tiered storage facility (and also skip the need for a RAID controller), then you can basically use an SSD as a high-speed cache for a small array of cheaper large SATA drives, hopefully giving you both performance and capacity.

Posted
If it was me I would get one nice new server with raid and then use the old server for replication? I don't think I could live with the stress of losing an entire server just because one disk has failed.
Posted

I think you'd have to be crazy not use RAID on a production server.

 

Losing a server and having to rebuild the host, restore all services, down time and angry staff for outweigh the costs of couple of extra disks for a 'Redundant Array of Inexpensive Disks' in your servers.

 

As others have said, I've lost disks over the years, but never a server, it happens and probably will happen to you at some point, this is why we take proactive/preventative measures.

Posted

@dhicks response makes the most sense.

 

RAID is about providing fault tolerance when you have a disk issue like others here you expect disks to fail at some point (mechanical or otherwise)

having to swap out a disk is much easier than a whole server because of a disk issue (or recovering a VM).

a replica server to cover another server is going to cost more than a disk and raid.

 

If the school are expecting a resilient system to cover a disaster event they need to give you the budget to do it (IE fail over server etc)

Posted

The way I look at it is like this -

 

Implementing a Hyper-V arrangement, means you're saving money elsewhere having so many physical servers in your organisation (including the likes of disks), so why go to the whole expense of setting up a redundant/fail over solution, yet cut corners on the medium where data is stored? It makes absolutely no sense at all.

 

I'd rather have a single/robust Hyper-V solution with RAID and not have the fail over solution at all, rather than two physical hosts with one physical disks. It's not something I've ever encountered or considered as an option.

Posted
None in particular, I'm just theorising with the options.

Software RAID is a no no, since its far to slow.

 

Why do you think it's slow? Can easily do 3Gbps on a single core.

Posted
Why do you think it's slow? Can easily do 3Gbps on a single core.

 

Add that to the throughput required for say 3 x VM's and it slows things down.

 

Like I say, I'm not against any suggestion, I'm theorising and budget planning.

Posted

But as you already said, raid card costs money, needs its own backup battery, ram. It's just a slow cpu anyway, might as well spend the money on a slightly faster cpu, we're living in the future now.

 

Seems like my servers barely use their cpus anyway.

 

Mirrored hdds won't use much cpu at all, could do 2 8TB hdd drives in raid 1 + ssd cache?

 

Or tiered storage spaces in Windows? Or zfs?

 

You're going to use much more cpu, disk, network and electricity to run a HA failover vs software raid.

 

So, first question, how much storage do you need, next how many MB/sec, how many IOPs? SSD or HDD?

Posted

The other consideration of course is the overall throughput - RAID5 delivers better performance than a single drive, as of course the data is split across the array.

 

Software RAID could be an option, but running a Hyper-V configuration would typically be slow - even worse if running a database of some kind!

Posted

First...everyone...especially SLT and governors need to know an understand that IT systems can and will fail. And it doesn’t really matter how much you spend or how well it’s managed. A SAN between 2 hyperV servers with dual SAS controllers can and have been known to fail. And unless you have some sophisticated transaction logging for redundant databases you will lose data too....so office staff need to keep track of what they are doing between backups. Software and firmware updates can be lethal and unexpectedly end up with data loss on a live seemingly fully redundant system even when you thought you tested them first. The ONLY backstop is backups...and you might need some archived data because you might not know you have lost data from one day to the next...possibly not for a week or longer. So single backups - like virtual images - might get you back a working server, but not your data. Coreswitch failure can take your network down for a couple of days by the time you get a replacement and reprogram it...especially if it’s a different model...as you wouldn’t want to spend new money on an old model.

 

Once everyone is happy that you have backups ..and that there is the potential for ...possibly several days disruption...then focus on things that are likely to fail...and disks must surely be the least reliable component. I think a RAiD is a must...and I wouldn’t be without a dual psi...although I have never lost one on a live server. I shut a server down at half term .and a day later the PSU wouldn’t start...no one knew..but it served to remind me that you can’t be certain these things won’t let you down.

Posted
Could be worth looking at a "thin" hypervisor, for instance running ESXi or similar on a USB or SD card; you can make as many cheap copies of those as you like, and invest in a reasonable NAS box such as a Synology or Qnap connected to the hypervisor with iSCSI. Start with 3 or 4 drives in it, and you'll have enough performance and storage space on a good budget whilst not limiting future plans; it will be a piece of cake to add storage should you need to in future, just as easy to replace drives should one fail. Can't agree more with the post above; get everything outlined and make damned sure SLT know what is in store should something go wrong, give them some examples of failures and how long it would take recovery from a backup in each case.
Posted
None in particular, I'm just theorising with the options.

Software RAID is a no no, since its far to slow.

 

Software RAID is always faster on modern hardware.

 

You're comparing a relatively slow CPU on a RAID card to the full blown CPU of your server. Unless you have massive CPU bottlenecks already, which is a separate problem to address, software RAID is faster.

Posted

You have various aims,

1. Data integrity, once it's written, can I get it back, and be sure it's correct? If not, how many hours of data will I lose?

2. Speed: cpu, disk bandwidth/iops, network

3. Availability, if something fails how long is it acceptable to be out for? How long does a restore from backup take?

 

A system that needs to be up for 23 hours and 59 mins per day every day is completely different from one that can be down for 24 hours

 

You could have duplicate everything, generators, dual power inputs to the building, 5 ISPs, SSDs everywhere etc etc, if you were Google.

 

 

Not sure why a database would lose data though, everything has transaction logs and atomic commits, that's the whole point.

 

Whether you have local or network storage depends on the number of physical servers and VMs, and the bandwidth you want, of course your SAN/NAS is another point of failure, so do you buy 2 of those now? It's just another computer after all.

 

SATA 3 is 6Gbps, so a 5 disk array could do 30Gbps max, do you need one of those per server, or one shared on a 40Gbps SAN network? But if they're HDDs not SSDs then that's not an issue, you'd need 20 or more HDDs to saturate 40Gbps.

 

Tiers of storage seem to be the way to go, last years photos/videos move to the NAS with a few 8TB drives in, if it goes down no one really cares. Hourly backups to that too.

 

N servers with software mirrored ssds (used as boot and caching, large enough to store all active data for the day) and software mirrored 8TB hdds for normal use (scale for school size), the ability to run everything from N-1 servers in case one fails/needs maintenance.

Posted
The other consideration of course is the overall throughput - RAID5 delivers better performance than a single drive, as of course the data is split across the array.

 

Software RAID could be an option, but running a Hyper-V configuration would typically be slow - even worse if running a database of some kind!

 

In a round about way but only because you have more disks, the performance penalty is high because for every read or write operation there are 4 operations happening. Read the data, read the parity, write the data, write the parity.

 

RAID 5 should only be recommended for SSDs as the URE rate is phenomenally small. If HDDs look at RAID6 or RAID10, depending on your performance and storage needs.

Posted
Not just disk performance, how important rebuild performance is too. RAID10 rebuild performance is static and doesn't change as the amount of disks scales. RAID6 is slow and gets slower with more disks.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...