Jump to content

Recommended Posts

Posted

I don't, but my friend has started using it. 3 nodes with 2 40Gbps network cards in each to form a loop. With 4 nodes you'd probably want a switch. 100Gb switch with 2 100Gb NICs per server, so each write can write to two other servers at 10GB/s speeds, which gives you a real world write speed of about 6GB/s

 

Works with any hardware at all, I tested it on 3 desktop machines with 100Mbps networking and it's slow, but works fine.

 

Optimising the speed is the complex part.

 

Not sure it's really worth it for the ability to restart a VM in seconds if the hardware fails. Hardware doesn't fail that often, and server class computers make that even less likely.

 

My own solution is hourly zfs syncs to a backup computer, so at most 1hr of data is lost, Few hours to restore to brand new hardware.

Posted
It boils down to risk analysis I suppose. I wont do anything other than HA. Ive lived through a hyper-v failure plus restore to new hardware (after waiting for the new hardware) and do not wish to endure that ever again. I also only use VEEAM as it works and I have faith with it. With my stretch cluster a node failing or one of the SAN failing isnt an issue. The hardware is aging though and the system is overkill for a full replacement (it was previously incrementally updated) hence a 3 node HCI would be a better (resilient) option. I even HA our pfsense firewall as the hardware is cheap but we are resilient. Sourcing new hardware timely might be an issue and if you already have a cold server then why not HCI it? For us it is definitely worth it to have a VM down for only a minute but also not losing the delta of data since the last backup. My head would be on the block if I lost an hour of MIS data or academic reports etc "by choice" rather than necessity, I have continuous delta SQL backup to a NAS but only half daily backups of other stuff. Yes we use sharepoint and onedrive for a lot of things but quite a bit is still onsite due to cost effectiveness - our media is saved locally as it is unworkable to have our drama, art, music, DT and engineering have their gigantic files working via sharepoint and onedrive. Yes I have costed three lumps of iron onsite + local backup vs SaaS. Going OT though.
Posted

Of course, cost vs risk, problem is each extra increase in 9s of uptime at least doubles the cost. What are they willing to pay for.

 

SQL data is easier to sync in real time by its nature

 

Topics aren't a thing without hierarchical reply systems.

 

How often has a node or san failed?

 

3 node seems to be the easy option, don't need a 100Gb switch, just DAC A to B, B to C, and C to D

Posted
SAN? I have never had a SAN go down totally, I have had a number of drives fail (in RAID configuration) and a management card die in an MD3220, but they are replicated so it was a simply yank out, put new one it - no downtime. Nodes? I had a hyper-v host R610 raid controller fail - the BBU went bad and took the RAID controller with it, magic smoke escaping and all. At the time we had a replicated hyper-v setup (local storage), not a cluster, and the replication failed when starting up - it had previously been fine when manually brought online for hypervisor updating plus a simulated fail had also worked. This ended up with around 8Tb of restoration which took a couple of days to get running. We had (and still have) a standalone physical DC so no rollbacks or AD sync issues. Since ive had the stretch cluster we had the UPS management card shutdown the cluster (power fault on one site) and one of the nodes did not wake up again, there was a motherboard fault (an R630, fans would simply spool to maximum, very noisy!). It should be noted that our "second site" nodes on the stretch cluster has no votes so cannot continue without manual intervention, this prevents split brain. The rest of the cluster came on just fine and no loss of service was seen (other than the time to reboot), this was all automatic and happened over a weekend. More recently we had a memory fault on an R740, we lost half of the memory, the poor thing carried on running, the cluster rebalanced automatically and the show carried on running, not a full fail though just not enough RAM to run the whole VM load - it would have been interesting on a single host system. Cluster Aware Updating works in 2019 (it didnt really work in 2016, it was always better to manually drain) so updating the hypervisors is a doddle.
Posted

Yeah, I always had trouble with hyper-v replication, would just stop for no reason. I can see why you're slightly paranoid though.

 

I've had a raid card die, don't trust them, single point of failure. Software raid all the way.

 

You'd probably miss some features you have if you moved to Proxmox, automatic cluster VM balancing requires a 3rd party script for example.

Posted (edited)

Ive had a look at proxmox and ceph. It does seem to do what I want it to and im downsizing our server load to 3 nodes which is the minimum. Hardware will arrive in September for lab "test and soak" before I migrate the VMs across, the hardware is S2D certified as im still on the fence with regards to

 

*S2D + full fat server + hyper-v cluster. Familiarity, im happy with windows clustering, S2D is new to me but the concept is similar. Im happy with powershell configuration (it was better to do this for stretch clustering anyway). Licensing is 3x EES DC licenses and covers cluster + S2D

*starwind + full fat server + hyper-v. Starwind is new to me, I would be running the free version simply because the paid one is massively expensive even in education licensing. I would still need DC for the cluster due to VM count.

*proxmox+ceph (with windows server guests). I have a good few months to play with the new hardware in lab configuration first. I would still need DC licensing for the VM count.

 

Ive been playing with starwind in a set of VMs but have hit quite a few gotchas. No CHAP on the free version. Not a total game changer as I will be running in direct connect between the servers (40Gb interconnect between the 3 nodes) so wont need CHAP really, but if you are traversing switches (for whatever reason - perhaps you have a backup volume presented?) then you wont have CHAP! Caching + HA means a full resync of all data to unsync nodes, that should be fun with a pair of 8Tb volumes and pretty much puts your 3 node HA into a no redundancy until this completes. S2D flushes the sync cache then syncs delta changes, it does not full resync, so whilst again you are in a single node redundancy it isnt for nearly as long (minutes compared to hours/days possibly - I have yet to perform this on the real hardware). However, you can use hardware RAID for extra drive redundancy with starwind thus having a bit more resiliency for hardware failure, S2D can only run in HBA plus nested resiliency (mirror accelerated parity for example) is only available in 2 node configuration. Im yet to even start lab work on proxmox+ceph, this will be done on real hardware in the lab.

 

Im not a fan of having a single VM perform all tasks hence DC licensing, it is cheap enough for schools to warrant such extravagance. Updating numerous VMs isnt an issue, they have been automated updating for years with scripts (always after the weekly archive backup!)

Edited by KK20
  • 4 months later...
Posted
I have tried 3 times to sign up to chest. Not once have they gotten back to me, my emails have gone unanswered. I even check 365 logs to make sure 365 isnt eating them. Ive tried with our .sch.uk and our .co.uk without issue. We are a genuine secondary school!

 

I remember registering online, paying the membership fees and calling them to confirm it and then proceeded with the actual order.

 

https://www.chest.ac.uk/registering-on-the-chest-website/

 

Phoenix Software gave us the heads up on this, i believe there where like 8 companies that you could choose from.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...