Jump to content

Recommended Posts

Posted

Hi all,

 

So I'm just getting my feet under the table as a newly minted secondary school NM and have inherited a NetApp SAN which is 6 months out of support, the production disk shelf is turned OFF, the DR disk shelf is powered on and seems to be running the show and there's NO documentation regarding access to the admin panel to see just what is going on with it.

 

Good times.

 

So, the school have already migrated their main data storage to Google Workspace and MIS has been migrated to Hosted SIMS so that helps ever so slightly. The former Senior Tech had been looking into a like-for-like replacement from various vendors but looking at the likely quantities of on-premise data and the remaining workload (effectively DCs, Papercut, Net2 and MDT) I think it's time to scale down and continue the march cloudwards; in 2017 everything was on-prem so I could see the logic from that point of view. My instinct is to cut the SAN from the equation completely, return to on-host storage (2x new Hyper-V hosts with full SSD storage) and focus on the creaking infrastructure across site.

 

Thoughts? :o

Posted
I would look to migrate to something like VMware VSAN or Microsoft S2D (Hyper-V Route) and get rid of the physical SAN. Replace your hosts with all SSD or a hybrid HDD / SSD cache.
Posted
We had a Fibre Channel SAN at the last place but had over 150 VMs, multiple blade chassis's and redundancy but hyper converged or other solutions mentioned above might suit the business as well and remove complexity.
Posted
I'd recommend going down the S2D route as well, we moved from a traditional "VM servers and separate storage" setup recently, and now have a hyperconverged setup with Hyper-V and S2D on 3 servers. Works great. No need for an external SAN, one less potential bottleneck or single point of failure. My recommendation would be to totally forgo "new" servers. Get 3 refurbished servers from ICT Direct, slap new SSDs in them (ensuring you have a storage controller with the right capabilities) and set them up as I did. We bought 3 servers with dual 28 core CPUs, 192GB RAM in each, and 4 10GbE ports, plus 4x 4TB enterprise SSDs, and HBA modules for the servers, all with 3 year warranties, for less than £10k. More than enough to run our entire Trust.
  • Thanks 4
Posted
Unless you have a really specific need , I don’t see the need for a SAN in a edu setting . As others have mentioned s2d is a good solution
Posted (edited)

Same as others have said, S2D all the way...

 

We went for a new 2 node setup on NVMe for local storage and SSD for the CSV storage with 5 year warranty from Dell with 24k for assisted install (was £400 cheaper with assisted install, #shrug), that was 3 years ago now.

 

We then overhauled our network to 10gb with POE at every cab with Ubiquiti Unifi switches for £60k including cabling costs and then another 10k for a WiFi replacement to WiFi 5 gen 2, that was last year.

 

This year I replaced our 12 year old backup server and 2 NAS boxes with a Dell optiplex on NVMe local with 10k SAS disks and a 35tb Wasabi plan and veeam licensing for 15k all in, the licensing is 3 years on Veeam and Wasabi.

 

We now have 4u worth of physical server on site, go back 5 years ago this was 2 full server racks of servers and sans. You have to spend money to save money.

 

The energy consumption, carbon footprint, security, backup/restore RTO's and RPO's.

 

During all of these I got to play with and learn all the new kit, install it, configure it and tweak it.

 

I can drain a host, migrate vm's, patch the host and have everything moved back in less than 30 minutes during a live day with no impact.

I also watched my backups tonight take 45 mins on-prem for 5tb of data and then offload it to the cloud within an hour, all happening at the same time so within an hour I have on-prem, off-site and immutable for 30 days with off-site archiving for 5 years.

The way technology now is crazy, monitoring my network and servers from mobile apps at home on my iPad or phone while on the sofa and making changes with 3 taps.

 

We then use Office365 for usual bits, teachers and SLT starting to see the investment in SharePoint, OneDrive and Teams more to make their lives easier... Sims going the way of Nextgen in the next X years.

 

Next step for me is looking to modernize the desktop infrastructure.

Edited by Tefters
Posted
For our three recent clusters replacements SAN came in cheaper than S2D believe it or not. Going from what we had to an all-flash SAN was still a big performance boost so were happy either way.
Posted (edited)

Take a look at StarWind VSAN. It's more flexible than S2D. There's a paid support avenue if you want it, or you can just use the free version (up to 3 nodes). Essentially the same thing as S2D, vmware vsan etc but with more wiggle room and a cheaper support option if you want it vs the other two's support options.

 

https://www.starwindsoftware.com/starwind-virtual-san

 

I use the paid version at the moment with two nodes.

Edited by mrbios
Posted
Some Edu settings might warrant a SAN but we (Uni) are moving to a hyper converged setup now using Greenlake. My last org was a college with over 150 VMs and had a san with multiple hosts/chassis and even our Citrix VDI shared the same SAN.
  • 4 weeks later...
Posted

Sorry for being ignorant and old but what is S2D?? an option to share drives across servers but they are located on Azure?

I just had our SAN crap its pants, 2 drives seem to have died at exactly the same time and despite having 1 more hot swap it caused all of our servers to crash and display a data write error until I manually went into the vhosts and rebooted them.

 

Luckily as we are cloud based a lot it didn't affect too many people, 5 years ago it would have been mayhem.

But each drive is approaching £500 and the unit is out of warranty so replacing it will be £££. So I am wondering what to do now.

 

What is S2D??? storing all this data on a virtual drive in Azure? what is the latency like? I noticed when onboarding our e-mail to o365, Outlook will hang now when you attach a 20mb+ attachment (yes I know not good practice) so I can imagine user frustration with large files as our LEA link is pretty poor quality.

Or am I missing the point on what it is/

Posted (edited)
Sorry for being ignorant and old but what is S2D?? an option to share drives across servers but they are located on Azure?

I just had our SAN crap its pants, 2 drives seem to have died at exactly the same time and despite having 1 more hot swap it caused all of our servers to crash and display a data write error until I manually went into the vhosts and rebooted them.

 

Luckily as we are cloud based a lot it didn't affect too many people, 5 years ago it would have been mayhem.

But each drive is approaching £500 and the unit is out of warranty so replacing it will be £££. So I am wondering what to do now.

 

What is S2D??? storing all this data on a virtual drive in Azure? what is the latency like? I noticed when onboarding our e-mail to o365, Outlook will hang now when you attach a 20mb+ attachment (yes I know not good practice) so I can imagine user frustration with large files as our LEA link is pretty poor quality.

Or am I missing the point on what it is/

 

S2D is where the storage is local to the host, so each host has quantity of disks your san usually would have (unless you get bigger disks, you can cache or not cache) and it creates a “software raid” on each node.

 

So effectively the storage is right there next to the cpu/ram so its much quicker than a SAN and also if a node fails all the data is on each node so you continue to operate, where as in the very unlikely event a SAN entirely dies you would be out of service.

Edited by CrootUK
Posted

I am aware that many members on this site advocate for hyper-converged systems like S2D (Storage spaces direct), but I wonder whether anyone else experiences the same problems as we do.

In July 2021, we switched to a two-node S2D setup from an existing 3-server + san setup using VMware. What I've discovered is that every time we perform a host update, replication on both servers undergoes a storage check when each server is restarted.

The 18TB of SSD storage we have on each host seems to take up to two hours to finish this. It makes the virtual machines using the storage run slowly. I typically shut down any virtual computers that are not absolutely necessary to speed things up. Is this typical for S2D? I must admit that I miss hardware raid on a SAN and being able to restart a server without needing to verify the surplus storage.

Posted
I am aware that many members on this site advocate for hyper-converged systems like S2D (Storage spaces direct), but I wonder whether anyone else experiences the same problems as we do.

In July 2021, we switched to a two-node S2D setup from an existing 3-server + san setup using VMware. What I've discovered is that every time we perform a host update, replication on both servers undergoes a storage check when each server is restarted.

The 18TB of SSD storage we have on each host seems to take up to two hours to finish this. It makes the virtual machines using the storage run slowly. I typically shut down any virtual computers that are not absolutely necessary to speed things up. Is this typical for S2D? I must admit that I miss hardware raid on a SAN and being able to restart a server without needing to verify the surplus storage.

 

What connectivity do you have between the hosts? Are they direct? 40gb? Rdma?

 

Is it one big volume or multiple volumes?

Posted (edited)
Funny enough I have moved on to a new venture in the time I posted this. It has 2 x 25gb qsfp direct connection between hosts and 2x10gb sfp connections to the core switches. 1 20tb big volume, with a mix of nvme and SSD storage. Edited by jslate1980
Posted
Funny enough I have moved on to a new venture in the time I posted this. It has 2 x 25gb qsfp direct connection between hosts and 2x10gb sfp connections to the core switches. 1 20tb big volume, with a mix of nvme and SSD storage.
Does it do the same as your old one?
Posted
Does it do the same as your old one?

 

Nope moved to a group of colleges and looking after 3 data centre’s, with 2 x vsan clusters and 1 x traditional VMware ISCSI san solution.

  • Thanks 1
Posted
I am aware that many members on this site advocate for hyper-converged systems like S2D (Storage spaces direct), but I wonder whether anyone else experiences the same problems as we do.

In July 2021, we switched to a two-node S2D setup from an existing 3-server + san setup using VMware. What I've discovered is that every time we perform a host update, replication on both servers undergoes a storage check when each server is restarted.

The 18TB of SSD storage we have on each host seems to take up to two hours to finish this. It makes the virtual machines using the storage run slowly. I typically shut down any virtual computers that are not absolutely necessary to speed things up. Is this typical for S2D? I must admit that I miss hardware raid on a SAN and being able to restart a server without needing to verify the surplus storage.

That's odd. If checks are happening like this, I've not noticed them causing any slow-downs on our setup.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...