Jump to content

Kings College London SAN failure and data loss post mortem


Recommended Posts

Posted
I've been saying this for years and I'll keep saying it. No one needs a SAN. The only people that need sans are the SAN companies that want you to sell them to you.

 

What's your preferred solution if not a SAN or the cloud?

Posted
Wow, just wow! Makes me glad I am about to change mine! I have inherited a backup that backs up to a NAS and then creates an 'off site copy' to a USB3 HDD! This was 'best practice' the cowboys my charity used to use before I started!

 

What is the system you are implementing?

Posted
What is the system you are implementing?

 

Going to keep the NAS as nothing wrong with that. Debating between an LTO drive, which would require the tapes to be taken off site or another NAS offsite that I can replicate to. Think I will go with the other NAS, as I can install it in one of my shops and replicate across a VPN. Will also look at either Veeam or DPM, budget dependant

Posted
Can't believe such a large institution thought it was acceptable practice to have the backup and the live data on the same machine. Even in our school of only 300 pupils we've got a redundant server and a NAS backup (all in separate buildings) and I'm currently looking into an off-site option. Not sure if this will be a replicated NAS or some other system. However, if we've had enough damage at school to affect all three buildings, then data recovery won't be our highest priority!
Posted
What's your preferred solution if not a SAN or the cloud?

 

Local Storage ideally DAS via 12GB SAS HBA. Can also use newer tech like Windows Storage Spaces or VMware VSAN to create a cluster of storage servers with only local storage to leverage local storage in each host.

 

Either way I'd avoid the SAN as the moral of this story is not just about backup procedures is the fact that the SAN market is sold as "buy this and all your storage woes are over" which is clearly absolutely not the case. Who'd have thought such an expensive ultra reliable mission critical super duper SAN could be totally taken out by just a rogue firmware update? And what service do we get from these devices? Just the blame. Sorry, not blame, accountability.

Posted

I don't understand how they ended up with this situation - their setup just plain ignores any and all industry standard practices!

 

I mean, we have our main server data on 2 servers, redundant. We then have an on-site backup, in a different building, for the day to day backup/restore jobs, and then we have an off-site backup also.

 

So, to lose everything would take some serious effort!

Posted

The real shock to me was this :

 

"The IT team has grown from c. 115 members in 2014 to approximately 350 members."

 

350!

 

Sure they have more than 27,600 students (10,500 postgraduate students) and 6,800 employees while the organisation has an income of £684 million.

 

So they are ~15 times bigger than us in terms of students, ~20 time as many staff and they have ~40 times our income but their IT department is nearly 90 times our size.

 

This IT thing just doesn't scale!

Posted (edited)
The real shock to me was this :

"The IT team has grown from c. 115 members in 2014 to approximately 350 members."

350!

Sure they have more than 27,600 students (10,500 postgraduate students) and 6,800 employees while the organisation has an income of £684 million.

So they are ~15 times bigger than us in terms of students, ~20 time as many staff and they have ~40 times our income but their IT department is nearly 90 times our size.

This IT thing just doesn't scale!

 

The demands on IT in a university are much larger than a school though, so some extra is to be expected.

 

I mean, they are 23x the size of us in terms of students, and 27x the size in terms of staff, but 175x our IT support team... They just aren't comparable.

Edited by localzuk
Posted
The real shock to me was this :

 

"The IT team has grown from c. 115 members in 2014 to approximately 350 members."

 

350!

 

Sure they have more than 27,600 students (10,500 postgraduate students) and 6,800 employees while the organisation has an income of £684 million.

 

So they are ~15 times bigger than us in terms of students, ~20 time as many staff and they have ~40 times our income but their IT department is nearly 90 times our size.

 

This IT thing just doesn't scale!

 

I would hazard a guess that their IT has a lot more specialist components unique to various departments going far beyond what a normal school or college would have.

Many of the IT positions would be beyond just normal desktop support and dev positions found elsewhere.

They may also have a lot more it staff duplicating roles over their 5 campuses throughout London. Central support is fine, but you will probably still need an IT presence locally at each.

It's still a hell of a size though!

Posted
I would hazard a guess that their IT has a lot more specialist components unique to various departments going far beyond what a normal school or college would have. Many of the IT positions would be beyond just normal desktop support and dev positions found elsewhere.

Sure there will be a lot of specialisation - IME that tends to happen anyway; as you get larger teams it makes sense to narrow peoples focus and increase the depth of their expertise.

 

However, our support ratio is ~ 1:540, theirs is 1:100. That is quite a stark difference. I wonder how it compares to an industry of comparable size with specialist IT needs.

Posted

Don't forget our IT budgets in schools are well below the average across the IT industry given the number of users and devices, and our lack of staffing is symptomatic of that. Universities generally have proper IT funding. I bet they have multimillion yearly IT budget plus capital spending.

 

There's only so much we poor schools can do with £40-80k/year including staff!!

Posted

I think a more useful metric is budget:user with a stated student:staff ratio for context.

 

An obvious challenge when comparing schools to university is the blurring of student activities and staff.

 

It is clear that the team responsible for the fileserver backups did not appreciate that the shared drives and home folders were used by staff and students, and that some of the "students" work actually represented the product multi-million pound third party research grants.

 

One does have to also consider what activities the organisation is engaged in.

 

Take three examples:

1) a school that is teaching programming needs to provide services for students to have a development environment, and this dev environment needs to either be backed up, or quickly recreatable with user data being stored somewhere that is backed-up.

In a cloud-first environment this might need a physical server infrastructure to host student VMs.

 

2) A school that is very big on digital creative arts, with huge levels of video and photos. Again, requiring large shared drives/home folders on the LAN where all other data might be in the cloud.

 

If either of these facilities were added after the initial design of the current service, was the cost of appropriate backup of these new service factored into the cost?

 

3) A Uni Desktop PC lab support team. Initial service: shared PC labs, transient undergrad student data, some departmental shared drives for providing students with files. Everything is relatively transient. Most 'important' documents are duplicates of others stored on the big reliable systems (VMS / Solaris) run by the grey-beards, with their own backup and restore procedures. Over time, everyone moves to Windows, the 'grey-beards' leave and responsibility falls under the PC team - who do not have a culture of caring particularly for the user data. Nobody reviews the first principles by which the service was designed, and thus a gap between use and expectation forms.

Posted
Sure there will be a lot of specialisation - IME that tends to happen anyway; as you get larger teams it makes sense to narrow peoples focus and increase the depth of their expertise.

 

However, our support ratio is ~ 1:540, theirs is 1:100. That is quite a stark difference. I wonder how it compares to an industry of comparable size with specialist IT needs.

 

It depends on what their figures cover. There was a question that popped up on here not too long ago asking about Uni IT. The consensus was it depended on if you were core IT or faculty based. You may have an IT person paid for by a grant for a specific research project or a faculty may have an IT support person that has no access to network settings but can manage the PCs in that faculty. It also depends on if they are a single site or multi campus. If they are counting every person that fiddles with IT I could well believe it getting that high, for core IT that figure sounds high but, again, if they are in separate campi you may have a network team, helpdesk team, server team, desktop support team, communications team and then overall teams for each of those as well as management and project teams. Soon adds up.

Posted
I sort of wonder why, out of 350 staff, not one was on backup duty. You would think that a person could be employed just for that on account of the scale of it.
Posted
I sort of wonder why, out of 350 staff, not one was on backup duty. You would think that a person could be employed just for that on account of the scale of it.

I'm guessing this was part of the problem. Who ever was allocated to be on backup duty would be backing up what they knew about, but with the way the systems were setup and the fact they were moving from 1 solution to another, I'm guessing systems/areas that were missed were not picked up. If everyone else thought "penfold" is looking after the backups I don't need to worry about it, then systems missing from backups would not be identified. Similarly, the backups that were thought to be good were not. Testing of the backups had not been carried out, but is that really the job of the person who carries out the backups or do they need each System admin to run through restore/testing with the backup person to ensure backups are good. It really is not a single IT Staff user issue, as there could quite easily be missing information that has never been provided.

 

Speaking as someone who inherited a backup system which was overly complex in management (too many steps to log & change tapes) and too simple in job requests (a colleague would say "backup system-A" but provide no requirements for backup schedules) I can completely understand how something like this can happen. Trying to get others to take ownership of their systems and perform testing against the backups is hard enough for me, take it up towards the scale of this college and I can see it could be almost impossible.

 

Still, everything that is on the network should have some form of redundancy (backups or secondary system) and if this is not available then it needs to be highlighted to the Senior Management to make a decision about if it is an acceptable risk or if some investment is needed.

 

Although it has made me realise I need to be sending out some more emails regarding backups & testing again.

Posted
Another interesting thing I noted in the report is that the #5 in the diagram, "Other backup options" Slough Data Centre was unaffected / praised, which appears to have be using Veeam. This ties in with my findings over the years that Veeam is an outstanding piece of software that everyone should be using! :)
  • Thanks 1
Posted
Not naming names as we're not BTRD but when our new infrastructure was set up (outsourced, really annoys me as we could easily have done it ourselves) Veeam saved the backup of the SAN to a separate LUN on the same SAN. Admittedly it was then copied over to tape but first thing I did once we got control back was move the backups off the SAN and add in some disk storage in a remote cab to duplicate the backups to (as well as having the tapes).

 

This was offered to one place I worked as a solution during a server refresh... I politely asked them to stop being stupid and to quote for repurposing one of our outgoing servers with a stack of disks instead, because we could use standard WD red in that server rather than EMC's (i can only assume based on the price) gold plated disks it was actually cheaper.

Posted (edited)
Another interesting thing I noted in the report is that the #5 in the diagram, "Other backup options" Slough Data Centre was unaffected / praised, which appears to have be using Veeam. This ties in with my findings over the years that Veeam is an outstanding piece of software that everyone should be using! :)

 

Kind of. I got the impression that the VEEAM backup repository was supposed to be replicated to Slough, but for performance reasons this had been disabled. If it had been replicated then the loss of the Stand datacenter SAN would have been completely mitigated as they would have been able to fail over to it.

 

Indeed if the SAN had been fully replicated (as was the original design) storing the VEEAM repository on the same SAN would have been a perfectly acceptable design as it gives 3 and 1 of 321. Slough could then handle streaming it off to tape to give 'two mediums' element of best practice. Actually this would have given a 3-2-2 backup plan, which is in theory even more robust.

 

Compromises were made at various points during design and operations, and the interaction between these compromises were never properly explored or risk assessed.

 

 

EDIT: I should add that the VEEAM replication should be handled by VEEAM not the SAN because "storage replication is not a backup"

 

Edited by psydii
Posted
Indeed if the SAN had been fully replicated (as was the original design) storing the VEEAM repository on the same SAN would have been a perfectly acceptable design as it gives 3 and 1 of 321. Slough could then handle streaming it off to tape to give 'two mediums' element of best practice. Actually this would have given a 3-2-2 backup plan, which is in theory even more robust.

If you store data on the same logical system you are at risk from design/code flaws locking you out of both the data and the backups. An example might be a date sensitive algorithm - so you pass an arbitrary date and bang - you have no access to any part of the system because all nodes that you have replicated to are subject to the same defect. These types of failures are rare but do occur. So IMO it is never a good idea to backup system states to the same system and unless you ignore the possibility of such failures, it cannot be said to be a secure copy.

Posted
After reading all of this (and the report) I'm glad so glad we use SCDPM to backup to disk & tape along with critical servers (like Exchange, Data stores, a DC and SIMS servers) replicated to our old servers in another building! The other week when we had a power cut that zapped one of our storage devices, the affected handful of servers were back online on our replication servers within minutes of us failing them over! :) We also test both disk and tape backups of random servers each and every month to ensure that the backups SCDPM creates actually work! :)
  • Thanks 2
Posted

I think some key points are.

 

1. Make someone responsible for backups including if applicable tape changing/rotation.

2. Make sure alerts for backups are sent to responsible people/team and applicable management. Alerts should be sent whether jobs are success or failure.

3. Make sure if the person who has been made responsible is not available they pass the task on to another capable/responsible person.

4. Document backup, restore procedures and any other relevant backup system setup i.e. Tape rotation, labelling scheme etc... .

5. 3-2-1 , Original source, Backed up data (on different storage) and off site/remote site backups for DR.

6. Test backups regularly. Veeam can do this in a Virtual lab.

 

IMO It doesn't matter what backup software is used as long as it is as reliable. We each have our favourites which may be dictated by available budget, existing organisation setup or organisational need.

  • Thanks 1
Posted

@Davit2005 - One thing to add to that. I get alerts for all the backup jobs but due to amount & other notifications I get I put a filter on the mail coming in so that it automatically moved backup reports to a backup folder. Any that failed stayed in my inbox. This has helped me keep on top of things as alerts for jobs failed/warnings were missed initially due to the sheer volume.

 

It's also on my daily task to check the backups. I have set up a daily email to be sent to spiceworks for Daily tasks. Using ticket rules it applies a checklist to the ticket which I use to ensure I have checked everything for the day which is the first job on my list. Should I be off (holiday or sickness) it allows someone else to go through the daily check list and ensure things are still being done. This is something we have only done recently but find it much easier to ensure things are checked as it's too easy to miss things when we are busy.

  • Thanks 1
Posted
I have set up a daily email to be sent to spiceworks for Daily tasks. Using ticket rules it applies a checklist to the ticket which I use to ensure I have checked everything for the day which is the first job on my list. Should I be off (holiday or sickness) it allows someone else to go through the daily check list and ensure things are still being done. This is something we have only done recently but find it much easier to ensure things are checked as it's too easy to miss things when we are busy.

 

This hadn't ever occurred to me, this sounds like a brilliant idea. How did you set the email up?

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...