Jump to content

Kings College London SAN failure and data loss post mortem


Recommended Posts

Posted (edited)

Following the loss of their SAN and then the discovery that their backups were not viable, here is the report.

 

https://regmedia.co.uk/2017/02/23/kcl_external_review.pdf

 

Every one here should read it.

 

Your managers must read it.

 

 

It does seem to put the blame on the technical team. What I would say is that if the technical team do not understand the system built by others sufficiently to manage the backups, it is a Management failing. If managers are not asking to see backup schedules logs and compliance against documented procedure on a Monthly basis, they cannot escape blame.

 

I have heard (down the pub with someone middlingly senior but not IT, so take with a pinch of salt) that large chunks of this system were outsourced, so depending on the particular element, the IT Technical team were not necessarily either the architects of the solution nor employees of KCL.

Edited by psydii
  • Thanks 2
Posted
They kept the main (complete) backups of the data on the SAN on the same SAN. Mind=blown. Yes, there were a catalogue of failures, but whoever decided to do that needs beating vigorously with a cluebat.
  • Thanks 1
Posted

People who haven't been given the budget to do 3-2-1 properly.

 

They made compromises that *should* have been fine, if the next system along wasn't also compromised in some fashion.

 

Nobody evaluated each decision in the wider context, and Senior Management were not engaged sufficiently to understand the risk IT were taking with the data.

Posted
Who stores their backups onto the same hardware, they played with fire and got massively burnt.

Not naming names as we're not BTRD but when our new infrastructure was set up (outsourced, really annoys me as we could easily have done it ourselves) Veeam saved the backup of the SAN to a separate LUN on the same SAN. Admittedly it was then copied over to tape but first thing I did once we got control back was move the backups off the SAN and add in some disk storage in a remote cab to duplicate the backups to (as well as having the tapes).

Posted
Not naming names as we're not BTRD but when our new infrastructure was set up (outsourced, really annoys me as we could easily have done it ourselves) Veeam saved the backup of the SAN to a separate LUN on the same SAN. Admittedly it was then copied over to tape but first thing I did once we got control back was move the backups off the SAN and add in some disk storage in a remote cab to duplicate the backups to (as well as having the tapes).

 

That's good you've done that, and this has only highlighted your cause to put in a more resilient backup solution.

Posted

3 out of 4 places I've worked in have had issues with backups for one reason or another, either budget, technical or jus plain ignorance. 2 had been lucky that no incidents had arisen, 1 lost data (, but not business critical data) after a SAN failure. As a field engineer many years ago I had to go on site to have a look at a server which had lost it's RAID controller and all the data with it, no backups and had to deal with a company who dialled in and set the backups to the local C drive where the original data was for troubleshooting but not remembered to set it back costing them £1000 in emergency data recovery for the customer.

 

Lesson learnt, unfortunately the hard way :getmecoat: .

  • Thanks 1
Posted
I think they are referring to the leadership here. The first mention is the IT Leadership team, then for the rest of the document it's the IT team, until page 17, where it specifies that their are multiple IT Teams. This makes me think they are talking about the IT Leadership team in particular.
Posted

The main thing that I get from this article is that if a major UK university has this sort of problem then the best solution for schools is a full scale migration of all data to a decent service provider!

I couldn't imagine presiding over such a huge data loss when I simply don't have the staff to patch every firmware update on SAN/Backup

Posted

I've been saying this for years and I'll keep saying it. No one needs a SAN. The only people that need sans are the SAN companies that want you to sell them to you.

 

Wow the report contains so many things I would like to comment on but in this day and age of litigation etc I don't think I'll bother. Guys, if you've spent such massive amounts on the SAN system that you can no longer afford to instate appropriate backup procedures - you realty need to have a long hard think.

 

As this episode has shown - all it takes is one little firmware bug - and all of a sudden the original promise of your "ultra reliable totally redundant SAN" suddenly turns into "catastrophic data loss". Complexity is the enemy here.

Posted
I've been saying this for years and I'll keep saying it. No one needs a SAN. The only people that need sans are the SAN companies that want you to sell them to you.

 

Wow the report contains so many things I would like to comment on but in this day and age of litigation etc I don't think I'll bother. Guys, if you've spent such massive amounts on the SAN system that you can no longer afford to instate appropriate backup procedures - you realty need to have a long hard think.

 

As this episode has shown - all it takes is one little firmware bug - and all of a sudden the original promise of your "ultra reliable totally redundant SAN" suddenly turns into "catastrophic data loss". Complexity is the enemy here.

 

I think that this is a good time to show a commercial for the SAN in question:

 

Posted
I've been saying this for years and I'll keep saying it. No one needs a SAN. The only people that need sans are the SAN companies that want you to sell them to you.

 

Wow the report contains so many things I would like to comment on but in this day and age of litigation etc I don't think I'll bother. Guys, if you've spent such massive amounts on the SAN system that you can no longer afford to instate appropriate backup procedures - you realty need to have a long hard think.

 

As this episode has shown - all it takes is one little firmware bug - and all of a sudden the original promise of your "ultra reliable totally redundant SAN" suddenly turns into "catastrophic data loss". Complexity is the enemy here.

 

Couldn't agree more. I think even the largest schools should be moving to Google/Azure/AWS now.

Posted
Following the loss of their SAN and then the discovery that their backups were not viable, here is the report.

 

https://regmedia.co.uk/2017/02/23/kcl_external_review.pdf

 

Every one here should read it.

 

Your managers must read it.

 

 

It does seem to put the blame on the technical team. What I would say is that if the technical team do not understand the system built by others sufficiently to manage the backups, it is a Management failing. If managers are not asking to see backup schedules logs and compliance against documented procedure on a Monthly basis, they cannot escape blame.

 

I have heard (down the pub with someone middlingly senior but not IT, so take with a pinch of salt) that large chunks of this system were outsourced, so depending on the particular element, the IT Technical team were not necessarily either the architects of the solution nor employees of KCL.

 

Thanks for sharing an interesting article....focuses the mind.

Posted
Couldn't agree more. I think even the largest schools should be moving to Google/Azure/AWS now.

 

Not so sure I agree!!

 

It's still riskly as yet again, your devolving reliability from your own team to a third party that promises "everything will all be ok". So you'd better hope Google/Azure/AWS don't have any firmware issues or anything catastrophic happen! You'd also better hope that if and when it does happen they have adequate backup procedures! Do they? You'll never find out - until something goes horribly wrong, then you'll be able to learn about it in the report written about your departments failure, see first post.

Posted

Having read the report there are several things in there that are happening in many schools right now, I'm not talking about the hardware I'm talking about the staff attitude to IT in general (mgmt., IT and users). A recurring theme in the report is that of users not knowing where to store data, IT staff not having enough time/budget to do things properly, management not being aware of the overall risk to the organisation. I bet most schools would not be able to read that report and say "yeh we're fine we're doing it all right".

 

It's sober reading, I've noted several improvements I can make to my systems, of course those will still depend on money and time being made available.

  • Thanks 1
Posted

Wow what a mess. I can't help but feel sorry for the IT team. Yes they made some mistakes but reading between the lines it sounds like they had a bunch of new guys in that were given very little training or hand over. Then they had to deal with a system that was bodged together over the course of a few years meaning that the people setting things up may not have even still been there.

 

Also, what the hell HP! Your engineers will replace a part that you know will tank the system (a system that is sold as being super reliable) but it's not your fault because the customer only had "proactive support". Surely proactive support should include checking that firmware has been applied to fix a known issue with the part you are about to replace???

 

The thing that really bugged me is that the IT team is getting the blame for people storing vital data in non-backed up areas. Do you all think it is IT's job to check what data is or isn't important for every single person or is it the persons job to check with IT that the data they know is vital is being kept secure? I'm struggling to get just 24 staff members to go though our data here and tell me what is and isn't important, can't imagine how hard it would be at a much larger and diverse organisation such as that. Yes the data storage system should have been better explained to users so it was clear what was and wasn't backed up but honestly I doubt most of them would have paid attention anyway.

 

Also these:

• IT do not understand the needs of all the different user groups and are perceived to lack empathy with some.• The IT culture is not aligned with user expectations hence they disengage.

 

I'm sure the majority of you have had to deal with trying to get information from users about what their requirements are and also have to deal with users who's "expectations" are completely impossible to work with due to budget or technical limitations. I think this is pretty unfair.

Posted

The thing that really bugged me is that the IT team is getting the blame for people storing vital data in non-backed up areas. Do you all think it is IT's job to check what data is or isn't important for every single person or is it the persons job to check with IT that the data they know is vital is being kept secure?

 

IT offer network drives to end users. The end users saved work to the network drives. It is reasonable for the end users to assume that the network drives are backed up.

 

Literally every KCL employe with a desk saved **ALL** their work to the network drives.

 

IT not knowing how they were being used is not an excuse. They should have been backed up 321. Decisions to skip backups or reduce recoverability should have passed through senior management for approval and from there effectively communicated to Departments.

 

 

Do not offer a service to users unless you can do a bare metal restore of any work they store within the service.

Posted

It highlights that tape is still king when it comes to restoring data. Data on disk for convenience but also tape for DR situation. The one key area that I picked out of the that is that not enough testing had gone on. As the person that is primarily responsible for the backups here it is my main gripe that we don't do enough testing. I do what I can but I don't know all of the systems so can only do so much. Systems that are managed by other members of staff don't get tested at all and I would much rather pass the backups for those systems to them....I can only imagine if we lost a system here the blame would be pushed directly to me even though all I do is effectively run the job.

 

It just confirms backups are one thing that often gets forgotten about as no one seems the benefit of it...until something goes wrong.

  • Thanks 1
Posted
IT offer network drives to end users. The end users saved work to the network drives. It is reasonable for the end users to assume that the network drives are backed up.

Literally every KCL employe with a desk saved **ALL** their work to the network drives.

IT not knowing how they were being used is not an excuse. They should have been backed up 321. Decisions to skip backups or reduce recoverability should have passed through senior management for approval and from there effectively communicated to Departments.

Do not offer a service to users unless you can do a bare metal restore of any work they store within the service.

 

I 100% agree. The problem here is that the way I read it the IT team (i.e. the techs, not the management) are the ones taking the fall for this not happening when they did not have the training or resources to do any different. As you said decisions to skip backups etc should go through senior management and it is their responsibility to make sure the data is being handled correctly.

 

The way I read this though is that it was all the tech's fault and they should know how the data is being used. Baring the guy that was lying about the tape backups I'm not sure what else most of these techs were supposed to do? If they are told by their superiors "we only back up X, Y and Z as we don't have capacity to do more" what else are they supposed to do? Without knowing exactly what the working conditions were like maybe I shouldn't comment but from the way I read the report it sounds like they had a lot of new hires, no decent hand over or training and a system that had been patched together over time with no real oversight (which are all management problems IMO).

 

You are right though it shouldn't be down to the user to check what is backed up and it is fair that they assume it all is unless clearly told otherwise. Just the way some of it was worded got my back up a bit and I get the impression that a lot of the blame was being dumped on techs that were probably under trained and had no ability to change anything anyway.

Posted
Also, what the hell HP! Your engineers will replace a part that you know will tank the system (a system that is sold as being super reliable) but it's not your fault because the customer only had "proactive support". Surely proactive support should include checking that firmware has been applied to fix a known issue with the part you are about to replace???

 

Storing backups on the same SAN as the original data is bad. But this ^ is the kick in the balls. Report states the firmware was out for weeks - if it was out for weeks from say 1st Jan I wouldn't be scheduling in the upgrade until the half term anyway as I'm not able to be onsite out of hours when sufficient downtime can be planned for during term time.

 

And the they sensibly bought enhanced support when buying the 3PAR but did HP come to them in 2015 and said we now do this better support, give us money. Bet they didn't.

 

The rest of it, ok get your backups in order but they haven't put a fair enough amount of blame at HP's door in my opinion.

Posted
You would expect even with the cheapest lowest level of support that they would check the firmware version before swapping parts, at a guess the new part was probably running the latest firmware and the rest of the kit was on old firmware so it got upset. Think HP need to go back to school and learn what the word "proactive" means.
Posted
Wow, just wow! Makes me glad I am about to change mine! I have inherited a backup that backs up to a NAS and then creates an 'off site copy' to a USB3 HDD! This was 'best practice' the cowboys my charity used to use before I started!
Posted (edited)
Sounds like someone should have got a Nimble. I hate backups that don't make sense like backup exec why don't backup file names match the job so you can move or delete and have a clue what you deleted. Edited by nicholab
  • Thanks 1

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...