Jump to content

Recommended Posts

Posted

I'm often told we don't shout about our achievements enough.

 

For example we had a disk failure in the SAN at the weekend, all covered by resilient systems and resolved by Tuesday BUT our SLT are none the wiser

 

Do any of you complete an IT Incident Report (form?) that goes to maybe your SLT that effectively demonstrates or records all the incidents we deal with of say at least a severe level (e.g. likely to impact T&L)

That the SLT would otherwise be completely unaware of!

 

I guess I'm either looking at an MS form or a word template?

 

Regards

 

Steve

 

We have a ticketing system, this is more management reporting similar to a cyber incident action report etc

Posted (edited)

I don't do a form but I do send after action emails if we had an outage or incident. Normally only if it has affected operations, but something like weekend work would definitely get one to highlight it was done. Nothing grandiose, just the facts with timings and work carried out.

 

EDIT: If your help desk is ITIL based it should have a incident section that can log this.

Edited by TechMonkey
Posted
we do a postmortem of every incident, the process is described here: https://sre.google/sre-book/postmortem-culture/

 

I wouldn't class that as an incident though - A disk failure is expected behavior and you already mitigated against it.

While you are right... the context is a school. SLT need to be reminded that the money spent by IT *actually* prevented an outage today. Otherwise they forget that resilient and reliable IT requires ongoing investment and next year slash your budget and hire another teacher.

  • Thanks 3
Posted
While you are right... the context is a school. SLT need to be reminded that the money spent by IT *actually* prevented an outage today. Otherwise they forget that resilient and reliable IT requires ongoing investment and next year slash your budget and hire another teacher.

 

Where would you draw the line though? Systems are designed to be resilient and failover. You could end up just reporting everything as an incident and not focussing on where money needs to be spent.

Posted
Where would you draw the line though? Systems are designed to be resilient and failover. You could end up just reporting everything as an incident and not focussing on where money needs to be spent.
It really depends on the attitude and level of technical knowledge of the management.

 

In a school, I had to fight to replace a failed disk in a raid array in because the server was due to be replaced in four months and in another larger organisation where the server infrastructure had become horrendously outdated, the management intention was to get rid of the people who'd fixed it as soon as possible. They virtualized everything and centralised hardware so that it was less visible, so management assumed there was almost nothing left to do.

Posted

At the company I work for, every incident is a ticket then an email is sent to those in the need to know and involved.

 

The email basically says 'who, what, where and resolution' with a link to the ticket for a more technical view. Also any questions/suggestions are then put in that ticket rather than the email.

 

Also there is a culture here to promote what you have done that you feel should be shouted about.

 

I recently tidied up the 12 meetings rooms we had so every room had Google Meet, Apple tv and a tv instead of a projector. I also made sure that there was tips on how to use it or simple trouble shooting. I didn't say anything about it, just got on with it as a job and I got 'told off' for not telling the company that they now are all identical etc.

  • Thanks 1
Posted
It really depends on the attitude and level of technical knowledge of the management.

 

In a school, I had to fight to replace a failed disk in a raid array in because the server was due to be replaced in four months and in another larger organisation where the server infrastructure had become horrendously outdated, the management intention was to get rid of the people who'd fixed it as soon as possible. They virtualized everything and centralised hardware so that it was less visible, so management assumed there was almost nothing left to do.

 

Perhaps I was lucky then. I would put forward a budget plan at the beginning of the year and spend the money as I saw fit. Never needed to argue over infrastructure as it was (rightly) seen as a vital component of teaching and learning. A thing like a single server disk drive would never have been an issue to fund. Our SLT weren't particularly IT savvy, but they did trust us to get on with it.

Posted

Post-incident reports go to the appropriate headteacher (and/or CEO, depending on impact). Typically a "this happened, this is why it happened, this is the current mitigation in place to prevent, this is (if applicable) additional mitigation being put in place".

 

Major changes (ISP swap, big migration etc) are communicated to headteachers ahead of time and they get a post-change report (typically any variance from planned changes).

Posted
Perhaps I was lucky then. I would put forward a budget plan at the beginning of the year and spend the money as I saw fit. Never needed to argue over infrastructure as it was (rightly) seen as a vital component of teaching and learning. A thing like a single server disk drive would never have been an issue to fund. Our SLT weren't particularly IT savvy, but they did trust us to get on with it.

 

Yes you were very very lucky. That is not how it is in most schools still.

 

I also have been very very lucky (for the most part).

 

But it can pivot overnight, with people saying "what do you even do" "psydii doesn't need any more toys" (when I was advocating the replacement of the aging server infrastructure) "we don't need the rolls-royce solution" (it was just a mid-tier flat panel for the classrooms) "stop being so inflexible" (the deployment plan had very specific critical path, and trying to implement D before B was a disaster).

 

Also to look to your link to the Google SRE handbook. Translate and scale that back to a single site school.... it is clear (to me) that some of the equivalent reporting would indeed have pro-active interventions and incidents that were serious but customer impact was avoided, written up and shared with the relevant teams. In schools its just that some of the relevant teams are not IT but instead SLT / Governors.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...