Ex-MGSTech Posted September 28, 2023 Posted September 28, 2023 I'm often told we don't shout about our achievements enough. For example we had a disk failure in the SAN at the weekend, all covered by resilient systems and resolved by Tuesday BUT our SLT are none the wiser Do any of you complete an IT Incident Report (form?) that goes to maybe your SLT that effectively demonstrates or records all the incidents we deal with of say at least a severe level (e.g. likely to impact T&L) That the SLT would otherwise be completely unaware of! I guess I'm either looking at an MS form or a word template? Regards Steve We have a ticketing system, this is more management reporting similar to a cyber incident action report etc
TechMonkey Posted September 28, 2023 Posted September 28, 2023 (edited) I don't do a form but I do send after action emails if we had an outage or incident. Normally only if it has affected operations, but something like weekend work would definitely get one to highlight it was done. Nothing grandiose, just the facts with timings and work carried out. EDIT: If your help desk is ITIL based it should have a incident section that can log this. Edited September 28, 2023 by TechMonkey
Ex-MGSTech Posted September 28, 2023 Author Posted September 28, 2023 EDIT: If your help desk is ITIL based it should have a incident section that can log this. No we kept the ticketing system really simple and integrated into Office 365 https://www.ivero.net/solutions/HelpDeskPlus/index.html
Oaktech Posted September 28, 2023 Posted September 28, 2023 I send always send a post incident review to management, but it's just an email.
dmj Posted September 28, 2023 Posted September 28, 2023 we do a postmortem of every incident, the process is described here: https://sre.google/sre-book/postmortem-culture/ I wouldn't class that as an incident though - A disk failure is expected behaviour and you already mitigated against it.
psydii Posted September 29, 2023 Posted September 29, 2023 we do a postmortem of every incident, the process is described here: https://sre.google/sre-book/postmortem-culture/ I wouldn't class that as an incident though - A disk failure is expected behavior and you already mitigated against it. While you are right... the context is a school. SLT need to be reminded that the money spent by IT *actually* prevented an outage today. Otherwise they forget that resilient and reliable IT requires ongoing investment and next year slash your budget and hire another teacher. 3
dmj Posted October 1, 2023 Posted October 1, 2023 While you are right... the context is a school. SLT need to be reminded that the money spent by IT *actually* prevented an outage today. Otherwise they forget that resilient and reliable IT requires ongoing investment and next year slash your budget and hire another teacher. Where would you draw the line though? Systems are designed to be resilient and failover. You could end up just reporting everything as an incident and not focussing on where money needs to be spent.
jmak Posted October 1, 2023 Posted October 1, 2023 Where would you draw the line though? Systems are designed to be resilient and failover. You could end up just reporting everything as an incident and not focussing on where money needs to be spent.It really depends on the attitude and level of technical knowledge of the management. In a school, I had to fight to replace a failed disk in a raid array in because the server was due to be replaced in four months and in another larger organisation where the server infrastructure had become horrendously outdated, the management intention was to get rid of the people who'd fixed it as soon as possible. They virtualized everything and centralised hardware so that it was less visible, so management assumed there was almost nothing left to do.
hardtailstar Posted October 1, 2023 Posted October 1, 2023 At the company I work for, every incident is a ticket then an email is sent to those in the need to know and involved. The email basically says 'who, what, where and resolution' with a link to the ticket for a more technical view. Also any questions/suggestions are then put in that ticket rather than the email. Also there is a culture here to promote what you have done that you feel should be shouted about. I recently tidied up the 12 meetings rooms we had so every room had Google Meet, Apple tv and a tv instead of a projector. I also made sure that there was tips on how to use it or simple trouble shooting. I didn't say anything about it, just got on with it as a job and I got 'told off' for not telling the company that they now are all identical etc. 1
dmj Posted October 2, 2023 Posted October 2, 2023 It really depends on the attitude and level of technical knowledge of the management. In a school, I had to fight to replace a failed disk in a raid array in because the server was due to be replaced in four months and in another larger organisation where the server infrastructure had become horrendously outdated, the management intention was to get rid of the people who'd fixed it as soon as possible. They virtualized everything and centralised hardware so that it was less visible, so management assumed there was almost nothing left to do. Perhaps I was lucky then. I would put forward a budget plan at the beginning of the year and spend the money as I saw fit. Never needed to argue over infrastructure as it was (rightly) seen as a vital component of teaching and learning. A thing like a single server disk drive would never have been an issue to fund. Our SLT weren't particularly IT savvy, but they did trust us to get on with it.
pete Posted October 2, 2023 Posted October 2, 2023 Post-incident reports go to the appropriate headteacher (and/or CEO, depending on impact). Typically a "this happened, this is why it happened, this is the current mitigation in place to prevent, this is (if applicable) additional mitigation being put in place". Major changes (ISP swap, big migration etc) are communicated to headteachers ahead of time and they get a post-change report (typically any variance from planned changes).
psydii Posted October 2, 2023 Posted October 2, 2023 Perhaps I was lucky then. I would put forward a budget plan at the beginning of the year and spend the money as I saw fit. Never needed to argue over infrastructure as it was (rightly) seen as a vital component of teaching and learning. A thing like a single server disk drive would never have been an issue to fund. Our SLT weren't particularly IT savvy, but they did trust us to get on with it. Yes you were very very lucky. That is not how it is in most schools still. I also have been very very lucky (for the most part). But it can pivot overnight, with people saying "what do you even do" "psydii doesn't need any more toys" (when I was advocating the replacement of the aging server infrastructure) "we don't need the rolls-royce solution" (it was just a mid-tier flat panel for the classrooms) "stop being so inflexible" (the deployment plan had very specific critical path, and trying to implement D before B was a disaster). Also to look to your link to the Google SRE handbook. Translate and scale that back to a single site school.... it is clear (to me) that some of the equivalent reporting would indeed have pro-active interventions and incidents that were serious but customer impact was avoided, written up and shared with the relevant teams. In schools its just that some of the relevant teams are not IT but instead SLT / Governors.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now