dantheitman Posted January 18, 2024 Posted January 18, 2024 I'm working through the DfE standards for servers and storage, and one standard says: "Your senior leadership team will need to decide the maximum downtime it is willing to accept. This should be based on your specific needs, including the sensitivity of the system and data, and the importance of the system and data to running your school or college". What do you do in your schools regarding allowed periods of network downtime? Do you 'budget' for a certain amount of downtime for unplanned/planned? 1
PotNoodleTech Posted January 18, 2024 Posted January 18, 2024 I guess what it's angling at is, do you have plans and the equipment in place to minimize downtime because of hardware failure? if you have one server and you skimped on it so it only has 1 power supply and it goes down your whole network may be down for a day or two whilst a power supply is sent to you - is that acceptible? What about if it was the motherboard and the downtime now edges up to 4-5 days? What about if the server was unrepairable and you did not have it in warranty and had to order another server then transfer everything - what if the downtime edged up to 5-10 days? Would the school be able to be open with no access to IT or MIS or possible heating/building/gates/door access etc etc? If SLT pick a number of down time it is wiling to accept - say 2 days, then that means you need to get more than one server operating maybe in a cluster, and keep power supplies, hard disk and fans in stock, possible have a same/next day warranty, etc etc. So the number of days SLT pick - entirely dictates the kind of hardware and software setup and the cost you need to absorb to get all this reliability? 2
dmj Posted January 18, 2024 Posted January 18, 2024 I think it's referring to Service Level Objectives (SLO), ie you set a realistic SLO and measure uptime. There's a good write up about this here Essentially, you pick a number that everyone is happy with; say 99.0% uptime during working hours over a month period. This means in a one month period of 160 working hours you have to be up for 158.4 hours. The other 1.5 hours can be used for maintenance or emergencies (hardware failure etc). If you don't think you can guarantee that then it can be adjusted. If management want higher uptimes they have to pay with hardware and out of hours support. It also means that if there is an emergency that eats up an hour or so, you can't plan working hours maintenance as you have already used the error budget. 4
Davit2005 Posted January 18, 2024 Posted January 18, 2024 Never heard of SLO, lol. "SLOs: What's the Difference? SLAs are used externally to define an agreement between a company's service and its paid users. SLOs are objectives that are measured internally to determine whether the SLA is being met. If an SLO's terms are violated, teams must respond and react quickly to prevent from breaking the SLA." 3
dmj Posted January 18, 2024 Posted January 18, 2024 I should also mention there's some fairly intricate calculations involved for some services: for example, a network share can't have an SLO higher than it's constituent components - the network, the storage system and the service that serves the files. If the Network has 99.9% the storage server 99.0%, the storage system (SAN/NAS etc) 99.0% and the service itself has 99.0%, the total SLO for the network share would be (99.9% * 99% * 99% * 99% = 96.9% I think that's why the DfE have simplified it with 'servers' and 'storage' which sounds good on paper, but arguing that the servers met their SLO whilst the networks was down probably won't wash. 3
DrBeaker Posted January 18, 2024 Posted January 18, 2024 All of these DFe standards link into the thread on devaluing of jobs here i.e. More and more and MORE is being asked of school IT Technicians - demands, responsibility, SLA's etc. Now the DFE is saying about 'acceptable levels of downtime'. Yet most IT staff get to overtime, no TOIL, no flexible working etc compared to teachers or estates teams to be able to sort out 'downtime'. Yet in most cases the pay level and compensation packages of school IT Technicians aren't matching suit. No wonder why there's a recruitment crisis in the public sector. 3
GrumbleDook Posted January 19, 2024 Posted January 19, 2024 The standards are not linked to the devaluing of IT roles. In many cases the standards are reworking areas from the old FITS 1.5 and v2 guidance and levels. Yes, some of it is dumbed down in language (as already pointed out in this thread), but it all comes back to establishing good working practices… something that has been worked on since 2001. Salary and terms of employment are always going to be an issue, as they are for many other roles outside of teaching, but the standards are simply part of establishing good practices.
enjay Posted January 19, 2024 Posted January 19, 2024 Now the DFE is saying about 'acceptable levels of downtime'. Yet most IT staff get to overtime, no TOIL, no flexible working etc compared to teachers or estates teams to be able to sort out 'downtime'. But we do work year-round, so there's a maintenance window every 6 weeks. Staff here know we often work on things over the holidays, so (are meant to) contact me and Site Manager if they plan on coming in and needing particular rooms/resources/systems so we can try to accommodate them. Anything urgent during term-time we give as much notice as we can, then do it at the end of the day with agreed flexible working for my team, or I connect in from home that evening. We have some redundancy on systems but don't have an agreed availability percentage. As has been said, there's an analysis of cost, benefit and impact to be done on that. Of course, elsewhere in the updated Standards is an SLT technology lead, so none of this will be our jobs to worry about ;-) 1
Popular Post pete Posted January 19, 2024 Popular Post Posted January 19, 2024 (edited) I'd ignore "downtime" and concentrate on service availability. If authentication keeps working but you had a staggered reboot of all 3 DCs in the last hour, no-one cares. Make your services sufficiently redundant and perform (scheduled, vs cowboy ad-hoc) maintenance during the working day. Edited January 19, 2024 by pete 5
dmj Posted January 19, 2024 Posted January 19, 2024 I'd ignore "downtime" and concentrate on service availability. If authentication keeps working but you had a staggered reboot of all 3 DCs in the last hour, no-one cares. Make your services sufficiently redundant and perform (scheduled, vs cowboy ad-hoc) maintenance during the working day. ^ definitely this. We go further and have server maintenance automated, at any one point in the day I couldn't tell you how many servers are up or down. Load balancers and cluster management handle it all. Service uptime is all anyone cares about.
enjay Posted January 19, 2024 Posted January 19, 2024 While we're talking downtime, have you seen "Servers containing critical data should have ... backup servers (onsite or in the cloud) and network cards that switch over seamlessly when one fails"? Hands up who has that setup in their school.
TechMonkey Posted January 19, 2024 Posted January 19, 2024 While we're talking downtime, have you seen "Servers containing critical data should have ... backup servers (onsite or in the cloud) and network cards that switch over seamlessly when one fails"? Hands up who has that setup in their school. Sounds like a cluster. Had it at last 2 places for everything not just critical and will get there here. N+1 infrastructure, nothing fancy. Having it transition semi-seamlessly to the Cloud would be nice for a temporary fix.
Davit2005 Posted January 19, 2024 Posted January 19, 2024 All of these DFe standards link into the thread on devaluing of jobs here i.e. More and more and MORE is being asked of school IT Technicians - demands, responsibility, SLA's etc. Now the DFE is saying about 'acceptable levels of downtime'. Yet most IT staff get to overtime, no TOIL, no flexible working etc compared to teachers or estates teams to be able to sort out 'downtime'. Yet in most cases the pay level and compensation packages of school IT Technicians aren't matching suit. No wonder why there's a recruitment crisis in the public sector. Even if you get TOIL it is hard to get it off as if you are short on staff anyway. I'd rather get the money, I know as a manager you might be expected from some to do long hours and not get any overtime pay but it can take the mick.
Davit2005 Posted January 19, 2024 Posted January 19, 2024 (edited) While we're talking downtime, have you seen "Servers containing critical data should have ... backup servers (onsite or in the cloud) and network cards that switch over seamlessly when one fails"? Hands up who has that setup in their school. Most things can be achieved by investment but that is sometimes the issue. The org needs to weigh up acceptable downtime against investment I feel then they are answerable. Edited January 19, 2024 by Davit2005
PotNoodleTech Posted January 19, 2024 Posted January 19, 2024 I guess it depends if you are able to get the funding to double/triple up on servers/switches/storage etc. Don't forget a cluster that uses just one storage tray or just one fibre switch or just one UPS etc etc is not fault tolerant. 2
TechMonkey Posted January 19, 2024 Posted January 19, 2024 I guess it depends if you are able to get the funding to double/triple up on servers/switches/storage etc. Don't forget a cluster that uses just one storage tray or just one fibre switch or just one UPS etc etc is not fault tolerant. It's about assessment, judgement and planning. If you are being fed by one substation you aren't fault tolerant, so do you get the power company to give you two separate power feeds? A storage tray that has duplicate controllers, power supplies and redundant drives is pretty bullet proof, so do you need to invest in a duplicate? If your supplier/support can get you replacement within 8 hours is that enough? The way to look at these things is that it gives you ammunition to go to SLT and the governors to say, this is what the Government say we should have and be doing, when inspection comes round what are you going to say to them? It helps us to get that funding as we have the case, rather than just the IT guy says. 1
Davit2005 Posted January 19, 2024 Posted January 19, 2024 (edited) I guess you should also have ample monitoring to tell you when you have lost fault tolerance rather be in the scenarion where you don't know one switch, a drive, a controller, etc. has failed. Edited January 19, 2024 by Davit2005
msi_school Posted January 19, 2024 Posted January 19, 2024 Most things can be achieved by investment but that is sometimes the issue. The org needs to weigh up acceptable downtime against investment I feel then they are answerable. I always love these conversations, once when I am in the wrong for wanting redundancy because it "never goes wrong and we don't need it" and then once when I am in the wrong for being cavalier because "it always goes wrong and its essential" and both times it's too expensive. 2
dmj Posted January 19, 2024 Posted January 19, 2024 I guess it depends if you are able to get the funding to double/triple up on servers/switches/storage etc. Don't forget a cluster that uses just one storage tray or just one fibre switch or just one UPS etc etc is not fault tolerant. It's not typically double or triple for redundancy unless you are running really outdated applications, or only have a couple of servers. With containerised apps you typically run them in a cluster of servers with enough redundancy to cover one or two servers going down and enough to take peak loads (in cloud the servers will just provision themselves and scale to near zero with no traffic). It's not super expensive, the larger the org the cheaper it is. With only one or two servers on site it's double but if you already have 10 then you probably already have capacity to provide redundancy if the applications can transition their workloads correctly. Same with storage if you're clustering it across servers.
TechMonkey Posted January 19, 2024 Posted January 19, 2024 I always love these conversations, once when I am in the wrong for wanting redundancy because it "never goes wrong and we don't need it" and then once when I am in the wrong for being cavalier because "it always goes wrong and its essential" and both times it's too expensive. If you have requested redundancy and you are overruled then you have CYA. If you lay out the options and state why you have chosen your preferred option, then you have done your best and been diligent. If management decide they do not feel the risk is big enough and go the cheaper option then it is down to them and you can point to your recommendations. It is why, even if I know the answer, I always provide the safe option as at least I have given the option. Also comes under the rule of "don't ask, don't get". If they use the "it never goes wrong and we don't need it" excuse, ask them if we have insurance for anything? Same principal, we pay the money hoping not to use it. And we thank our lucky stars when we do have to.
PotNoodleTech Posted January 19, 2024 Posted January 19, 2024 It's not typically double or triple for redundancy unless you are running really outdated applications, or only have a couple of servers. With containerised apps you typically run them in a cluster of servers with enough redundancy to cover one or two servers going down and enough to take peak loads (in cloud the servers will just provision themselves and scale to near zero with no traffic). It's not super expensive, the larger the org the cheaper it is. With only one or two servers on site it's double but if you already have 10 then you probably already have capacity to provide redundancy if the applications can transition their workloads correctly. Same with storage if you're clustering it across servers. Not typically double? Well - it sure isn't single! I've never seen a fault tolerant single device. Again all comes down to budget. Some schools have a single server with 5 VMs on. Some school have 3 VM hosts with 50VMs on 2 SANS and 2 Fibre switches 2 xUPS diesel backup generator etc etc 1
dmj Posted January 19, 2024 Posted January 19, 2024 Not typically double? Well - it sure isn't single! I've never seen a fault tolerant single device. Again all comes down to budget. Some schools have a single server with 5 VMs on. Some school have 3 VM hosts with 50VMs on 2 SANS and 2 Fibre switches 2 xUPS diesel backup generator etc etc Of course if you have a single server any redundancy will double the cost!
enjay Posted January 19, 2024 Posted January 19, 2024 Some school have 3 VM hosts with 50VMs on 2 SANS and 2 Fibre switches 2 xUPS diesel backup generator etc etc You forgot to mention two fibre runs into each building by totally different routes, so when a digger/tree/squirrel goes through one, they can still connect (and even then, the internal cabling is still a single point of vulnerability). Oh, and two server rooms of course. 1
psydii Posted January 19, 2024 Posted January 19, 2024 (edited) You forgot to mention two fibre runs into each building by totally different routes, so when a digger/tree/squirrel goes through one, they can still connect (and even then, the internal cabling is still a single point of vulnerability). Oh, and two server rooms of course. I feel seen. (though the second server room doesn't actually have any servers or a rack, but we could stand it up in the length of time it would take to order them) But seriously - a single server can have redundancy- dual power supplies a RAID and ECC Memory, with two UPS's and have it be a VM Host with your critical infra-servers all being VMs on the box. You get almost all the uptime improvements of a 60K multi-host/san based solution at a fraction of the cost. I've run a couple of schools on such a set up. In schools, generally the services we run on our servers are typically pretty simple/basic affairs so uptime and availability of each vm does actually matter (1 vm per service, so VM uptime is a reasonable proxy for service uptime, and service uptime is a critical component in service availability). The effort to run clustering or re-architect the services so they run across multiple hosts with load balancers where the relevance of the individual vm is abstracted away is really not worth the effort in a in a single site school or even a small-medium MAT. We (school users) can accept *way* more downtime (after school-core hours, overnight, school holidays) than many businesses. As long as you have *some* resiliency/redundancy at the lowest levels (power, storage, connectivity) it is relatively trivial to meet availability needs (for example the Head is not going to sanction staff being on the MIS at 2am - so that's a safe time for updates to roll out and servers reboot automatically). Of course, running equivalent services at a global scale is a completely different proposition, and we can benefit from that but we need the whole stack to the data centers to be reliable. Google, Amazon and Microsoft can handle the service infrastructure complexity from there. However things start to get sticky when you land in the middle - Classcharts, CPOMS, Applicaa and Bromcom I think being pertinent examples here. They seem barely more reliable than running equivalents in house. I've not read the DfE guidance, but I wonder how these guys stack up against the service levels we (the school/MAT teams) are expected to deliver. Edited January 19, 2024 by psydii
dmj Posted January 19, 2024 Posted January 19, 2024 In schools, generally the services we run on our servers are typically pretty simple/basic affairs so uptime and availability of each vm does actually matter (1 vm per service, so VM uptime is a reasonable proxy for service uptime, and service uptime is a critical component in service availability). 1 VM per service was fairly sensible 10yrs ago but nowdays is just a massive waste of resources, especially if the VM is over provisioned in the first place to account for periods of high usage. I guess you can do some vertical scaling here, but it's probably still out of reach for single server schools. The effort to run clustering or re-architect the services so they run across multiple hosts with load balancers where the relevance of the individual vm is abstracted away is really not worth the effort in a in a single site school or even a small-medium MAT. For sure, but some of the larger MAT's approach the size of mid-level universities, who absolutely do re-architect services and attempt to run them in efficient ways. I guess it's just one of those things that the more services you run the cheaper each service is.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now