SchoolsBroadband Posted May 21, 2020 Posted May 21, 2020 Afternoon all, i'm on holiday today but from a couple of whatsapp messages i've had from one of our chief engineers it looks as though a software bug in the FortiOS caused HA to fail on a specific cluster pair of Fortigates. We run them in Active / Passive, Passive / Active mode (a hybrid of Active / Active) and those that were connected to one of them were not automatically failed over to the other firewall. We are investigating with Fortinet TAC but this looks like a software bug. We performed an emergency reboot of the device which forced HA to kick in and move affected clients onto the other firewall. Once the device rebooted it rejoined the HA cluster and took back half of the traffic of that pair of Fortigates. We are rebooting the other Fortinet device in that pair as a precaution this evening. The Fortigate that the issue occurred on had 630 days of uptime since it was last rebooted :-/ We'll wait to see what Fortinet TAC says and then issue the RFO via our normal channels. The Fortinets are normally pretty solid so this is surprising. The customers on the other side of the pair continued to function as normal. Our other cluster pairs of Fortinets were also not affected. In the background prior to this, we've already engaged with Fortinet professional services to upgrade all of our core Fortinet estate over the summer to a version of V6 from v5 so they'll be a host of features enabled as well as permanently fixing some bugs in the current FortiOS software revision. Thanks Dave Morning all, the other Fortinet which hadn't had an issue in that pair was rebooted in the early hours of the morning and has successfully re-joined the HA cluster. We will continue to monitor the service and will update with an RFO to those affected shortly. Thanks Dave
BrotherSidious Posted May 21, 2020 Posted May 21, 2020 SBB are not the only ones to quote metrics so this part is not aimed solely at them but uptimes, stability, failovers, degradation versus failure and whatever other metric is flavour of the month; in the words of Shania Twain "that don't impress me much". The only metric I care about is the simple one that says: when needed, did the broadband connection allow for the school administration to function as expected and did it allow for the delivery of the curriculum as expected. Now I fully understand that ISP's are not in a position to provide such metrics, however, they are the only metrics that staff and students care about. They have no time for fancy names, graphs, technical explanations or anything else. They have a very basic expectation, that it works consistently. When you have changed nothing internal to your network and suddenly your broadband connection stops working, questions are rightly asked. Unfortunately answers, such as; X has changed and is better than Y, Z may have had a problem but W has been rock solid or B failed but not for everyone and the other time it was C and affected different people means zero to staff and students. I appreciate these responses are technical in their nature and aimed at us as support staff but our non-technical response to staff and students when they ask whose fault it is will be "our ISP". Now an ISP may pass the buck on to its partners or other services that they work with but I'm afraid the ISP chooses those partners and services so from our point of view, we bought our broadband/ filtering package from you so the buck stops with you. If the fault is with partners/services then pick better partner/services. It is also somewhat concerning that on discovering an issue, the stock answer tends to be that ‘we are not seeing any issues our end.’ Why does my free monitoring software pick up issues with an internet connection within seconds of it going down yet the ISP cannot ‘see anything?’ It goes back to the point of expectations, I am not concerned about whether X or Y is working but simply does my internet connection and the level of filtering that as a school we are obliged to provide, work. There comes a point when it is out of the technicians hands, as I feel it should be, and our customers (staff and students) get to vocalize their concern. They ask for change because regardless of the provider's metrics they are not happy, their lessons are being disrupted, admin functions cannot be completed, paid for online resources cannot be accessed etc. Over the last 12 months, the feedback from staff and students is that our broadband connection has become less reliable. Rightly so, we have to answer to them. We are, after all, service providers ourselves. As such, I would imagine that at the end of our current contract we will not be renewing with SBB as the only real metric we care about, that it is working as expected, is not being met consistently enough. The perception is that it has been getting progressively worse despite statements that it is better and more stable. Yesterday we experienced another outage which only serves to reinforce that perception. Now, 60 minutes may be insignificant to an ISP's metric but that has a major impact on schools, especially at the moment with all of the home working and the remote accessing of resources going on by staff and students. It is already a stressful time for staff and students without us adding to it. Whatever these enhancements and improvements are, whatever these updates and patches are and whatever new systems and hardware is being bought online is of zero interest to staff and students and they get the deciding vote. I may well get to say where we go next but they definitely get to say enough is enough. Regardless of this, to any and all reading this post I hope that you stay well, be safe and take care 3
Steve21 Posted May 21, 2020 Posted May 21, 2020 It is also somewhat concerning that on discovering an issue, the stock answer tends to be that ‘we are not seeing any issues our end.’ Why does my free monitoring software pick up issues with an internet connection within seconds of it going down yet the ISP cannot ‘see anything?’ It goes back to the point of expectations, I am not concerned about whether X or Y is working but simply does my internet connection and the level of filtering that as a school we are obliged to provide, work. Yep, this 100x over, and may I say very well said! Schools want a service that will work, can cope, and when "the odd" downtime occurs a quick and decent response It really does seem like they can't actually tell who's being affected on these setups which is the most worrying part... and then there's the constant disregard to any of these questions from their representatives Don't really think we have much option but to look at all of our schools that joined SBB previously and move them off when we can Steve
caffrey Posted July 2, 2020 Posted July 2, 2020 Anyone getting a 404 on the blockpage ? Or is it just me ? Server Error 404 Web server error '404' encountered when trying to process request to 'http://denied.schoolsbroadband.net/webadmin/deny/defaultblockpage.php? That URL doesn't work either
RobMason Posted July 2, 2020 Posted July 2, 2020 Anyone getting a 404 on the blockpage ? Or is it just me ? Server Error 404 Web server error '404' encountered when trying to process request to 'http://denied.schoolsbroadband.net/webadmin/deny/defaultblockpage.php? That URL doesn't work either Good Morning Caffrey Is this happening consistently? Our monitoring isn't flagging anything and I haven't heard of any other reports. FYI, you can't browse to the URL directly, that is by design. Thanks,
caffrey Posted July 2, 2020 Posted July 2, 2020 Good Morning Caffrey Is this happening consistently? Our monitoring isn't flagging anything and I haven't heard of any other reports. FYI, you can't browse to the URL directly, that is by design. Thanks, Just the report up to now, I can't load that page either (not sure if I'm supposed to be able to) I get a "your internet access is blocked" . My PC is set to bypass netsweeper entirely through a different gateway though.
sippo Posted August 18, 2020 Posted August 18, 2020 Not sure how true but: Apparently they've had a bad power cut at their data centre in London which powers most of the UK. They're due an update in 20 minutes, but apparently the fire brigade is involved so they don't have an eta on that.
SchoolsBroadband Posted August 18, 2020 Posted August 18, 2020 (edited) Not sure how true but: Morning all, Equinix LD8 has a major power issue that is affecting numerous ISP's including our normal competitors. Please see a copy of the comms we were sent this morning from Equinix. Equinix is one of the UK's major Internet hubs. IBX(s): LD8 IBX Address: 6/7/8/9 Harbour Exchange Square London E14 9GE United Kingdom Ticket#: 5-200094218580 Date and Time of Occurrence: 18-AUG-2020 04:40 Site Local Time Date and Time Update Reported: 18-AUG-2020 09:01 Site Local Time INCIDENT SUMMARY: Fire Alarm - IBX was Evacuated UPDATE: Equinix IBX Site Staff reports that IBX Engineers are currently migrating customers' supplies into new infrastructure. IBX Engineers will start to investigate the cause of the failure on UPS. The next update will be provided in approximately 30 mins. IBX(s): LD8 IBX Address: 6/7/8/9 Harbour Exchange Square London E14 9GE United Kingdom Ticket#: 5-200094218580 Date and Time of Occurrence: 18-AUG-2020 04:40 Site Local Time Date and Time Update Reported: 18-AUG-2020 08:34 Site Local Time INCIDENT SUMMARY: Fire Alarm - IBX was Evacuated UPDATE: Equinix IBX Site Staff reports that IBX Engineers and the specialist vendor have begun restoring services to customers by migrating to newly installed and commissioned infrastructure. IBX Engineers continue to work towards restoring services to all customers and further updates will follow when more information is available. The next update will be provided in approximately 30 mins. IBX(s): LD8 IBX Address: 6/7/8/9 Harbour Exchange Square London E14 9GE United Kingdom Ticket#: 5-200094218580 Date and Time of Occurrence: 18-AUG-2020 04:40 Site Local Time Date and Time Update Reported: 18-AUG-2020 06:44 Site Local Time INCIDENT SUMMARY: Fire Alarm - IBX was Evacuated UPDATE: Equinix IBX Site Staff reports that fire alarm was triggered by the failure of output static switch from Galaxy UPS system supporting levels 1, 2, 3, 4 in building 8/9 at LD8. This has resulted in a loss of power for multiple customers and IBX Engineers are working to resolve the issue. The next update will be provided in approximately 30 mins. IBX(s): LD8 IBX Address: 6/7/8/9 Harbour Exchange Square London E14 9GE United Kingdom Ticket#: 5-200094218580 Date and Time of Occurrence: 18-AUG-2020 04:40 Site Local Time INCIDENT SUMMARY: Fire Alarm - IBX was Evacuated INCIDENT DESCRIPTION: Equinix IBX Site Staff reports that a fire alarm has been triggered and the IBX was evacuated. IBX Engineers are currently investigating the issue. The next update will be provided in approximately 60 mins or earlier if available. Latest update is they have a new UPS on site and are installing that to bring customers back online. As to why this has happened we will find out in due course. Needless to say we take resilient power feeds to all of our cabinets so as you can imagine we're not best impressed. We've put out comms as usual and have been calling as many customers as possible. We don't have a resolution time as yet but we believe some services are starting to be restored within the data centre. Thanks Dave Edited August 18, 2020 by SchoolsBroadband 1
Opendium_Steve Posted August 18, 2020 Posted August 18, 2020 Latest update is they have a new UPS on site and are installing that to bring customers back online. As to why this has happened we will find out in due course. Needless to say we take resilient power feeds to all of our cabinets so as you can imagine we're not best impressed. I've seen a few complaints from people who have lost all of the diverse power to their racks, and the updates from Equinix seem few and far between as well. Think there will be some hard questions at the end of this. 1
RobFuller Posted August 18, 2020 Posted August 18, 2020 I've seen a few complaints from people who have lost all of the diverse power to their racks, and the updates from Equinix seem few and far between as well. Think there will be some hard questions at the end of this. This isn't good timing with results downloads tomorrow, good job for our backup line (different provider)! 1
DalekSec Posted August 18, 2020 Posted August 18, 2020 (edited) Is it ever coming back? I'm bored now I'm doing my least favourite thing while i wait.... tidying! Edited August 18, 2020 by DalekSec
caffrey Posted August 18, 2020 Posted August 18, 2020 I'm doing my least favourite thing while i wait.... tidying! I've just caught up with my shredding...
andy_b Posted August 18, 2020 Posted August 18, 2020 It does make me question the point of my redundant line if the outages are always further up the chain.
Opendium_Steve Posted August 18, 2020 Posted August 18, 2020 It does make me question the point of my redundant line if the outages are always further up the chain. Redundancy is always a balancing act between reliability, cost and complexity (not forgetting that complexity can harm reliability). You can get your backup lines from completely independent suppliers, but even then making sure their routes are as diverse as possible is hard - when a big datacentre goes down, multiple otherwise independent suppliers are affected. In fairness to the service providers who are affected, what I'm hearing is that many of them were paying for diverse power supplies and have lost power on both supplies. So there are going to be some pointed questions directed at Equinix as to whether they were actually providing what they were paid for, and what could have failed so catastrophically to take out supplies that were supposed to be independent. 1
sippo Posted August 18, 2020 Posted August 18, 2020 I really hope it is backup by the end of the day!
nathan Posted August 18, 2020 Posted August 18, 2020 You have a redundant line with the same ISP? It's not a single ISP that is affected though.
SchoolsBroadband Posted August 18, 2020 Posted August 18, 2020 Hi all, it looks like all of our services are coming back up and running. We need to check all kit has come back safely but we can now see connections back online. Apparently 2 of the 3 floors (floors 1 and 2) which had no power are now back online too. Hopefully none of our carriers are on floor 3...... We'll update everyone through the normal channels.... Dave
SchoolsBroadband Posted August 18, 2020 Posted August 18, 2020 Not here, oh well hometime anyway we are seeing some carriers as down still (Virgin in particular) so it's likely their equipment is not yet back on. All of our kit is up and now powered but if the carrier we interconnect to doesn't have power then obviously their circuit won't be working. THis is happening for other ISP's too. I've seen comms about everything should be fixed from a power point of view by 9 o'clock this evening from the data centre. Regards David
RobFuller Posted August 18, 2020 Posted August 18, 2020 It's not a single ISP that is affected though. VMB for us was unaffected, any idea who else was as could be useful for future procurement decisions.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now