Jump to content

Recommended Posts

Posted
Can’t see how an RFO affecting 25 schools or so can take days to write. I’ve seen some companies publish RFO’s within hours. It must be fairly comprehensive perhaps.
  • Thanks 1
Posted (edited)
Can’t see how an RFO affecting 25 schools or so can take days to write. I’ve seen some companies publish RFO’s within hours. It must be fairly comprehensive perhaps.

 

While I do jest, partially, I can pretty much sum it up!

 

Customer - I have issues, please fix

SBB - No issues reported, will look

Customer - Still broke, please fix!

*Insert long wait*

SBB - It's fixed

Customer - So what was wrong?

SBB - RFO to follow

Customer - Waiting... waiting...

SBB - Intermittent (insert company) issue, not us, will tell them off in future

 

Repeat every few (insert days/week)

 

Still feel that despite SBB's opinion that it should only be sent to "affected schools" it's very much time they start to publish issues and resolutions publicly. The fact that schools might be getting the same issues at different times and won't know about it as they weren't in the first bunch of affected schools is still rather poor attitude imho.

 

Whether it's 25 or 25,000 school an issue affects, it's still an issue and one that schools should be able to check, both in terms of QoS and that if they had issues they know it's resolved fully and not going to happen more (or replicated around each box/filter etc)

 

Steve

Edited by Steve21
  • Thanks 1
Posted
Can’t see how an RFO affecting 25 schools or so can take days to write. I’ve seen some companies publish RFO’s within hours. It must be fairly comprehensive perhaps.

 

 

Evening all,

 

We always try to get an rfo out within a week of any issue happening if not then we will send a holding communication. A week from the point of issue I believe is tomorrow so we are within our normal time parameters which is defined in our service handbook.

 

A draft of the comms has already been sent from our noc team to our operations team at the end of last week. If we can give it to those affected then we will send just to them. If we can't or it's too difficult / takes too long to accurately identify those affected then we will send it to everyone or a larger sub set of customers.

 

The comms will be sent out shortly.

 

Thanks

 

Dave

Posted

Com was sent out today:

 

Netsweeper Connectivity Issue – Monday 11th May 2020 Reason for Outage Report

Fault Summary

Fault Detection Time:   08:58 11/05/2020

Fault Resolution Time: 17:00 11/05/2020

P0 Case Reference - 01144666

 

Impact Level

Some customers experienced degraded connectivity when trying to browse the internet.

Incident Timeline

08:58 – First customer case raised relating to no Internet

09:58 – Escalated to NOC

10:15 – Interim fix put in place

10:26 – Majority of customers are back online, some are still down due to configuration issues

11:54 – Further fix put in place to resolve configuration issues

13:34 – Operations asked to check small number of remaining customers that reported were offline are now back online

13:41 – Operations confirm they see customers as being online

15:01 – Information passed to NOC following feedback from a very small subsection of customers (sub 20) stating they are still down

15:07 – Operations escalate customer down cases to NOC

16:20 – NOC determine issue and root cause with remaining customers and start fixing reported customers downs one by one to finally resolve the issue

17:00 – Customers all report they are back online

Incident Summary

A misconfiguration within one of our automation scripts responsible for new customer deployment unfortunately caused a flush of our IP rules upon the inline platform. These IP rules are critical to how Netsweeper functions and as such it was necessary for engineers to manually restore the IP rules from a backup. Without these IP rules customer traffic will be blocked at the Netsweeper inline proxy server. Initial detection of the issue was difficult due to reduced traffic levels and customer feedback, due to COVID-19. The code in question was initially removed from the platform and has now been patched to ensure this cannot happen again.

 

Conclusion

Regretfully, this incident was self-inflicted. The issue was not down to a software bug or Netsweeper configuration but to human error. We understand that this should not be able to happen and as such we have reviewed our policies and procedures relating to code deployment as well as implementing safeguards to prevent this specific issue from happening again. We were however able to restore service quickly once the loss of IP rules was detected in most cases.

It is important to note that the Netsweeper platform is currently stable and has been for many months, this incident was in no way related to those at the end of last year and the start of this one. We are confident you will not see any repeat issues.

Once again, we sincerely apologise for the inconvenience suffered during these already trying times.

If you have any queries or concerns raised from reading this RFO please contact our support department in the usual way.

Kind regards

Network Operations Centre

Posted (edited)
Regretfully, this incident was self-inflicted. The issue was not down to a software bug or Netsweeper configuration but to human error.

Ah. I see.

 

Edit:

 

We were however able to restore service quickly once the loss of IP rules was detected in most cases.

I'd disagree a tad. 10.15am until nearly 5pm.

Edited by Edu-IT
Posted
10:26 – Majority of customers are back online, some are still down due to configuration issues

2:24 PM - we had an issue affecting about 25 customers. I don't have any more information than that at the moment but will pass on once I've an update from the NOC team.

15:01 – Information passed to NOC following feedback from a very small subsection of customers (sub 20) stating they are still down

 

Can you clarify then, your initial comment about 25 customers being affected:

 

a) Did the "Majority of customers" refer to the 10 out of 25 that were fixed? (25 less "sub 20"), which is an odd majority

b) Or there were more than 25 affected, and your post is incorrect?

 

Again, not trying to be a pain but these stories don't add up again, considering your post was at 2:24 PM?

 

Steve

  • Thanks 1
Posted

hi @Steve21,

 

to clarify when i sent the post I believe we had about 25 schools still affected. This was then whittled down over the afternoon until all cases were cleared and everyone was back online.

 

We are unsure on the exact number of customers that were initially impacted in the early morning as it was very difficult to ascertain and the number one priority was to get people back online when we knew how the issue occurred. It wasn't everyone (as per discussion on this thread some people saying they were fine) but it was a number of customers. The majority of the initial batch of customers that were offline were back online by 10:26.

 

if you need any more information or want a call to discuss please send me a PM.

 

Thanks

 

Dave

Posted
to clarify when i sent the post I believe we had about 25 schools still affected. This was then whittled down over the afternoon until all cases were cleared and everyone was back online.

 

We are unsure on the exact number of customers that were initially impacted in the early morning as it was very difficult to ascertain and the number one priority was to get people back online when we knew how the issue occurred. It wasn't everyone (as per discussion on this thread some people saying they were fine) but it was a number of customers. The majority of the initial batch of customers that were offline were back online by 10:26.

 

if you need any more information or want a call to discuss please send me a PM.

 

Thanks

 

Dave

 

To be frank I don't think taking it to PM is correct, and feel rather irritated that despite my previous comments you even PM me directly to discuss it off this topic, as this refers directly to my previous comment about things not being very public facing with SBB

 

15th May 2020, 07:51 PM - Steve21 - Still feel that despite SBB's opinion that it should only be sent to "affected schools" it's very much time they start to publish issues and resolutions publicly. The fact that schools might be getting the same issues at different times and won't know about it as they weren't in the first bunch of affected schools is still rather poor attitude imho.

 

^ironic hey?

 

When you're officially posting a response to an issue schools have, and tell them you had an issue that has "about 25 customers", that's not the same as "exact number of customers that were initially impacted in the early morning as it was very difficult to ascertain", followed by "25 schools still affected"

 

Else you are effectively saying that your RFO is inaccurate to say the least, you don't know how many were fixed in the morning, as the majority of an unknown figure were resolved? If there were 26 customers with an issue, then fixing 1 isn't a majority. If there were 25,000, I think the other 24,975 might want to know...

 

Even by stating that you must have a rough idea of how many paying customers this affected?

 

Steve

Posted (edited)

Just a couple of questions?

 

08:58 – First customer case raised relating to no Internet

09:58 – Escalated to NOC

What happened during this hour? Why was it not escalated to the NOC sooner or did it take an hour for somebody to realise it was a NOC issue?

 

16:20 – NOC determine issue and root cause with remaining customers and start fixing reported customers downs one by one to finally resolve the issue

Was this the same issue - flush of our IP rules upon the inline platform - or another issue? Reason I ask is if it's the same issue why was was it not resolved during the 10.15am interim fix or the 11.54am further fix? The RFO says 'and as such it was necessary for engineers to manually restore the IP rules from a backup' and therefore if it's the same issue I would have thought when the IP rules were fully restored from the backup it would have solved the issue?

 

I don't follow the timeline fully and would appreciate further guidance on these points.

 

Edit: As a follow up.

 

We are unsure on the exact number of customers that were initially impacted in the early morning as it was very difficult to ascertain and the number one priority was to get people back online when we knew how the issue occurred

From a technical point of view, if it was the IP rules had been deleted, could you not compare the IP rules from those in the backup against those on the platform to work out what was missing and narrow down the sites affected as I would assume those IP rules relate to specific sites?

 

 

Edited by Edu-IT
Posted

Evening Steve,

 

I'm sorry you feel that way. I got in touch and asked by PM so I could get a NOC engineer who has a lot more knowledge than I do (as they dealt with the issue technically and I didn't as i'm not as technical as they are) to discuss. If we go into the tiny intricacies as to what we do at every action point along the way of an RFO or any issue then you'd have a very very long piece of documentation to read some of which probably isn't for public consumption as it would go into specifics of our solution design which could compromise the security of our platform. The RFO is a high level reason for the issue.

 

At the moment we live in very odd times and network traffic utilisation isn't anywhere near what it normally is due to people not being at school at all so there connections aren't being used. We do not have easy to access statistics which shows how many schools are actually using our Internet connections (unless we go through each graph per connection individually or write some new code to produce the stats). if it was in normal time and we saw a drop in network traffic by 50% then it's likely 50% of our customers would be offline but at the moment that's not correct as many aren't in use.

 

We believe the majority of customers were fixed because A) the number of calls into our call centre we had went down significantly and overall traffic throughput increased to near normal levels which would indicate most customers were back online. Also we had cases open and closed each individual one down in the afternoon. To be any more technical than that I'd need a NOC guy / gal to answer.

 

I'm very sorry this affected you and others and that I can't be more concise as to your requests in a public forum. Again though, I'd be happy for one of the NOC engineers that personally dealt with the issue to give you a call to discuss as I would with any other customer also.

 

I think i'll leave it at that. Have a good evening.

 

Dave

Posted

There is one part i really take issue with.

 

"It is important to note that the Netsweeper platform is currently stable and has been for many months"

 

24th Feb to 11th May is not many months, its 2 months and about 2 weeks..

Posted
I'm sorry you feel that way. I got in touch and asked by PM so I could get a NOC engineer who has a lot more knowledge than I do (as they dealt with the issue technically and I didn't as i'm not as technical as they are) to discuss. If we go into the tiny intricacies as to what we do at every action point along the way of an RFO or any issue then you'd have a very very long piece of documentation to read some of which probably isn't for public consumption as it would go into specifics of our solution design which could compromise the security of our platform. The RFO is a high level reason for the issue.

 

With all due respect that's not a very good answer, nothing I asked about requires a high detailed answer or any private information, and using that as a reason not to answer seems again a rather poor response.

 

If you have an issue that affected, lets say just Filter8 and your engineer "broke" (for lack of a better phrase) the Rules that affect every customer on that filter8, all the customers on that filter8 are affected, therefore a restore of a filter8 backups would be that number (Lets say 50 for argument sake). Whether or not they had a user using the internet at said time doesn't affect what schools were affected by that change.

 

If it's a filter higher in the chain, and it affects multiple filters lets say 8-10, that's still a number you would be able to know (Now 3 filters at 50 schools)

 

Or if it affected all filters (lets say 20) that's 1000 etc.

 

Therefore if you restored filter8 backups at 10am, you know 50 are resolved out of 150. If you restored filter 8-10 backups at once, you would know all 150 were resolved.

 

If you're doing manual restores of backups one-by-one to phrase your RFO, surely again you know how many were restored by your technicians...

 

I'm not sure why it seems unreasonable to you to expect to hear about an issue that affected a school without having to dig through and pull out every bit of information to get a simple answer without having to check in with staff/students to ask if they happened to not have internet at a certain time...

 

Steve

Posted
24th Feb to 11th May is not many months, its 2 months and about 2 weeks..

 

Realistically, its a shade under a month. It hasn't been tested in any significant way since then.

Posted
Even by stating that you must have a rough idea of how many paying customers this affected?

People pay for this service? For an ISP that doesn't have any kind of proactive monitoring other than looking at a traffic graph?

 

um...

  • Thanks 2
Posted
People pay for this service? For an ISP that doesn't have any kind of proactive monitoring other than looking at a traffic graph?

 

um...

 

The connection didn’t drop, just web browsing was affected.

 

Not sure how proactive monitoring would be able to check that?

Posted

@DrCheese We've lots of metrics we look at other than traffic graphs. Libre is one of the tools we use to achieve this although we use other products as well.

@snagrat We also do some HTTP traffic request monitoring to each server too so we can see if HTTP and HTTPS traffic is being passed specifically for the reason you mention.

 

Again I'm not going to say exactly what we do and don't monitor on a public forum but it is extensive. You can see for yourself what Libre is capable of as a minimum.

 

Thanks

 

Dave

Posted
@DrCheese We've lots of metrics we look at other than traffic graphs. Libre is one of the tools we use to achieve this although we use other products as well.

@snagrat We also do some HTTP traffic request monitoring to each server too so we can see if HTTP and HTTPS traffic is being passed specifically for the reason you mention.

 

We do not have easy to access statistics which shows how many schools are actually using our Internet connections (unless we go through each graph per connection individually or write some new code to produce the stats)

 

And seeing you seem to have skipped my post:

 

That's two very different stories?

 

If you know and have access to more monitoring/filtering, you must know how many paying customers this downtime affected and whether they've actually been told of an issue...

 

So which is it, you do have knowledge of how many were affected and are refusing to say (Which makes it sound very much like you're trying to hide it from those who didn't log it directly with you), or you don't have knowledge? (and your comments above are again inaccurate, and leads to the question why don't you know/have access to this data?)

 

Steve

  • Thanks 1
Posted
We are back now but apparently SBB didn't see any issues at their end. Nice to hear we weren't the only ones that went down as you start to doubt yourself (not good that you were down too though). I guess you did nothing like us and it suddenly fixed itself. Not really sure which magic elves to thank for fixing it as it wasn't us and apparently not SBB either.
Posted
We are back now but apparently SBB didn't see any issues at their end. Nice to hear we weren't the only ones that went down as you start to doubt yourself (not good that you were down too though). I guess you did nothing like us and it suddenly fixed itself. Not really sure which magic elves to thank for fixing it as it wasn't us and apparently not SBB either.

 

We were told the memory on their firewall was clogged up. 😐 They couldn't access it when we called them initially.

 

When we tried calling back when the Internet came back on; we couldn't get through as they were experiencing a high level of calls. So I don't think it was just us.

Posted (edited)

Afternoon all,

 

i'm on holiday today but from a couple of whatsapp messages i've had from one of our chief engineers it looks as though a software bug in the FortiOS caused HA to fail on a specific cluster pair of Fortigates.

 

We run them in Active / Passive, Passive / Active mode (a hybrid of Active / Active) and those that were connected to one of them were not automatically failed over to the other firewall. We are investigating with Fortinet TAC but this looks like a software bug. We performed an emergency reboot of the device which forced HA to kick in and move affected clients onto the other firewall. Once the device rebooted it rejoined the HA cluster and took back half of the traffic of that pair of Fortigates. We are rebooting the other Fortinet device in that pair as a precaution this evening. The Fortigate that the issue occurred on had 630 days of uptime since it was last rebooted :-/

 

We'll wait to see what Fortinet TAC says and then issue the RFO via our normal channels. The Fortinets are normally pretty solid so this is surprising. The customers on the other side of the pair continued to function as normal. Our other cluster pairs of Fortinets were also not affected.

 

In the background prior to this, we've already engaged with Fortinet professional services to upgrade all of our core Fortinet estate over the summer to a version of V6 from v5 so they'll be a host of features enabled as well as permanently fixing some bugs in the current FortiOS software revision.

 

Thanks

 

Dave

Edited by SchoolsBroadband
  • Thanks 1

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...