lou-ashtonhurst
Members-
Posts
3 -
Joined
-
Last visited
Reputation
60 ExcellentAbout lou-ashtonhurst

Personal Information
-
Occupation
Senior Network Consultant
- X
Employer (optional)
-
Company Represented
FAELIX
-
Schools Broadband issues
lou-ashtonhurst replied to nicholab's topic in Internet Related/Filtering/Firewall
Hi folks I thought I'd just drop you all a note with a bit more of an explanation of today’s incident, the timeline of events so far and why we are in the position we are now in. I'll leave the commercial discussions to Dave - this post is purely about the tech side of things for those that are interested. The first incident occurred on 23/11 at 2pm, the second on the 24th at 1pm, the third at near enough midday on the 25th - a separation of approximately 23 hours each time. After the first incident, we reported it through to Fortinet and awaited their debug - but we weren't overly concerned. These things do happen and initially this looked like a random, one off occurrence - all our standard troubleshooting showed no problems and while we awaited Fortinet's diagnosis, we put this down to a random unit crash. As I mentioned in the earlier post, we were made aware of problems with the HA after the second occurrence which we resolved that night in an emergency maintenance slot. The next day we had the third occurrence - after which we replaced the hardware. Each of these steps was taken as soon as possible in an emergency maintenance window and on the basis of decisions of myself, our Technical Director, our senior engineers and Fortinet's senior engineers. We had been aware of a potential issue with a process called miglogd, which since the update to firmware 5.2.8 had been regularly crashing (but not with significant impact to the box). Over the last few weeks we have implemented a workaround on Fortinet's recommendation, knowing that a new firmware release was required to resolve it and was in the works. As with any vendor, Fortinet quite rightly run significant testing before any firmware upgrade is rolled out. The debug output from the first night led us to miglogd, the second night showed HA problems, the third we suspected a hardware fault on the ASIC as traffic seemingly wasn't being forwarded from the WAN port. These were all approx 23 hours apart. As you know, today at 10.30 we experienced a fourth occurrence. This didn't match the previous pattern of 23 hours and given we had already replaced the hardware, this was looking more like a software problem. The decision was made to allow the downtime to continue longer than usual to allow Fortinet TAC to debug on the live unit, in the hope of nailing the cause once and for all. We had a cut off of 1.30pm, at which point we would reboot the box regardless of whether we had gotten to the bottom of the issue, so as not to disrupt afternoon lessons. One of our senior engineers held a teleconference with a senior TAC engineer and conducted the live troubleshooting - while myself, Mat and our Product Manager Rob went through the debug files. This time, the problem was different again - we were seeing our edge router was receiving no traffic from the WAN port of the Fortigate, however traffic captures on the Fortigate unit were showing ARP-Reply packets being sent. These tcpdumps capture from the kernel - not the ASIC/NIC (based on NP6 architecture) and so this left us with the possibility of a software issue affecting the ASIC/NIC, a problem with the SFPs in the ports or potentially an issue with the input on the edge router. I'm not ashamed to say that this point (close to 5pm), we all took a break (service was restored, and we were awaiting handover from the European service centre to the American service centre as this was still a P1 ticket). During the break, firmware version 5.2.10 was released, along with its release notes. This is the firmware version we were waiting for in relation to the miglogd issue, however, please find below some of other issues resolved in this version: 385115 Miglogd constantly crashing after upgrade to 5.2.8 386683 FG-1500D kernel panics after roughly 24 hours of uptime 387212 HA gets out of sync frequently and hasync becomes zombie 388032 Corrupted packets may cause malfunction of NP6, which causes NP ports to be unable to accept and forward traffic. Affected models: All NP6 platforms. 387675 ARP-Reply packets drops in NP6 For those who aren't aware of our setup - the device that has been having a problem is a 1500D. The 1500D platform is based on NP6 architecture. We have this evening updated to 5.2.10 after discussion internally and with TAC. At the moment, all looks stable and we are hopeful we will not see further issues - however I'm sure you can appreciate I cannot guarantee this. We will be monitoring closely tomorrow and if we do experience a further issue, the platform will be rebuilt on new routing hardware on Sunday evening. I do apologise for the inconvenience these issues have caused, we are working as hard as we can to resolve this as soon as possible. Kind Regards Lou Ashtonhurst NOC Manager Talk Straight/Schools Broadband -
Schools Broadband issues
lou-ashtonhurst replied to nicholab's topic in Internet Related/Filtering/Firewall
Unfortunately we have had a recurrence of the issue this morning at 10.30. We are currently working with senior engineers at Fortinet who are logged live into the kit - this does mean that the service will be impacted for a slightly longer period of time while we allow them to inspect what is happening. We appreciate this is not an acceptable level of service and we are doing everything in our power to restore stability as soon as possible. Thank you for your patience. Lou Ashtonhurst NOC Manager Talk Straight/Schools Broadband -
Schools Broadband issues
lou-ashtonhurst replied to nicholab's topic in Internet Related/Filtering/Firewall
Hi everyone Just to let you know, we did replace the hardware we suspected to be problematic with Filter 11 last night in a scheduled maintenance window. As with any maintenance, this is now under close supervision by the NOC team. While as standard we would not disclose information regarding major incidents in such a public forum, we do feel that given the severity of the situation everyone here does deserve some further information about what has happened, and what we are doing to resolve it. Yesterday we experienced an issue with one of our Fortigate clusters, as part of which some of our customers will have experienced downtime of up to 20minutes. After investigation by our NOC team, we discovered a problem with the High Availability (HA) setup of the cluster. We already had a maintenance booked in for last night to replace the hardware associated with the Filter 11 cluster issues, and so we decided to complete the work at the same time. The maintenance completed in the early hours and we were not expecting further issues - but of course this, as always, remains under observation and review. This afternoon we unexpectedly experienced a recurrence of the issue, and the HA for one of our Fortigate clusters failed. This wasn't expected, and we agree that is unacceptable. We collected all the information that Fortinet would require, and resolved the situation as quickly as we could. The situation is now escalated with Fortinet and we are working around the clock to get ensure we are all confident that this issue is fully resolved. Both myself as NOC Manager and Mat (our Technical Director) are currently in close proximity to one of our main datacentres to resolve these issues as quickly and efficiently as we can. Please be assured, we are here and we are listening - many of you will have spoken to one or both of us and we want to be clear that those avenues of communication are always open (but please, we are working overnight over the next few days - please give us until at least 10am). We are returning calls when we can, however, datacentres are giant Faraday cages which does inhibit signal somewhat - it may be best to leave us a message either via VM or our CTS team, who are passing your concerns on. If I haven't covered any of your specific concerns here, please do let me know. Kind Regards Lou Ashtonhurst NOC Manager - Schools Broadband
