Hi folks
I thought I'd just drop you all a note with a bit more of an explanation of today’s incident, the timeline of events so far and why we are in the position we are now in. I'll leave the commercial discussions to Dave - this post is purely about the tech side of things for those that are interested.
The first incident occurred on 23/11 at 2pm, the second on the 24th at 1pm, the third at near enough midday on the 25th - a separation of approximately 23 hours each time. After the first incident, we reported it through to Fortinet and awaited their debug - but we weren't overly concerned. These things do happen and initially this looked like a random, one off occurrence - all our standard troubleshooting showed no problems and while we awaited Fortinet's diagnosis, we put this down to a random unit crash. As I mentioned in the earlier post, we were made aware of problems with the HA after the second occurrence which we resolved that night in an emergency maintenance slot. The next day we had the third occurrence - after which we replaced the hardware. Each of these steps was taken as soon as possible in an emergency maintenance window and on the basis of decisions of myself, our Technical Director, our senior engineers and Fortinet's senior engineers.
We had been aware of a potential issue with a process called miglogd, which since the update to firmware 5.2.8 had been regularly crashing (but not with significant impact to the box). Over the last few weeks we have implemented a workaround on Fortinet's recommendation, knowing that a new firmware release was required to resolve it and was in the works. As with any vendor, Fortinet quite rightly run significant testing before any firmware upgrade is rolled out.
The debug output from the first night led us to miglogd, the second night showed HA problems, the third we suspected a hardware fault on the ASIC as traffic seemingly wasn't being forwarded from the WAN port. These were all approx 23 hours apart.
As you know, today at 10.30 we experienced a fourth occurrence. This didn't match the previous pattern of 23 hours and given we had already replaced the hardware, this was looking more like a software problem. The decision was made to allow the downtime to continue longer than usual to allow Fortinet TAC to debug on the live unit, in the hope of nailing the cause once and for all. We had a cut off of 1.30pm, at which point we would reboot the box regardless of whether we had gotten to the bottom of the issue, so as not to disrupt afternoon lessons.
One of our senior engineers held a teleconference with a senior TAC engineer and conducted the live troubleshooting - while myself, Mat and our Product Manager Rob went through the debug files. This time, the problem was different again - we were seeing our edge router was receiving no traffic from the WAN port of the Fortigate, however traffic captures on the Fortigate unit were showing ARP-Reply packets being sent. These tcpdumps capture from the kernel - not the ASIC/NIC (based on NP6 architecture) and so this left us with the possibility of a software issue affecting the ASIC/NIC, a problem with the SFPs in the ports or potentially an issue with the input on the edge router. I'm not ashamed to say that this point (close to 5pm), we all took a break (service was restored, and we were awaiting handover from the European service centre to the American service centre as this was still a P1 ticket).
During the break, firmware version 5.2.10 was released, along with its release notes. This is the firmware version we were waiting for in relation to the miglogd issue, however, please find below some of other issues resolved in this version:
385115 Miglogd constantly crashing after upgrade to 5.2.8
386683 FG-1500D kernel panics after roughly 24 hours of uptime
387212 HA gets out of sync frequently and hasync becomes zombie
388032 Corrupted packets may cause malfunction of NP6, which causes NP ports to be unable to accept and forward traffic. Affected models: All NP6 platforms.
387675 ARP-Reply packets drops in NP6
For those who aren't aware of our setup - the device that has been having a problem is a 1500D. The 1500D platform is based on NP6 architecture.
We have this evening updated to 5.2.10 after discussion internally and with TAC. At the moment, all looks stable and we are hopeful we will not see further issues - however I'm sure you can appreciate I cannot guarantee this.
We will be monitoring closely tomorrow and if we do experience a further issue, the platform will be rebuilt on new routing hardware on Sunday evening.
I do apologise for the inconvenience these issues have caused, we are working as hard as we can to resolve this as soon as possible.
Kind Regards
Lou Ashtonhurst
NOC Manager
Talk Straight/Schools Broadband