Jobos Posted September 23, 2021 Posted September 23, 2021 Around 4-5 years ago we started using a managed VoIP from Orbtalk. At first it gave problems with bad call quality but after changing our ISP to Wave9 and using a Sophos XG things were rock solid. Since June we have started having problems again but this time lines go offline when viewed in the Orbtalk portal and any calls to the offline numbers are either transferred to a mobile number that we set for VoIP outages or to voicemail depending on the line. Not all lines go offline and sometimes lines that went offline stay on when others go off. Strange thing is the offline number can make calls and the portal shows them as making a call but they cannot receive calls. They come back on their own or when the the handset is rebooted. Orbtalk support have investigated and said this We can see that some of your handsets became unreachable around that period and have experienced similar drops since, please see below. [Jul 1 10:34:01] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now UNREACHABLE! Last qualify: 23 [Jul 1 10:34:03] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now UNREACHABLE! Last qualify: 28 [Jul 1 10:34:21] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now UNREACHABLE! Last qualify: 37 [Jul 1 10:34:25] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now UNREACHABLE! Last qualify: 58 [Jul 1 10:34:28] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now UNREACHABLE! Last qualify: 38 [Jul 1 10:34:31] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now UNREACHABLE! Last qualify: 44 [Jul 1 10:34:42] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now Reachable. (59ms / 2000ms) [Jul 1 10:34:44] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now UNREACHABLE! Last qualify: 36 [Jul 1 10:34:46] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now UNREACHABLE! Last qualify: 42 [Jul 1 10:34:46] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now UNREACHABLE! Last qualify: 24 [Jul 1 10:34:56] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now UNREACHABLE! Last qualify: 71 [Jul 1 10:34:57] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now Reachable. (57ms / 2000ms) [Jul 1 10:35:36] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now Reachable. (64ms / 2000ms) [Jul 1 10:36:07] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now Reachable. (64ms / 2000ms) [Jul 1 10:47:23] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now Reachable. (36ms / 2000ms) [Jul 1 11:02:42] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now Reachable. (27ms / 2000ms) [Jul 1 11:03:59] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now Reachable. (29ms / 2000ms) [Jul 1 11:08:56] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now Reachable. (25ms / 2000ms) [Jul 1 11:09:39] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now Reachable. (31ms / 2000ms) [Jul 1 11:19:17] NOTICE[1803] chan_sip.c: Peer 'xxxxxxx' is now Reachable. (35ms / 2000ms) The unreachable status indicates that the handset has failed to respond to the Keep Alives the PBX sends and will result in the PBX marking them as offline. This is typically indicative of a network or connectivity issue affecting traffic. Wave9 cannot see anything wrong with the XG. For much of the time everything is fine just now and again lines dropout. Any pointers on where to start looking?
psydii Posted September 23, 2021 Posted September 23, 2021 We had similar with another provider. So many variables in the mix, but the VoIP provider were confident it was a local network problem. One day I caught a switch misbehaving - packets just stopped flowing for a few minutes. Nothing in the logs and no NMS alarms to indicate there was any problem - just devices on that switch couldn't get packets in or out for a bit. Swapped out that switch and the problems went away. Swapped out other switches that had the misbehaving phones and those problems also went away. Really irritating. (but luckily we were in year where switch refresh was in the budget) I should say that when we first had problems it was also a local network problem - a link port would occasionally flap for a minute and cause spanning tree recalculations. Most devices didn't exhibit noticeable issues, but the phones in that part of the network would drop. That one was easier to spot.
mavhc Posted September 23, 2021 Posted September 23, 2021 Could be something weird on the switch, true. I'd run continuous ping tests internally to all the phones, see if you can catch them in the act. What switches do you have? Do they have error logs, make sure the times and synced and then try to find concurrent events.
TechMonkey Posted September 23, 2021 Posted September 23, 2021 Interestingly we are with OrbTalk and get something similar occasionally with our SIP lines. 3CX will say it isn't available, their support says everything is OK and no amount of reconnecting does anything. If you ignore it (which is hard when everyone is complaining about no phones) it eventually reconnects. Often a reboot sorts it.
mavhc Posted September 23, 2021 Posted September 23, 2021 If it's started in the past month, there's a lot of DDoS on VoIP going on currently https://www.uctoday.com/unified-communications/3-uk-voip-providers-hit-by-ddos-attacks/ There's no cloudflare ddos type protection for voip service available, and the attackers have realised this.
Jobos Posted September 23, 2021 Author Posted September 23, 2021 Most of the cabling dates back to when the school was with Virgin and goes back to one cab. It made sense to keep the existing cabling and add a new switch, an Aurba 2920-48. There are 12 phones connected and nearly all the other connections are used for data so maybe the switch can't cope with the throughput. As time went on and another handset was needed we just connected it to the nearest network point so just spent a little time double checking the network paperwork and one of the problem phones is connected to a different switch in a different part of the school. The handsets are Gigabit T46, T42 and the cheap and nasty Gigaset N300 which funny enough seem to stay online when the others have not.
psydii Posted September 23, 2021 Posted September 23, 2021 The Aruba 2920 is a beefy switch and more than capable of handling VOIP and data. Probably worth looking at the link to that other switch, both physically and have a peek at the traffic flowing over it - perhaps is becomes full momentarily? Does that switch have a 1gb+ uplink to the rest of the network? Is broadcast storm conntrol / traffic limiters configured? The ping test would be a good place to start. I would add pings to each of the switches in the chain, and also your FW and DNS servers - just to see if there are any co-incident issues.
Jobos Posted September 23, 2021 Author Posted September 23, 2021 Sorry I gave the wrong information. The Aruba 2920 is the core and switch in question is a HP 1910-48.
psydii Posted September 23, 2021 Posted September 23, 2021 (edited) The HP 1910-48 is in fact the HPE A5120-SI with restricted UI. Should also be more than capable of handling the traffic. The only potential limitation that it is possible to hit (but only in the largest of schools) is the mac address limit of 8000. The Aruba 2920 has a 16000 mac address limit (and by the by, not that it is relevant here, so does the A5120-SI's big brother the HPE-A5120-EI). That said the 1910 is quite old now, so it might be feeling its age? Edited September 23, 2021 by psydii
Michael Posted September 23, 2021 Posted September 23, 2021 Maybe for a future consideration, but I can recommend Wave 9 VOIP
Jobos Posted September 23, 2021 Author Posted September 23, 2021 Maybe for a future consideration, but I can recommend Wave 9 VOIP We have thought about it but Wave9 only do VoIP with the handsets and as we have already bought decent Yealink handsets it seems like we would be paying twice. Had we looked around in the first place things might have been different. Incidentally, Orbtalk were recommended by Virgin as a VoIP supplier (I should have known it would end in tears).
Jobos Posted September 27, 2021 Author Posted September 27, 2021 Right more information and just wanted to get your opinion before I contact Orbtalk as this might have a bearing on our dropped lines. Came in school today and there is an email from the office manager explaining that she received a call on Friday from someone trying to call a food supplier. It appears the caller got to the food suppliers IVR and selected the option for accounts but was transferred to the office managers phone rather than the food suppliers. The food suppliers number is nothing like the schools so it wouldn't be a because of misdialing and the callers number is not listed in the call logs.
TechMonkey Posted September 27, 2021 Posted September 27, 2021 That sounds like a coincidence or red herring. Your log shows a peer uncontactable then contactable, so the peer is available some of the time, not just unavailable when someone dials a wrong extension.
Jobos Posted October 7, 2021 Author Posted October 7, 2021 Update: Finally got to the bottom of why the handsets were dropping connections and it turned out to the Sophos XG. It seems when teachers connected or disconnected using their IPSec VPN it caused the XG some sort reset and was enough to drop the phones. Most of the time it was in the evening but some teachers work at home in their PPA time and this was more noticeable. After I pointed it out to Wave9 they fixed it within minutes and it's been ok since.
Jobos Posted October 7, 2021 Author Posted October 7, 2021 set vpn conn-remove-tunnel-up disable Info about the command here https://community.sophos.com/sophos-xg-firewall/f/discussions/112066/sophos-connect-vpn-vs-voip 1
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now