Jaan Posted February 22, 2022 Posted February 22, 2022 It's absolutely pointless having this page if it doesn't reflect current service status. Just remove it...... Those using SSO to sign into Bromcom this morning might have noticed that you couldn't log in. Not sure what the issue was, i assume it was something to do with the update that links SSO accounts to bromcom accounts or the maintenance they have scheduled. So far we've had 3 instances where we have been unable to access our Bromcom instance since we went live in SEPT. But at least the service status pages says it's been 100%............. Rant over 2
E_G_R2 Posted February 22, 2022 Posted February 22, 2022 Agreed! You come into a 5**! storm of users not able to access the MIS, emailing, phoning and knocking on your door and by simply updating the status page will reduce the pressure as you know its not you.
Jaan Posted February 22, 2022 Author Posted February 22, 2022 (edited) Agreed! You come into a 5**! storm of users not able to access the MIS, emailing, phoning and knocking on your door and by simply updating the status page will reduce the pressure as you know its not you. I had my team checking we hadn't nuked the Google accounts (a sync was done late yesterday), Checking Google Services Status pages, Our Admin panel API security Settings as well as checking directly within Bromcom. Apparently Bromcom's Service status page needs to be Manually updated. I once asked support when we had a previous issue and the response was "I'll get the team to update the page" However, I would like to clarify........ I wouldn't go back to SIMs! Edited February 22, 2022 by Jaan
mthomas08 Posted February 22, 2022 Posted February 22, 2022 However, I would like to clarify........ I wouldn't go back to SIMs! Yes, count your blessings. SIMs have a habit of denying the problem exists and it must be your school, two hours later "oh yes we do have a problem" OR in the case of a random memory crash, just deny the issue exists for about 2 years then finally admit "oh it is a wider issue".
bobsmith Posted February 22, 2022 Posted February 22, 2022 Since we've just comitted to migrating to BromComm by the end of next month - can anyone from BromCom give us some techie talk about how you're expanding your infrastructure to deal with an influx of new schools?
ultimatepie Posted February 22, 2022 Posted February 22, 2022 I understand your Pain http://www.edugeek.net/forums/mis-systems/226127-bromcom-status-page-pointless.html
Bromcom_Hus Posted February 22, 2022 Posted February 22, 2022 I had my team checking we hadn't nuked the Google accounts (a sync was done late yesterday), Checking Google Services Status pages, Our Admin panel API security Settings as well as checking directly within Bromcom. Apparently Bromcom's Service status page needs to be Manually updated. I once asked support when we had a previous issue and the response was "I'll get the team to update the page" However, I would like to clarify........ I wouldn't go back to SIMs! Thanks for the feedbacks everyone. I like to clear few confusions around our status page. Our status page is not updated manually, so first of all apologies about the previous response given to you by our support team. The page is linked to independent monitoring tool that measures the platform's availability and response speeds. Availability monitor and speed checks are done every minute by the monitoring software with 100+ agent monitors through the UK and pushes the data to monitoring software and results are reflected within the the status page: There are configured triggers in place based on 100+ agent monitor stats collected, and results reflected automatically within status page - some basic triggers are as follow: - if the system is not accessible by 10% or more monitors in last 5 attempts (with 1 minute interval checks) - then Major Outage status created automatically - if the system is not accessible by 10% or less monitors in last 5 attempts (with 1 minute interval checks) - then Minor Outage status created automatically - if the system is accessible by all monitors, but response times that received are higher than usual, then system checks how long this slow response continues, if more than 10 minutes for 50% of more agent monitors, then system automatically set Degraded Performance status. The incidents you may have experienced in the past may not reach to the level to active the above triggers or the incidents were specific to your databases. Today's incident related to SSO authentication was not an outage, it was not effected all users, and system was operational. When software bug identified that may effect many users, our Customer Services department send email out to inform all IT Managers, Data Managers and registered contacts in our CRM system. Our status page is not designed to report about bugs or specific/local issues. Today at 8.29am, email sent out from Customer Services to inform about this issue, and 8.52am another email sent out to all users to confirm that issue is resolved. We will take on board all your comments and feedbacks where possible and operationally viable. Meanwhile, I hope this explains how our status page functions, and clears that it is not manually updated by someone at Bromcom. 1
Primus Posted February 23, 2022 Posted February 23, 2022 Thanks for the feedbacks everyone. I like to clear few confusions around our status page. Our status page is not updated manually, so first of all apologies about the previous response given to you by our support team. The page is linked to independent monitoring tool that measures the platform's availability and response speeds. Availability monitor and speed checks are done every minute by the monitoring software with 100+ agent monitors through the UK and pushes the data to monitoring software and results are reflected within the the status page: There are configured triggers in place based on 100+ agent monitor stats collected, and results reflected automatically within status page - some basic triggers are as follow: - if the system is not accessible by 10% or more monitors in last 5 attempts (with 1 minute interval checks) - then Major Outage status created automatically - if the system is not accessible by 10% or less monitors in last 5 attempts (with 1 minute interval checks) - then Minor Outage status created automatically - if the system is accessible by all monitors, but response times that received are higher than usual, then system checks how long this slow response continues, if more than 10 minutes for 50% of more agent monitors, then system automatically set Degraded Performance status. The incidents you may have experienced in the past may not reach to the level to active the above triggers or the incidents were specific to your databases. Today's incident related to SSO authentication was not an outage, it was not effected all users, and system was operational. When software bug identified that may effect many users, our Customer Services department send email out to inform all IT Managers, Data Managers and registered contacts in our CRM system. Our status page is not designed to report about bugs or specific/local issues. Today at 8.29am, email sent out from Customer Services to inform about this issue, and 8.52am another email sent out to all users to confirm that issue is resolved. We will take on board all your comments and feedbacks where possible and operationally viable. Meanwhile, I hope this explains how our status page functions, and clears that it is not manually updated by someone at Bromcom. That's absolute nonsense. 2
Bromcom_Hus Posted February 23, 2022 Posted February 23, 2022 That's absolute nonsense. Happy to elaborate on the points that does not make sense.
MicrodigitUK Posted February 23, 2022 Posted February 23, 2022 Happy to elaborate on the points that does not make sense. Isn’t it clear that Bromcom need to increase their monitoring points to pick up on stuff like SSO authentication mechanism breaking. Or just not publicly have the status page if they’re not willing to increase the points they monitor.
Jaan Posted February 23, 2022 Author Posted February 23, 2022 If a number of your customers aren't able to access a system (we'll use SSO in the case) because an authentication mechanism that's failed, and it isn't working as it is intended.....thats degraded performance. It doesn't matter if "Bromcom still works"..... I can't log into Bromcom , its not working....... I shouldn't have to reconfigure our login methods to get our staff up and running again as a temporary solution. Its degraded performance and has an impact on the running of the school day. Your response is the equivalent of saying "Well, its working here....." That doesn't help. You need to take into consideration that your customers have different setup and requirement needs. If there's a issue affecting 10% of your customer base, surely it's just common courtesy to say "We got issue X that affecting a small %age of our users" and change the status page to "Degraded for users using X" I'm curious, what number is 10% of your customer base? Are you able to provide what percentage of your customers use SSO as their login method? it's the norm this day and age?
localzuk Posted February 23, 2022 Posted February 23, 2022 (edited) Today's incident related to SSO authentication was not an outage, it was not effected all users, and system was operational. This is ridiculous. If an issue is affecting more than 1 customer, it is a partial outage. No-one else works like that? The system was not operational if people could not log in. This screams of trying to redefine what an outage is. To your customers? Its an outage. To elaborate - what the *cause* of the partial outage was is irrelevant to the end user. Bug, natural disaster, alien attack. Doesn't matter. The end result is the same - customers not being able to log in using methods they have paid for as part of their contract. You even explicitly list SSO as a method of logging in as part of your sales stuff, so you cannot claim the system was working if SSO wasn't - SSO is part of your system. Edited February 23, 2022 by localzuk
DrCheese Posted February 23, 2022 Posted February 23, 2022 This is ridiculous. If an issue is affecting more than 1 customer, it is a partial outage. No-one else works like that? The system was not operational if people could not log in. This screams of trying to redefine what an outage is. To your customers? Its an outage. To elaborate - what the *cause* of the partial outage was is irrelevant to the end user. Bug, natural disaster, alien attack. Doesn't matter. The end result is the same - customers not being able to log in using methods they have paid for as part of their contract. You even explicitly list SSO as a method of logging in as part of your sales stuff, so you cannot claim the system was working if SSO wasn't - SSO is part of your system. Localzuk is bang on here. A outage that affects some customers is STILL an outage. A partial one yes, but you can't state that the system is operational and is fine. That it's working for other customers doesn't matter a jot to the customer that's affected & having major operational issues because of this. This is basic stuff here, don't turn into Capita.
Bromcom_Hus Posted February 23, 2022 Posted February 23, 2022 If a number of your customers aren't able to access a system (we'll use SSO in the case) because an authentication mechanism that's failed, and it isn't working as it is intended.....thats degraded performance. It doesn't matter if "Bromcom still works"..... I can't log into Bromcom , its not working....... I shouldn't have to reconfigure our login methods to get our staff up and running again as a temporary solution. Its degraded performance and has an impact on the running of the school day. Your response is the equivalent of saying "Well, its working here....." That doesn't help. You need to take into consideration that your customers have different setup and requirement needs. If there's a issue affecting 10% of your customer base, surely it's just common courtesy to say "We got issue X that affecting a small %age of our users" and change the status page to "Degraded for users using X" I'm curious, what number is 10% of your customer base? Are you able to provide what percentage of your customers use SSO as their login method? it's the norm this day and age? [ATTACH=CONFIG]64714[/ATTACH] [ATTACH=CONFIG]64715[/ATTACH] Degraded Performance relates to the response times, rather than failure of functionality (i.e. SSO). Regarding to 10% benchmark, that's not related to customer base but configured monitor agents that polls the web application and collects the stats throughout the UK. I appreciate your feedback that you are not happy that you are not informed every issue via status page, but status page is designed to measure "availability" and "response performance", it is not designed to report every incident (currently), surely these feedbacks will be discussed internally and we will see how can be improved. Meanwhile, we will continue to inform the users with such incidents via email like we did yesterday, until we improve in this area.
bald_pig Posted February 23, 2022 Posted February 23, 2022 If some people can't access your service, that's a partial outage. To claim otherwise and try to dress it up is nonsense. We aren't stupid, and this is the kind of thing people remember when evaluating which service to use.
bald_pig Posted February 23, 2022 Posted February 23, 2022 Degraded Performance relates to the response times, rather than failure of functionality (i.e. SSO). Regarding to 10% benchmark, that's not related to customer base but configured monitor agents that polls the web application and collects the stats throughout the UK. I appreciate your feedback that you are not happy that you are not informed every issue via status page, but status page is designed to measure "availability" and "response performance", it is not designed to report every incident (currently), surely these feedbacks will be discussed internally and we will see how can be improved. Meanwhile, we will continue to inform the users with such incidents via email like we did yesterday, until we improve in this area.Well, it clearly wasn't available for users using SSO was it?
Bromcom_Hus Posted February 23, 2022 Posted February 23, 2022 Well, it clearly wasn't available for users using SSO was it? True. It was not available for some of the SSO users.
localzuk Posted February 23, 2022 Posted February 23, 2022 True. It was not available for some of the SSO users. Precisely, so it was an outage. Put it this way. If we as customers turn round to our staff if they can't get in because of such an issue and say "oh, the system is working fine, you just can't log in due to a bug", the staff would not use kind words in response. Its 1984 newspeak basically. 2
Popular Post Garacesh Posted February 23, 2022 Popular Post Posted February 23, 2022 (edited) I appreciate your feedback that you are not happy that you are not informed every issue via status page, but status page is designed to measure "availability" and "response performance", it is not designed to report every incident (currently) lmao we don't even use Bromcom here so I got no dog in this fight but this is laughably incompetent. "Availability" Y'know what not-being-able-to-log-in makes your service? The exact opposite of available. Now I don't know the details of this outage (and yes, this is a partial outage), whether it was your end or whoever handles the SSO (e.g. Sign in with Google etc), but it's clear that the automation providing the status page isn't testing enough stuff. And I mean to some degree, I get it. I don't always remember to take the time to tell my users when there's a problem, usually because I'm too busy trying to fix it, and any time an engineer is taking to explain a problem to somebody to update the website is time they're not taking to fix it. But when you're an organisation your size, with (presumably) considerably large teams, and you're running a product that you know is such an essential, core service to your clients.. Yeah, you should probably either add more automatic checks, or get some manual bits too. Or both. Both is good. Edited February 23, 2022 by Garacesh 5
Popular Post CyBeRkId2002 Posted February 28, 2022 Popular Post Posted February 28, 2022 Playing devils advocate here I can see both sides. The service status page should be a central tool to find out any issues that the platform is experiencing, and I would think it would be important to have someone in the customer care team with the responsibility of ensuring it is updated, even when automated measures are not available (I would imagine this is extremely hard to choose for every scenario, particularly something that relies on third parties, such as SSO. Having said that I read @HuseiynGuryel response as an explanation of how that service page works currently. I didn't read it as an attempt to minimise the situation, nor did I read it as someone who was playing semantics on what an outage was, just what the service page lists as an outage (this comes from someone who really did have trouble with a company playing semantics with their SLA in the past. My feedback going forward would be: There should be a named member of the team everyday with the responsibility of adding service status that significantly impact a large proportion of your customers. The uptime/availability stats do need a rethink to try and somehow incorporate these kind of events to give a true reflection. Having said all of that an email on this occasion did go out promptly and overall we have been very happy with the uptime in our three years with the system (I think I remember 3 separate occasions - all less than 30 minutes or so) 7
ultimatepie Posted March 1, 2022 Posted March 1, 2022 Status page again not working on database performance issues on VM-UKS-WebL03, any other schools on this host?
JRowley Posted March 8, 2022 Posted March 8, 2022 I also don't have a horse in this race but it strikes me that the service page needs a rethink and a redesign. Currently it clearly doesn't do what customers expect it to do - specifically to report changes in service status. Why not just add an area to display the notices which are also sent out by email, that would be a good start.
Oaktech Posted March 8, 2022 Posted March 8, 2022 I'd like a change in threshold to see a change to the status. If maybe 10 customers report a broadly similar issue then change the status to something like 'investigating' and a brief description. You can easily cancel it out if it proves to be a coincidence or not what you thought. It gives people an indication that something might be going on, rather than just waiting until there is a fat outage to change status, by which time it is self evident there's a problem.
Jaan Posted March 16, 2022 Author Posted March 16, 2022 So the service status page has had a update? https://status.bromcomcloud.com/ 3
supportman Posted March 16, 2022 Posted March 16, 2022 So the service status page has had a update? https://status.bromcomcloud.com/ Looks very nice! Loving the susbscribe to updates feature, it even lets you add planned maintance into a calendar. Very cool. 1
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now