Jump to content

Recommended Posts

Posted

It's absolutely pointless having this page if it doesn't reflect current service status.

 

Just remove it......

 

Those using SSO to sign into Bromcom this morning might have noticed that you couldn't log in.

 

Not sure what the issue was, i assume it was something to do with the update that links SSO accounts to bromcom accounts or the maintenance they have scheduled.

 

 

So far we've had 3 instances where we have been unable to access our Bromcom instance since we went live in SEPT.

 

But at least the service status pages says it's been 100%.............

 

Rant over

  • Thanks 2
Posted
Agreed! You come into a 5**! storm of users not able to access the MIS, emailing, phoning and knocking on your door and by simply updating the status page will reduce the pressure as you know its not you.
Posted (edited)
Agreed! You come into a 5**! storm of users not able to access the MIS, emailing, phoning and knocking on your door and by simply updating the status page will reduce the pressure as you know its not you.

 

I had my team checking we hadn't nuked the Google accounts (a sync was done late yesterday), Checking Google Services Status pages, Our Admin panel API security Settings as well as checking directly within Bromcom.

 

Apparently Bromcom's Service status page needs to be Manually updated. I once asked support when we had a previous issue and the response was "I'll get the team to update the page" :rolleyes:

 

However, I would like to clarify........ I wouldn't go back to SIMs! :)

Edited by Jaan
Posted

 

However, I would like to clarify........ I wouldn't go back to SIMs! :)

 

Yes, count your blessings. SIMs have a habit of denying the problem exists and it must be your school, two hours later "oh yes we do have a problem" OR in the case of a random memory crash, just deny the issue exists for about 2 years then finally admit "oh it is a wider issue".

Posted
Since we've just comitted to migrating to BromComm by the end of next month - can anyone from BromCom give us some techie talk about how you're expanding your infrastructure to deal with an influx of new schools?
Posted
I had my team checking we hadn't nuked the Google accounts (a sync was done late yesterday), Checking Google Services Status pages, Our Admin panel API security Settings as well as checking directly within Bromcom.

 

Apparently Bromcom's Service status page needs to be Manually updated. I once asked support when we had a previous issue and the response was "I'll get the team to update the page" :rolleyes:

 

However, I would like to clarify........ I wouldn't go back to SIMs! :)

 

Thanks for the feedbacks everyone. I like to clear few confusions around our status page. Our status page is not updated manually, so first of all apologies about the previous response given to you by our support team.

 

The page is linked to independent monitoring tool that measures the platform's availability and response speeds.

Availability monitor and speed checks are done every minute by the monitoring software with 100+ agent monitors through the UK and pushes the data to monitoring software and results are reflected within the the status page:

 

There are configured triggers in place based on 100+ agent monitor stats collected, and results reflected automatically within status page - some basic triggers are as follow:

- if the system is not accessible by 10% or more monitors in last 5 attempts (with 1 minute interval checks) - then Major Outage status created automatically

- if the system is not accessible by 10% or less monitors in last 5 attempts (with 1 minute interval checks) - then Minor Outage status created automatically

- if the system is accessible by all monitors, but response times that received are higher than usual, then system checks how long this slow response continues, if more than 10 minutes for 50% of more agent monitors, then system automatically set Degraded Performance status.

 

The incidents you may have experienced in the past may not reach to the level to active the above triggers or the incidents were specific to your databases.

 

Today's incident related to SSO authentication was not an outage, it was not effected all users, and system was operational. When software bug identified that may effect many users, our Customer Services department send email out to inform all IT Managers, Data Managers and registered contacts in our CRM system. Our status page is not designed to report about bugs or specific/local issues. Today at 8.29am, email sent out from Customer Services to inform about this issue, and 8.52am another email sent out to all users to confirm that issue is resolved.

 

We will take on board all your comments and feedbacks where possible and operationally viable. Meanwhile, I hope this explains how our status page functions, and clears that it is not manually updated by someone at Bromcom. :)

  • Thanks 1
Posted
Thanks for the feedbacks everyone. I like to clear few confusions around our status page. Our status page is not updated manually, so first of all apologies about the previous response given to you by our support team.

 

The page is linked to independent monitoring tool that measures the platform's availability and response speeds.

Availability monitor and speed checks are done every minute by the monitoring software with 100+ agent monitors through the UK and pushes the data to monitoring software and results are reflected within the the status page:

 

There are configured triggers in place based on 100+ agent monitor stats collected, and results reflected automatically within status page - some basic triggers are as follow:

- if the system is not accessible by 10% or more monitors in last 5 attempts (with 1 minute interval checks) - then Major Outage status created automatically

- if the system is not accessible by 10% or less monitors in last 5 attempts (with 1 minute interval checks) - then Minor Outage status created automatically

- if the system is accessible by all monitors, but response times that received are higher than usual, then system checks how long this slow response continues, if more than 10 minutes for 50% of more agent monitors, then system automatically set Degraded Performance status.

 

The incidents you may have experienced in the past may not reach to the level to active the above triggers or the incidents were specific to your databases.

 

Today's incident related to SSO authentication was not an outage, it was not effected all users, and system was operational. When software bug identified that may effect many users, our Customer Services department send email out to inform all IT Managers, Data Managers and registered contacts in our CRM system. Our status page is not designed to report about bugs or specific/local issues. Today at 8.29am, email sent out from Customer Services to inform about this issue, and 8.52am another email sent out to all users to confirm that issue is resolved.

 

We will take on board all your comments and feedbacks where possible and operationally viable. Meanwhile, I hope this explains how our status page functions, and clears that it is not manually updated by someone at Bromcom. :)

 

That's absolute nonsense.

  • Thanks 2
Posted
Happy to elaborate on the points that does not make sense.

 

Isn’t it clear that Bromcom need to increase their monitoring points to pick up on stuff like SSO authentication mechanism breaking. Or just not publicly have the status page if they’re not willing to increase the points they monitor.

Posted

If a number of your customers aren't able to access a system (we'll use SSO in the case) because an authentication mechanism that's failed, and it isn't working as it is intended.....thats degraded performance.

 

It doesn't matter if "Bromcom still works"..... I can't log into Bromcom , its not working....... I shouldn't have to reconfigure our login methods to get our staff up and running again as a temporary solution. Its degraded performance and has an impact on the running of the school day.

 

Your response is the equivalent of saying "Well, its working here....." That doesn't help. You need to take into consideration that your customers have different setup and requirement needs.

 

If there's a issue affecting 10% of your customer base, surely it's just common courtesy to say "We got issue X that affecting a small %age of our users" and change the status page to "Degraded for users using X"

 

I'm curious, what number is 10% of your customer base?

 

Are you able to provide what percentage of your customers use SSO as their login method? it's the norm this day and age?

 

 

Sixty-Percent-of-the-Time.jpg

 

05onfire1_xp-superJumbo-v2.jpg

Posted (edited)
Today's incident related to SSO authentication was not an outage, it was not effected all users, and system was operational.

This is ridiculous. If an issue is affecting more than 1 customer, it is a partial outage. No-one else works like that? The system was not operational if people could not log in. This screams of trying to redefine what an outage is. To your customers? Its an outage.

 

To elaborate - what the *cause* of the partial outage was is irrelevant to the end user. Bug, natural disaster, alien attack. Doesn't matter. The end result is the same - customers not being able to log in using methods they have paid for as part of their contract. You even explicitly list SSO as a method of logging in as part of your sales stuff, so you cannot claim the system was working if SSO wasn't - SSO is part of your system.

Edited by localzuk
Posted
This is ridiculous. If an issue is affecting more than 1 customer, it is a partial outage. No-one else works like that? The system was not operational if people could not log in. This screams of trying to redefine what an outage is. To your customers? Its an outage.

 

To elaborate - what the *cause* of the partial outage was is irrelevant to the end user. Bug, natural disaster, alien attack. Doesn't matter. The end result is the same - customers not being able to log in using methods they have paid for as part of their contract. You even explicitly list SSO as a method of logging in as part of your sales stuff, so you cannot claim the system was working if SSO wasn't - SSO is part of your system.

 

Localzuk is bang on here. A outage that affects some customers is STILL an outage. A partial one yes, but you can't state that the system is operational and is fine. That it's working for other customers doesn't matter a jot to the customer that's affected & having major operational issues because of this.

 

This is basic stuff here, don't turn into Capita.

Posted
If a number of your customers aren't able to access a system (we'll use SSO in the case) because an authentication mechanism that's failed, and it isn't working as it is intended.....thats degraded performance.

 

It doesn't matter if "Bromcom still works"..... I can't log into Bromcom , its not working....... I shouldn't have to reconfigure our login methods to get our staff up and running again as a temporary solution. Its degraded performance and has an impact on the running of the school day.

 

Your response is the equivalent of saying "Well, its working here....." That doesn't help. You need to take into consideration that your customers have different setup and requirement needs.

 

If there's a issue affecting 10% of your customer base, surely it's just common courtesy to say "We got issue X that affecting a small %age of our users" and change the status page to "Degraded for users using X"

 

I'm curious, what number is 10% of your customer base?

 

Are you able to provide what percentage of your customers use SSO as their login method? it's the norm this day and age?

 

 

[ATTACH=CONFIG]64714[/ATTACH]

 

[ATTACH=CONFIG]64715[/ATTACH]

 

Degraded Performance relates to the response times, rather than failure of functionality (i.e. SSO).

 

Regarding to 10% benchmark, that's not related to customer base but configured monitor agents that polls the web application and collects the stats throughout the UK. I appreciate your feedback that you are not happy that you are not informed every issue via status page, but status page is designed to measure "availability" and "response performance", it is not designed to report every incident (currently), surely these feedbacks will be discussed internally and we will see how can be improved. Meanwhile, we will continue to inform the users with such incidents via email like we did yesterday, until we improve in this area.

Posted
If some people can't access your service, that's a partial outage. To claim otherwise and try to dress it up is nonsense. We aren't stupid, and this is the kind of thing people remember when evaluating which service to use.
Posted
Degraded Performance relates to the response times, rather than failure of functionality (i.e. SSO).

 

Regarding to 10% benchmark, that's not related to customer base but configured monitor agents that polls the web application and collects the stats throughout the UK. I appreciate your feedback that you are not happy that you are not informed every issue via status page, but status page is designed to measure "availability" and "response performance", it is not designed to report every incident (currently), surely these feedbacks will be discussed internally and we will see how can be improved. Meanwhile, we will continue to inform the users with such incidents via email like we did yesterday, until we improve in this area.

Well, it clearly wasn't available for users using SSO was it?
Posted
True. It was not available for some of the SSO users.

Precisely, so it was an outage.

 

Put it this way. If we as customers turn round to our staff if they can't get in because of such an issue and say "oh, the system is working fine, you just can't log in due to a bug", the staff would not use kind words in response. Its 1984 newspeak basically.

  • Thanks 2
Posted

I also don't have a horse in this race but it strikes me that the service page needs a rethink and a redesign. Currently it clearly doesn't do what customers expect it to do - specifically to report changes in service status.

 

Why not just add an area to display the notices which are also sent out by email, that would be a good start.

Posted
I'd like a change in threshold to see a change to the status. If maybe 10 customers report a broadly similar issue then change the status to something like 'investigating' and a brief description. You can easily cancel it out if it proves to be a coincidence or not what you thought. It gives people an indication that something might be going on, rather than just waiting until there is a fat outage to change status, by which time it is self evident there's a problem.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...