Jump to content

Recommended Posts

Posted (edited)

Yup :( same error!

And of course bromcom service status page says AOK.

 

Status pages should be proactive, not reactive - it should know before we do, and that isn't difficult.

Edited by synaesthesia
Posted

I've only used Bromcom for about 3 years but from my memory every outage/poor performance... of which there has been too many... has been down to capacity limits.
Search for "root cause" on the bromcom community and you'll notice a pattern.
 

Posted
12 minutes ago, Synkrox said:

I've only used Bromcom for about 3 years but from my memory every outage/poor performance... of which there has been too many... has been down to capacity limits.
Search for "root cause" on the bromcom community and you'll notice a pattern.
 

I'm fairly sure that, if it was capacity, they'd solve it really easily and really quickly. Having heard what the CTO has to say on it and having read previous RCAs, Bromcom's challenge is similar to what SIMS went through (and arguably still is) - old and badly written / unfuture-proof code and coding practices. SIMS suffered more because they delayed doing anything about it for too long. At least Bromcom seems to be doing something about this early. 

You can throw hardware at software issues - but it only takes you so far. 

Posted

& again this morning, was fixed yesterday but happen again this morning

 

 

Their status page has been updated to acknowledge the issue 

Posted

If each non-green pip were representing 99% uptime that day, then you could round up favourably.

 

Am tempted to cobble together a script that will sign into Bromcom and check that pages are loading as expected, and plug that into our monitoring system. Do pages load as expected, and how long did it take. I don't use Bromcom much day to day, so often the first I know of an issue is several hours into the day, hearing people saying "Bromcom is down".

Posted

their "up time" is caculated by compleate outages (when its marked in red) not orange or light green

 

i want to know what they class as a compleate outage and marked in red as i would say yesterday and today morning was a compleate outage but still says that it was partual

Posted

Are you all on the same ISP? 429 is rate limiting, it could be that multiple schools sharing the same external IP (via a proxy or carrier-grade NAT) are collectively triggering the limit, rather than it being a Bromcom infrastructure problem per se. We've seen similar in the last few weeks with Google Admin Console.

Posted
Just now, paulkerton said:

Are you all on the same ISP? 429 is rate limiting, it could be that multiple schools sharing the same external IP (via a proxy or carrier-grade NAT) are collectively triggering the limit, rather than it being a Bromcom infrastructure problem per se. We've seen similar in the last few weeks with Google Admin Console.

We experienced it and our external IP is our own and not shared.

Posted

Last week or so we've seen it now about 4 days of issues. Usually back up within a short time but it is a bit worrying.

 

Before that though.. reliability had greatly improved since the Microsoft problem. The performance issues were gone. Generally reliability was there......  maybe me thinking about it.... jinxed things.

 

I'll always add big companies worth billions go down, last 2 years have seen some of the biggest providers cloudflare prime example. I'll always add that note when things like Bromcom suffer... but I just hope what ever the issue is.. they get it permanently resolved.

Posted
1 hour ago, bscott said:

their "up time" is caculated by compleate outages (when its marked in red) not orange or light green

 

i want to know what they class as a compleate outage and marked in red as i would say yesterday and today morning was a compleate outage but still says that it was partual

That's not how "proper" Cloud SaaS provider calculates the uptime for sure. 

Posted
1 minute ago, BlackCat80 said:

That's not how "proper" Cloud SaaS provider calculates the uptime for sure. 

It's how github do it. Since they moved to Azure (like bromcom), it's been so terrible that somebody created their own version https://mrshu.github.io/github-statuses/

 

Quote

GitHub stopped updating its status page with aggregate uptime numbers some time ago — if you use it regularly, you might have a feeling why. This is the missing version.

We rebuild platform‑wide and per‑service uptime from archived status updates, derive minute‑level downtime windows, and map incidents to services whenever the source data allows. Everything is open source, and PRs are very welcome!

Perhaps we could do one for bromcom

Posted (edited)
2 hours ago, dmj said:

It's how github do it. Since they moved to Azure (like bromcom), it's been so terrible that somebody created their own version https://mrshu.github.io/github-statuses/

 

Perhaps we could do one for bromcom

If we count Bromcom's degraded status as system outage, uptime reduces to 76%, which represents accurately I guess.

Edited by BlackCat80
  • Like 1

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...