Jump to content

Recommended Posts

Posted
I have spoken directly with your CSM and the case has been escalated further and the case replied to. Rest assured we are working to resolve any outstanding issues.

 

Thanks

Darren Moody

Community Manager

 

Saying: "Rest assured we are working to resolve any outstanding issues" sounds like an empty promise at this point.

Could you please give us some dates when this will be sorted? Someone pointed out that the Bromcom slow performance was already reported in 2021...

"https://www.edugeek.net/forums/mis-s...eak-times.htmlDate of thread."

Posted
Saying: "Rest assured we are working to resolve any outstanding issues" sounds like an empty promise at this point.

Could you please give us some dates when this will be sorted? Someone pointed out that the Bromcom slow performance was already reported in 2021...

"https://www.edugeek.net/forums/mis-s...eak-times.htmlDate of thread."

 

From what I've heard they (Bromcom) have invested in dynamic infrastructure that allows them to deal with the workload better. At BETT they said that they have identified the issues occuring during peak times of AM and PM registration. Hopefully after February Half Term things will improve as that would have given Bromcom a few weeks to sort eveything out. The new schools from Northern Ireland apparently will be hosted on their own data centre so shouldnt affect the existing England schools.

Posted
Saying: "Rest assured we are working to resolve any outstanding issues" sounds like an empty promise at this point.

Could you please give us some dates when this will be sorted? Someone pointed out that the Bromcom slow performance was already reported in 2021...

"https://www.edugeek.net/forums/mis-s...eak-times.htmlDate of thread."

 

This is the official response for the performance issues last week, apologies if you have already seen it.

 

Major Incident Response - 22nd and 23rd January (bromcomcloud.com)

Thanks

Darren Moody

Community Manager

Posted

Could anyone post a copy here for us non-Bromcom folks (if allowed)?

 

Even though we aren't in the Bromcom ecosystem, I like to keep up-to-date with things such as these, and I'd like to see what their response was.

Posted
Could anyone post a copy here for us non-Bromcom folks (if allowed)?

 

Even though we aren't in the Bromcom ecosystem, I like to keep up-to-date with things such as these, and I'd like to see what their response was.

 

Major Incident Response - 22nd and 23rd January

Background

 

In the week prior to the issues that were experienced on Monday 22nd and Tuesday 23rd January 2024, Bromcom undertook work to improve the overall capacity, performance, and availability of our services through a planned series of infrastructure & software upgrades and improvements. These upgrades and improvements are necessary as we continue to scale our organisation and introduce new features and functionality.

 

Cause of incident(s)

 

Despite extensive testing and approval prior to release into production, the introduction of the additional capacity, designed to automatically scale as demand increases during peak hours, did not deliver the anticipated outcome during the morning of Monday 22nd January. Service was restored by manually increasing resources to meet demand. On Tuesday morning, similar issues occurred, but with a differed root cause, an application timeout in the Redis servers which were introduced to improve performance of availability and application response.

 

Service restoration

 

On Tuesday morning, the team noticed errors in the performance monitoring around 07:15 which could not be resolved with increased resources; therefore, a decision was taken to revert the system to its last known good configuration – a previous state in the week prior. This restored the system to a stable state.

 

Further precautions

 

To prevent any further issues with availability and to ensure stability of the system, a change freeze was enacted from Tuesday to avoid any further impact to customers whilst the incidents were investigated. This freeze lasted until Friday where our normal release schedule resumed.

 

Ongoing investigation

 

We are committed to ensuring our system capacity, availability and performance meets the demands of our customers, and we plan to continue with the programme of improvements. We have engaged additional support from our partners at Microsoft and Redis to assist us further in the investigation into the issues that caused the incidents. Recommendations have arisen from this which we will be carefully considered for how we move forward.

 

Communication

 

The current reporting of service incidents and degradations in performance and availability are mostly limited to manual updates. These take time to complete, and we recognise that as a result, customers are not always kept up to date in a timely manner. We are investigating options available to enhance and automate our status page at status.bromcom.com by enabling it to report Incidents where page response rate goes over a specific threshold.

 

Bromcom would like to apologise for the significant disruption last week. We understand it had a severe impact on teachers, students and wider teams associated with supporting them. The causes and future prevention of such issues in future are being carefully reviewed at the highest levels in Bromcom.

  • Thanks 1
Posted

[h=2]FYI - this is what the statement was:

 

Major Incident Response - 22nd and 23rd January[/h]

Background

In the week prior to the issues that were experienced on Monday 22nd and Tuesday 23rd January 2024, Bromcom undertook work to improve the overall capacity, performance, and availability of our services through a planned series of infrastructure & software upgrades and improvements. These upgrades and improvements are necessary as we continue to scale our organisation and introduce new features and functionality.

Cause of incident(s)

Despite extensive testing and approval prior to release into production, the introduction of the additional capacity, designed to automatically scale as demand increases during peak hours, did not deliver the anticipated outcome during the morning of Monday 22nd January. Service was restored by manually increasing resources to meet demand. On Tuesday morning, similar issues occurred, but with a differed root cause, an application timeout in the Redis servers which were introduced to improve performance of availability and application response.

Service restoration

On Tuesday morning, the team noticed errors in the performance monitoring around 07:15 which could not be resolved with increased resources; therefore, a decision was taken to revert the system to its last known good configuration – a previous state in the week prior. This restored the system to a stable state.

Further precautions

To prevent any further issues with availability and to ensure stability of the system, a change freeze was enacted from Tuesday to avoid any further impact to customers whilst the incidents were investigated. This freeze lasted until Friday where our normal release schedule resumed.

Ongoing investigation

We are committed to ensuring our system capacity, availability and performance meets the demands of our customers, and we plan to continue with the programme of improvements. We have engaged additional support from our partners at Microsoft and Redis to assist us further in the investigation into the issues that caused the incidents. Recommendations have arisen from this which we will be carefully considered for how we move forward.

Communication

The current reporting of service incidents and degradations in performance and availability are mostly limited to manual updates. These take time to complete, and we recognise that as a result, customers are not always kept up to date in a timely manner. We are investigating options available to enhance and automate our status page at status.bromcom.com by enabling it to report Incidents where page response rate goes over a specific threshold.

Bromcom would like to apologise for the significant disruption last week. We understand it had a severe impact on teachers, students and wider teams associated with supporting them. The causes and future prevention of such issues in future are being carefully reviewed at the highest levels in Bromcom.

  • Thanks 1
Posted
FYI - this is what the statement was:

 

Major Incident Response - 22nd and 23rd January

 

Background

In the week prior to the issues that were experienced on Monday 22nd and Tuesday 23rd January 2024, Bromcom undertook work to improve the overall capacity, performance, and availability of our services through a planned series of infrastructure & software upgrades and improvements. These upgrades and improvements are necessary as we continue to scale our organisation and introduce new features and functionality.

Cause of incident(s)

Despite extensive testing and approval prior to release into production, the introduction of the additional capacity, designed to automatically scale as demand increases during peak hours, did not deliver the anticipated outcome during the morning of Monday 22nd January. Service was restored by manually increasing resources to meet demand. On Tuesday morning, similar issues occurred, but with a differed root cause, an application timeout in the Redis servers which were introduced to improve performance of availability and application response.

Service restoration

On Tuesday morning, the team noticed errors in the performance monitoring around 07:15 which could not be resolved with increased resources; therefore, a decision was taken to revert the system to its last known good configuration – a previous state in the week prior. This restored the system to a stable state.

Further precautions

To prevent any further issues with availability and to ensure stability of the system, a change freeze was enacted from Tuesday to avoid any further impact to customers whilst the incidents were investigated. This freeze lasted until Friday where our normal release schedule resumed.

Ongoing investigation

We are committed to ensuring our system capacity, availability and performance meets the demands of our customers, and we plan to continue with the programme of improvements. We have engaged additional support from our partners at Microsoft and Redis to assist us further in the investigation into the issues that caused the incidents. Recommendations have arisen from this which we will be carefully considered for how we move forward.

Communication

The current reporting of service incidents and degradations in performance and availability are mostly limited to manual updates. These take time to complete, and we recognise that as a result, customers are not always kept up to date in a timely manner. We are investigating options available to enhance and automate our status page at status.bromcom.com by enabling it to report Incidents where page response rate goes over a specific threshold.

Bromcom would like to apologise for the significant disruption last week. We understand it had a severe impact on teachers, students and wider teams associated with supporting them. The causes and future prevention of such issues in future are being carefully reviewed at the highest levels in Bromcom.

 

The TLDR:

 

We don't do load testing, so had no idea we'd need to scale redis when we increased the cluster size. As we made a fix manually the same thing happened again, but this time we just increased the timeout so it takes longer to process things so we rolled back. Because our updates are manual, we don't have any oversight of the process and barely test them, we really need to get the senior devops post filled, but int he mean time we're going to pay MS and redis to suggest a redesign.

Posted
They mean https://status.bromcomcloud.com/ but yet again this is another rookie error on their side.

What bothers me is that all other countries are green, but their main UK market status page looks like a rainbow.

Also can't see Ireland on the list .... https://bromcom.com/news/bromcommis-northern-ireland

 

[ATTACH=CONFIG]70777[/ATTACH]

 

Did you realise everything before that red on Dec 13 for UK is 'No data' when you mouse hover. In fact the first three all green services are 100% no data.

So 'No data' is green :rolleyes:

  • Thanks 1
Posted

So 'No data' is green :rolleyes:

 

Despite extensive testing and approval prior to release into production,

 

"extensive testing"

  • Thanks 1
Posted
In fairness, NI probably falls under the "Bromcom MIS Cloud UK". I am also not surprised that the UK has the most downtime, as I presume it's by far the most busy of all of the other cloud instances.

The ink on the NI contract is barely dry, so I wouldn't expect it to appear on any dashboards yet, as any existing customers are almost certainly running out of the existing datacentres.

 

Moving forward if its contractually allowable, I'd expect the NI provision to run out of the datacentres south of the border, and have zero impact on system performance. Staffing of course is another matter entirely - but as observed earlier in this thread, they are hiring.

Posted
The ink on the NI contract is barely dry, so I wouldn't expect it to appear on any dashboards yet

 

NI schools are a job lot - they were all on SIMS and that contract win will move them all 1079 of them to Bromcom.

 

Should be a walk in the park for their first rate infrastructure team.

Posted
NI schools are a job lot - they were all on SIMS and that contract win will move them all 1079 of them to Bromcom.

 

Should be a walk in the park for their first rate infrastructure team.

 

Well when we signed up to Bromcom, they said that they constantly invest in the scalability of their cloud compute infrastructure. So when they add all 1079 of them to their system, the rest of us won't even notice...............

 

:getmecoat:

Posted
Well when we signed up to Bromcom, they said that they constantly invest in the scalability of their cloud compute infrastructure.

 

I think this is part of the problem, they shouldn't have to constantly 'do' anything - the infrastructure should scale with usage, it shouldn't be anything they even notice.

Posted
Well when we signed up to Bromcom, they said that they constantly invest in the scalability of their cloud compute infrastructure. So when they add all 1079 of them to their system, the rest of us won't even notice...............

 

:getmecoat:

 

Well you would hope so, I think someone else mentioned NI will be on separate infrastructure so it shouldn't have any effect in theory.

Posted
Yeah I think NI is on their own infrastructure and I'm sure they will have a lot of KPIs and quality controls to hit. The investment in quality to meet those targets should benefit all. Isn't this recent issue about a move from SaaS to PaaS? Which explains the complication somewhat. I'm hopeful that once that move is properly completed then these issues will massively abate.
  • Thanks 1
Posted
Well when we signed up to Bromcom, they said that they constantly invest in the scalability of their cloud compute infrastructure. So when they add all 1079 of them to their system, the rest of us won't even notice...............

 

:getmecoat:

 

Same - that's what the Sales guy told us when it was demo'd to us!

Posted
Yeah I think NI is on their own infrastructure and I'm sure they will have a lot of KPIs and quality controls to hit. The investment in quality to meet those targets should benefit all.

 

It won't benefit all if it's on separate "high uptime" infrastructure!

Posted
It won't benefit all if it's on separate "high uptime" infrastructure!

I was just meaning the investment in the software side will benefit all. I.e. a bigger budget for testing, dev, quality controls, etc.. Infrastructure alone isn't sufficient to meet stringent quality controls so I'm just seeing this as a potential positive for the rest of their client base.

Posted (edited)

From the outside In I think Bromcom need congratulating for the NI contract and I am sure provisions are being put in place their end to ensure it doesn't hurt existing customers.

 

The growth of both Arbor and Bromcom in the MIS market needs commending as both are clearly doing a lot right.

Edited by TheCookieMonster
  • Thanks 1
Posted
From the outside In I think Bromcom need congratulating for the NI contract and I am sure provisions are being put in place their end to ensure it doesn't hurt existing customers.

 

The growth of both Arbor and Bromcom in the MIS market needs commending as both are clearly doing a lot right.

 

To be fair i totally agree with this. They do get it Wrong sometimes, and clear communication when things "go wrong" can be lacking. That being said i could NEVER go back to SIMs. The difference in features and manageability ability is Day and Night.

 

This issue with expanding a customer base so fast mean so do the number of pitchforks!

  • Thanks 1
Posted
From the outside In I think Bromcom need congratulating for the NI contract and I am sure provisions are being put in place their end to ensure it doesn't hurt existing customers.

 

The growth of both Arbor and Bromcom in the MIS market needs commending as both are clearly doing a lot right.

 

This.

 

After speaking to a few key staff at Bromcom at BETT and I've been vocal of my critisism, I'm satisfied they also know where they have fallen short and what they need to do to correct their approach. Its how how they go about doing this that matters.

  • Thanks 3

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...