BlackCat80 Posted March 9 Posted March 9 2 minutes ago, Cat_Jam148 said: I think we all just need to give this small startup a chance... They've only been in the MIS market since 1986 so they haven't had all that long to fix the infrastructure scaling issues...😁 *sarcasm* Very true, we should probably stop expecting them to be ready for something as unexpected as schools taking registers at 9am.
Wizard101 Posted March 9 Posted March 9 19 minutes ago, jthompson said: If the Bromcom web app were to get a couple of tweaks, I wonder whether that would significantly ease the demand on their resources. Namely, not opening up new tabs for things so readily, and not loading the full list of students/staff/groups before the end user has had a chance to provide and filters or search terms. Yeah just how it was fundamentally built and the optimisation seems poor. Like they didn't build for any scalability or understand how much data their system would be working with and didn't plan for it. Everything requires a new window to be opened up, It feels clunky like its barely standing when I use it, you make a report it feels like it's about to die. 1
Jaan Posted March 9 Posted March 9 (edited) So basically somebody has fugged the CPU/Memory buffers and probably left the default mem allocation in the default range of 512mb - 1048576 mb EDIT: oh and don't forget to add another single vcore Edited March 9 by Jaan
Synkrox Posted March 9 Posted March 9 Remember when it was falling over in Sept and their early fix was to keep turning it off and on again? 😂 1
psydii Posted March 9 Posted March 9 6 minutes ago, E_G_R2 said: Same here, the catering firm remoted in and reduced the number of requests. ..it could be just that. A 3rd party (or parties) made changes that increased the load, and that was enough to push it over the edge. It may be that that previous guardrails protecting against that had never been tested, or simply that limits were not enforced and had never been an issue previously. I'm pretty sure those of us who've nursed SIMS along for years have experienced the occasional performance hit relating to one integration or another. As for the scaling of resources. Bromcom is a business that operates in cash, with no debt. This is a philosophical/existential position for them, and throwing money at the problem is the route to debt. Arbor and SIMS are debt laden financial vehicles who happen to generate income from annual leasing of access to an MIS product to an almost captive audience. They are less afraid of spending now and worrying about it later. By which I mean, there is a risk that (in my opinion) their reliability and performance comes at the cost of monstrous price hikes in the future. 2
BlackCat80 Posted March 9 Posted March 9 9 minutes ago, psydii said: ..it could be just that. A 3rd party (or parties) made changes that increased the load, and that was enough to push it over the edge. It may be that that previous guardrails protecting against that had never been tested, or simply that limits were not enforced and had never been an issue previously. I'm pretty sure those of us who've nursed SIMS along for years have experienced the occasional performance hit relating to one integration or another. As for the scaling of resources. Bromcom is a business that operates in cash, with no debt. This is a philosophical/existential position for them, and throwing money at the problem is the route to debt. Arbor and SIMS are debt laden financial vehicles who happen to generate income from annual leasing of access to an MIS product to an almost captive audience. They are less afraid of spending now and worrying about it later. By which I mean, there is a risk that (in my opinion) their reliability and performance comes at the cost of monstrous price hikes in the future. If Bromcom has built a platform where one integration can help tip core morning operations over, then the engineering is not robust enough. And if “we prefer not to spend too much on infrastructure” is part of the explanation, then schools are effectively being told reliability was sacrificed to protect Bromcom’s philosophy. That is not admirable, it is unacceptable.
psydii Posted March 9 Posted March 9 Just now, BlackCat80 said: If Bromcom has built a platform where one integration can help tip core morning operations over, then the engineering is not robust enough. And if “we prefer not to spend too much on infrastructure” is part of the explanation, then schools are effectively being told reliability was sacrificed to protect Bromcom’s philosophy. That is not admirable, it is unacceptable. This is all my opinion: Its a choice, though perhaps one that most Bromcom customers aren't actually bought into. I am predicting that when the founder of Bromcom passes (or he finally divests), a private equity firm will swoop in, restructure the finances (as per Arbor and SIMS), and within 12-18 months all MIS providers will have tripled their prices. Bromcom is the only thing keeping the prices down. Occasionally unreliable performance vs £150K per annum. Lets see where we are in another decade. 1
BlackCat80 Posted March 9 Posted March 9 (edited) 37 minutes ago, psydii said: This is all my opinion: Its a choice, though perhaps one that most Bromcom customers aren't actually bought into. I am predicting that when the founder of Bromcom passes (or he finally divests), a private equity firm will swoop in, restructure the finances (as per Arbor and SIMS), and within 12-18 months all MIS providers will have tripled their prices. Bromcom is the only thing keeping the prices down. Occasionally unreliable performance vs £150K per annum. Lets see where we are in another decade. I respect your opinion and my opinion is that's the wrong trade-off. Schools should not have to choose between an affordable MIS and one that actually works at 9am. And let’s not pretend Bromcom is cheap out of charity. Schools are still paying serious money, so expecting reliability during morning registration is hardly unreasonable. “Be grateful it is not even more expensive” is not much of a defence when the service has been failing repeatedly during a core safeguarding process. My pick would be the first one. Edited March 9 by BlackCat80
tom_newton Posted March 9 Posted March 9 I came here to say that performance problems can be a real bear in the cloud, it's not usually just a case of "throw some more infrastructure" or "this is a terrible architecture", there can be some real interesting interactions between services, arguably harder to deal with when you are using increasingly cloud native stuff. But if it turns out it's a scale-up-in-the-morning issue, that is hard to stomach, pre-warming (scaling ahead of time) your servers is a good idea when you know traffic peaks are coming.
Popular Post Marci Posted March 9 Popular Post Posted March 9 (edited) ^^^ There are also duff practices going on in user space that aren’t helping matters… Lots out there using exclusively (or over-using) livefeeds to feed BI stacks, which are generated using main CloudMIS app QuickReport engine… when same sources are available from API or oData - generated from separate apps with separate resources… inflates the load on CloudMIS app due to not knowing any better. Stack these up and you suddenly have a *lot* of recurring load going on unnecessarily… with BC support not clued up enough to discourage bad practice (or advise on best practice) around feeds/oData/api use, and no public documentation covering the same. This *could* all be from one large-ish MAT or LA with an unfortunately timed refresh driven solely by quickreport-livefeeds, with all their schools refreshing the same data domain at the same time. You’d hope logging would identify, but past experience here says not when it comes to CloudMIS app. Lack of awareness at both ends means no controls in place to minimise abuse = everyone suffers. Someone out there’s put *something* live in the last 6 days that’s smacking everything against the guard rails, and I don’t think BC have entirely considered that this could be a customer, rather than anything their own teams have pushed to production… hence the head scratching, straw clutching, and eventual brute force “we’ve turn off all the brakes” as a “solution”. Edited March 9 by Marci 5
Marci Posted March 10 Posted March 10 It's still struggling this morning, with 429 errors showing in DevTools console... 1
garethEds Posted March 10 Posted March 10 (edited) Another issue this morning - but with the pupil application on Android (maybe iPhone) - all the modules have gone. My resident Year 8 Geek took great pride in showing me his phone with a broken Student App - and it is the new one. As for the main cloudmis site - the stylesheet seems to be broken on our side this morning - but yip, broken on the Welsh side as well. Edited March 10 by garethedmondson
Olliedawg Posted March 10 Posted March 10 Yep broken again. This is what my home page currently looks like
TheRobins Posted March 10 Posted March 10 They have made changes over the last week for the new Student App. I would also check your student app's, as they have now broken timetabling on the web based student portal as it shows the teachers full names. This has been reported on the community forums, we have disabled all student modules apart from home learning as it was riddled with errors such as the teachers name. I have also seen that students are able to browse teacher timetables, not the worst thing to happen but still not right. Gareth re the above, it maybe someone has disabled all of your modules for this reason,.
m_m Posted March 10 Posted March 10 (edited) They posted here yesterday at 4.30pm saying they had added additional compute manually as it wasn't auto-scaling quick enough. https://community.bromcomcloud.com/announcements/post/performance-issues-04-03-2026-under-investigation-wUoQ0AVwuUat8FM Edited March 10 by m_m
Olliedawg Posted March 10 Posted March 10 🤔 Looks like their manual scaling of resources didn't resolve the problem Quote To address this immediately, we have removed the scaling element from the platform and manually set the capacity to be more than required for handling tomorrow (Tuesday 10th March) morning’s peak. This ensures the system is operating with higher capacity from the start of the school day rather than waiting for automatic scaling to respond. 1
Theblacksheep Posted March 10 Posted March 10 Just now, Olliedawg said: Looks like their manual scaling of resources didn't resolve the problem Again.
CrootUK Posted March 10 Posted March 10 Looks like they've gone for the turn it off and on again trick!
synaesthesia Posted March 10 Posted March 10 Have they tried feeding the hamster? Or upgrading to a gerbil perhaps? 1
Steve21 Posted March 10 Posted March 10 2 minutes ago, Olliedawg said: Looks like their manual scaling of resources didn't resolve the problem No no you misunderstand... Manual scaling DID fix it, but then the customers refreshed too often and broke it again! Steve
supportman Posted March 10 Posted March 10 Tempted to apply for a job there haha. I saw they were looking for infrastructure engineers recently. Seems they are wedded to Microsoft Azure though which can be a pain for things like this. I already run a high capacity API and can see the problems with Bromcom from a mile off: - 3rd party catering providers should be accessing a replica DB API. - Scaling servers is important, but if they cant scale quick enough then you need to over provision. - Split the services up into defined availability centers. Registration would be my number one target to isolate to guarantee performance. - Seems like a lot of technical debt and bloat over time to me, Go back to basics and spend 6 months on a performance sprint company wide. - Increase internal real time monitoring. I've been using Releem recently and its been amazing to see database metrics and use AI to fix slow queries. Also finally if Azure isn't working for you then run your own infrastructure. There is a great article on how basecamp saved millions by doing exactly that. 1
Alis_Klar Posted March 10 Posted March 10 (edited) This must be costing them a small fortune in Azure credits? Their code needs performance tuning badly. If their primarily having issues with the massive leap in demand during registration can't they laser focus on streamlining the register page and all the database calls it calls. Once they've done that, focus on common office staff tasks like calling up pupil profiles for phone calls home and managing attendance. All these operation spike between 8:30 and 9:00 in most schools. Then implement API fair use and trottleing as @Marci was alluding to. Edited March 10 by Alis_Klar
Cat_Jam148 Posted March 10 Posted March 10 My sympathies to all the end users who are having to put up with this on a fairly regular basis! They're making SIMS look reliable. Actually, I don't think SIMS was ever this unreliable... Most of the issues I encountered with SIMS ultimately stemmed from changes I made in my younger, and less experienced, years. 1
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now