Jump to content

Recommended Posts

Posted
Can anyone else confirm if it is affecting hosted / connected Sims? Do they use Azure servers? Ours is working very intermittently, and their status page appears to have gone completely! Causing chaos with cover.
Posted

Opened up my inbox this morning to a blizzard of SAN error alarms, and then it went silent.

 

SAN turns out was 'fine' (degraded but fine), it was just that 365 went down in the middle of the SAN telling me its problems making it look like the SAN had completely died.

 

Teams and Outlook were back up for us by about 0820UTC, most people have their files cached by the onedrive client so we are shielded from the current lingering problems on the sharepoint side of the farm.

 

SIMS sits on our SAN so I thought I was going to have a very bad day.

Posted

Microsoft says it has "rolled back a network change" it thinks may be linked to tens of thousands of worldwide users being unable to access its services, including Teams and Outlook.

 

Downdetector, which tracks website outages, showed more than 5,000 people in the UK had reported the email service Outlook was inaccessible.

 

Other services including Teams and Xbox Live were also reported as not working.

 

Microsoft said it was "monitoring the service" as the change takes effect.

Posted

"We’ve identified that a wide-area

networking (WAN) routing change caused impact to the service. We’ve rolled back

the change and monitoring the service as it recovers. Some of the customers who

had previously reported impact are also reporting recovery."

 

Sounds like someone is in trouble.

Posted (edited)

Generally that's not how things work in the big cloud providers. Mistakes happen. Systems and processes exist to reduce the chance of a mistake getting into production, but they are inevitable. What is learned is what needs to be done to prevent that same human/systems failure re-occurring. If a human pushed the wrong config to the wrong system, then a system needs to be put in place to prevent either that happening again, or happening at such scale so quickly again.

 

Individuals come and go, so there is no point in singling someone out because blaming a specific person prevents the *organization* from learning. Because what ever happens, the next outage will be be triggered by the actions of a different human.

Edited by psydii
  • Thanks 1
Posted
Yep you made a mistake if I get rid of you then that's our loss if I keep you this will not happen again. Also let's be fair here there is no one person involved in a reconfiguration of how the underlying routing works in the MS cloudy environment. This has passed a few people's eyes and probably some AI before they took down their entire operation!
Posted
Anyone having problems with Link to Windows today? My phone has just lost all my linked devices and won't connect to the one in front of me. Wondering if it's connected to these problems.
Posted
Anyone having problems with Link to Windows today? My phone has just lost all my linked devices and won't connect to the one in front of me.

 

Co-incidentally, my link (Android to Windows) broke a while ago and I re-made it last night.

Posted
Co-incidentally, my link (Android to Windows) broke a while ago and I re-made it last night.

 

 

Yeah, just had to do the same. Just my house and other two sites to go, I guess.

 

Hive (Heating) seems broken today too :(

Just checked. Mine seemed OK, then threw a wobbly and settled down. Doesn't look like the adjustment I made to the "work day" temp this morning as everyone is working "out" today took, though. The cat will have been happy, at least. Looking at the usage chart, it's not him wanting the door open to go out that's causing the high usage. Weeps.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...