mikes Posted September 22, 2023 Posted September 22, 2023 Hi guys looking for some perspective on wether our SAN might be on its way out? We have a Lenovo DE-4000H; with 3 VMWare Vhosts. yesterday afternoon I was in my office when all of a sudden the server I was working on via the VMWare console blue screened, and the RDP ones half I lost contact with - but others seemed to be working fine. Initially I suspected a Vhost had gone nuts but usually they should fail over perfectly fine. I checked each vhost manually and they were all fine but then each server had a black screen with DATA_WRITE_ERROR on it. I checked our SAN and it has errors showing Host Redundancy Lost 2 times (for both esxi SAN controllers) - and it is showing a drive failure. When I checked our drive management it is showing 2 failed drives all of a sudden - 1 hot swap one and 1 normal one. But in the event log all I can see is 21 Sep 2023 14:02:44 CriticalController FirmwareNoneLoss of host-side connection redundancy detected 21 Sep 2023 14:02:44 CriticalController FirmwareNoneLoss of host-side connection redundancy detected 21 Sep 2023 13:50:46 CriticalControllerShelf 99, Bay AVolume not on preferred path due to AVT/RDAC failover and tons of these going back to when the problem started. (I guess it only keeps 1000 rows of logs, teach me to not set up syslog or event notification...) 21/09/2023 13:48 Informational Controller Shelf 99, Bay A VDD repair completed 49188 201F 0/0/0 Internal 21/09/2023 13:48 Informational Controller Shelf 99, Bay A VDD repair started 49187 201E 0/0/0 Internal 21/09/2023 13:48 Informational Controller Shelf 99, Bay A VDD repair completed 49186 201F 0/0/0 Internal The SAN is back up and running fine now(but with 2 duff drives), and I was able to manually restart each vhost so everything is up again, problem is we are wondering what to do now - the DE-4000H SAN is out of warranty - each replacement drive is £500 each and stock is very hard to find. We are wondering - extend the warranty at a high price and get the drives replaced; or buy a new one and go through the work of swapping over the data? It's just wierd how the redundancy was lost on both SAN controllers when I check the cabling all seems fine. It is a mess in the server room Maths seem fit to use it as a cupboard and I saw a spider crawling over one of the power sockets so I wonder if that might be half of it... We were having electrolocks fitted on site at the same time so figured maybe some weird power issue happened though I doubt it nothing else was affected? Just wierd how both controllers lost redundancy which caused all our vhosts to crap their pants, and maybe 2 drives didn't come back up? (last time I checked the SAN was a month ago so I guess 1 drive could have failed 2 weeks ago and 1 fail yesterday as they are getting old now)
psydii Posted September 22, 2023 Posted September 22, 2023 It’s your SAN. The very core of the IT service. Buy a new one or get a warranty /support contract (including software updates) on the existing one. We ran our old San for 10 years, but recently migrated to a newer model. Only had two problems in all that time, but 24/7 4hr break-fix saved our bacon both times. 1
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now