Jump to content

Recommended Posts

Posted

Good morning,

 

Came in this morning to find one of the hosts off, powered back on and everything looks fine but can't ping it\access interface. Doesn't look like anything has changed in terms of switch config.

 

Not great with VmWare but vSphere shows the host as 'Not Responding' so I'm wondering where I go from here?

 

Thanks

Posted

@Julian

 

Yes, I cannot reach the problematic host but can reach the other one.

 

I'm starting to think it's some sort of hardware failure as it has powered off again but I don't have access to the iDrac without resetting all of that. I'm going to see if I can run some diagnostics.

 

Thanks

Posted

@CrootUK

 

Yes, checked with my predecessor and apparently the drives are SD cards so the error above is expected.

 

I think I've definitely got a hardware issue somewhere because the host keeps shutting down and a couple of times when powering on it has hit a loop.

Posted

This is happened to me before running ESXi on a SD card. Was fine for a few years though! We didn't have a backup of the ESXi configuration either :) that was a fun afternoon

 

I think best practice is running on an internal SSD/SSD's (BOSS RAID 1 if using dell servers)

Posted

@Olliedawg

 

Thanks, my situation isn't getting any better. The log files that VmWare are asking for don't open after being downloaded to a zip folder and I reset the iDrac only to now not be able to reach it at all.

Posted
If its a Esxi host thats part of a cluster, you’ll spend more time trying to get it working on SD Cards than starting again, id just throw a disk in and install ESXI on it and get it going. We had this with R630s and the chassis couldn’t support disks in a Raid Config so we just winged it on a single disk until we replaced the entire lot.
Posted

@CrootUK

 

It's booting up fine now but I'm confused about why VCentre shows it as not responding and why I cannot ping it when there are seemingly no network config issues.

 

Thanks

Posted

For anyone still following this thread, the upshot was this:

 

The power cut\surge caused the NIC card used for connecting to SAN to become faulty, hence server boot issues\looping etc. AND separately it would appear that when the server and switch have come back on that LACP has not detected\negotiated the trunk properly.

 

So in essence the fix was to remove three of the four network cables from the server and put them back one at a time.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...