Jump to content

Recommended Posts

Posted

So yesterday was a bad day!

Got in to find the Primary DC down and further more corrupt (.vmdk.), snapshot consolidation needed blah blah..

 

Consolidating the snapshots manually didn't work and had to restore from backup, eventually got all working.

 

More of a worry is that over the last weeks\months we have had probably 5/6 of our VM`s become corrupt.

 

Our setup :

 

3 Hosts running 15 VM`s all using local storage

Backups using Veeam and then replicating to a DR vCentre

 

The corrupt VM`s all happen over night in and around when the backups occur, it doesn't happen every night, pretty much a lottery on which on and when!

 

The only thing I can put it down to is two of the hosts are low on disk space (Datastore) 200/300gb remaining, but this also occurs on the host which has plenty of DS remaining (500gb)

 

Any advice would be beneficial.

Posted
are they dell servers with a perc raid card by any chance? if so theres an option on setting up the raid array that can cause issues (it literally wiped servers for me) unfortunately i cant remember what the option is off the top of my head as on newer ones its off by default thank f
  • Thanks 1
Posted
Out of interest why would you have snapshots? If you must use snapshots it should only be for the duration of the fix or update and you should immediately merge the data after. Snapshots causes all sorts of issues the main one being system slow down etc. Not sure if this could cause the corruption issue but I have been heavily warned against snapshots on live systems in the past.
  • Thanks 1
Posted

What does Veeam say about the VM in question when it happens?

 

I am suspicious of whether this is happening snapshot removal on your local storage through too many things happening at once. (we sometimes see it when a machine drops into wanting consolidation it's because the snapshot removal timed out or such).

 

You could try slowing Veeam down on the number of concurrent tasks, or possibly looking at proxy configuration.

 

Your post says you backup and then replicate, do you know which of the tasks it happens on? (I presume you're backing up and then replicating from the backup, or are you hitting the prod server again?)

 

What is your Veeam configuration in terms of proxies, etc?

  • Thanks 1
Posted
Out of interest why would you have snapshots? If you must use snapshots it should only be for the duration of the fix or update and you should immediately merge the data after. Snapshots causes all sorts of issues the main one being system slow down etc. Not sure if this could cause the corruption issue but I have been heavily warned against snapshots on live systems in the past.

I think Veeam creates a snapshot whilst doing the backup process, AFAIK

Posted
What does Veeam say about the VM in question when it happens?

 

I am suspicious of whether this is happening snapshot removal on your local storage through too many things happening at once. (we sometimes see it when a machine drops into wanting consolidation it's because the snapshot removal timed out or such).

 

You could try slowing Veeam down on the number of concurrent tasks, or possibly looking at proxy configuration.

 

Your post says you backup and then replicate, do you know which of the tasks it happens on? (I presume you're backing up and then replicating from the backup, or are you hitting the prod server again?)

 

What is your Veeam configuration in terms of proxies, etc?

 

Veeam config - backup console and proxy all on same server.

 

The Veeam error says "VM disk consolidation failure" - The failed VM yesterday seemed to happen within seconds of the scheduled replication Job starting.

The backup of this VM was a 5pm and ran successfully

 

I try to set them all up so theres no concurrent jobs running, maybe 2 at most but usually one starts after the other is complete.

Posted
I had this before veeam had locked the disk while taking a backup i think if you storage migrate the disk and vm off then back on i think that's how i freed the disk up. we a re talking 2 years ago tho
  • Thanks 1
Posted

The word Dell always makes me cringe! bad experiences put me off for life! I had more issues with an inherited dell (after a merge) than i have in my HPs of which some are still running 7 years on without a part failure!

 

Despite the old G5 servers only having support up to 2008 R2, i've got 1 happily running on server 2016

  • Thanks 1
Posted

So it appears it's not the backups that are the problem, it is the replications?

 

Have you contacted veeam support about this? They will be able to help you if you supply log files etc.

  • Thanks 1
Posted
So it appears it's not the backups that are the problem, it is the replications?

 

Have you contacted veeam support about this? They will be able to help you if you supply log files etc.

 

Yes upon further investigation It was a replication job that caused the issue, I have disabled the replication jobs whilst I look into it

Posted
We are thinking about moving to veeam, it seem scary.... I hope that you are able to sort this out before we are ready to move towards veeam....... Good luck
  • Thanks 1
Posted

Sounds interesting, the removal of the snapshot should be virtually the same on the backup as on the replication - it sounds like for whatever reason there's a difference in what's going on on the ESXi/Storage side when the removal happens and fails causing the consolidation need. For it to then fail with the snapshot hunter to consolidate would reveal some very interesting logs on the ESXi side. I wonder if it's some kind of storage reservation issue perhaps.

 

You say you have multiple hosts, I wonder if its as simple as a lock file issue on the datastore if two hosts simultaneously call for a removal of snapshot and just by luck it's not happened on a backup yet.

 

We could test this theory out quite easily with an orchestrated replication job from a single source datastore.

 

Nonetheless, you should take a vm-support from VMware's side and a log bundle from Veeam and send to both vendors to see whether they can see what's happening.

 

I'd be betting on an ESXi/storage issue myself, the actions being carried out by Veeam are all tasks issued to ESXi/vCenter but if Veeam can reliably cause it in your environment it might help provide insight on what is happening.

  • Thanks 1
Posted
Sounds interesting, the removal of the snapshot should be virtually the same on the backup as on the replication - it sounds like for whatever reason there's a difference in what's going on on the ESXi/Storage side when the removal happens and fails causing the consolidation need. For it to then fail with the snapshot hunter to consolidate would reveal some very interesting logs on the ESXi side. I wonder if it's some kind of storage reservation issue perhaps.

 

You say you have multiple hosts, I wonder if its as simple as a lock file issue on the datastore if two hosts simultaneously call for a removal of snapshot and just by luck it's not happened on a backup yet.

 

We could test this theory out quite easily with an orchestrated replication job from a single source datastore.

 

Nonetheless, you should take a vm-support from VMware's side and a log bundle from Veeam and send to both vendors to see whether they can see what's happening.

 

I'd be betting on an ESXi/storage issue myself, the actions being carried out by Veeam are all tasks issued to ESXi/vCenter but if Veeam can reliably cause it in your environment it might help provide insight on what is happening.

I`ve disabled the replication jobs for the time being (friday) and will run with this this week to see if its that, (None over the weekend, so far so good!) I can then be sure its the replication that`s causing the error and then take this up with VMware\Veeam.

Posted
Dell servers (or possibly others with rebranded LSI/Avago controllers) may need a firmware update to remove the T10 protection option from the RAID controller (assuming it's a PERC 730). I lost a whole array to this, just after moving a school to a virtualised setup. Handily it was backed up to tape so there was minimal downtime after the issue was identified and fixed.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...