Jump to content

Recommended Posts

Posted

I’m wondering if anyone can help us with the issue below.

 

We are currently running around 12 HA VMs, on a 2-node Windows Server 2012 R2 Hyper-V cluster. VM storage is housed on SMB 3.0 shares, which are running from a 2-node Windows Server 2012 Scale-out File Server cluster.

 

2 SMB 3.0 shares have been provisioned from the 2 CSVs presented by the SOFS cluster. Each SOFS cluster node is an owner of a CSV. Storage hardware is a Lenovo ThinkServer JBOD.

 

For VM backup, we are utilising Veeam Backup and Replication 8, with update 2b applied. The backup run begins at 10pm each evening, by means of a scheduled PowerShell script.

The backup completes successfully, but have found over the past week or so that upon arriving at the office the next day, there are multiple 1069 cluster events for the VMs which have rebooted at random.

 

The VMs in question, are random in terms of which ones reboot each evening.

 

In an effort to find the root cause of the problem, we disabled all Veeam VM backups one evening. The following morning, our Hyper-V cluster reported that it had gone the entire period without having any issues.

 

We then, manually ran the backup script during office hours, and waited for any issues. The backup ran without issue, until it came to one of the last few VMs.

What then happened, was that 3 of the VMs restarted. Specifically, EMAIL, PRINTSRV & TS4. The VMs rebooted during the TS4 backup.

 

These restarted between 12.47pm and 12.48pm. All 3 came back up. There doesn’t appear to be any link between the three (apart from the fact that all 3 were running on the same HV node). What’s more odd is that the reboots occurred way after 2 of the VMs.

I should add that there were other VMs running on that same HV node.

 

Backup completion times:

EMAIL – 10.49am

PRINTSRV – 12.08am

TS4 – 12.55am

 

The backup then proceeded, until reaching the penultimate VM. Suddenly, I then noticed that all VMs on all HV nodes lost connection to their storage, and were either turning off or starting up on another node of the HV cluster. A few seconds after seeing this, I checked the logs for our SOFS cluster, and noticed that RHS had stopped unexpectedly, which caused the file cluster to restart and VMs to bomb out.

 

Amazingly the Veeam backup proceeded to backup the last VM when both clusters returned to a normal running state.

 

Does anyone have any ideas what is causing this problem? I keep reading about disabling ODX in Server 2012, for storage hardware that doesn’t support it.

 

All I know is that running the backup, causes problems.

 

Many thanks.

Posted

Hi Steve21,

 

Thanks for the reply.

 

The thing is, we don't actually have any events in the logs that pertain to losing connection to the CSVs themselves. I have found multiple articles online which address Windows Updates. Problem is, I'd like to patch for the problems that can be seen in the event logs, as opposed to just patching with anything linked to clusters/CSVs etc, and potentially causing myself a bigger headache.

 

What issues did you have with your cluster? Was it the same setup, VMs over SMB 3.0 running from an SOFS cluster?

Posted

HyperV cluster but using SAN, multiple 1069 5120/1 5142 errors etc

 

Had some issues in regards to CSVs going offline thus pausing/restarting all machines etc.

 

It's quite a lot of issues these rollups fix though: https://support.microsoft.com/en-us/kb/2813630 another example in relation to your ODX:

 

To avoid CSV failovers, you may have to make additional changes to the computer after you install the hotfix. For example, you may be experiencing the issue described in the "Symptoms" section because of the lack of hardware support for Offloaded Data Transfer (ODX). This causes delays when the operating system queries for hardware support during I/O requests.

 

In this situation, disable ODX by changing the FilterSupportedFeaturesMode value for the storage device that does not support ODX to 1

 

You can always try them individually if you prefer, but most the individual patches were removed as they were put into the rollups

 

Steve

Posted

Hi Steve,

 

We've just found a useful Veeam post, detailing a list of recommended hotfixes/patches for 2012 clusters.

Going back to the one you originally mentioned, one of the final lines is, "After you install this hotfix on a Hyper-V server, you must update the integration components..."

It led me to ask the question, are we firing off backups from the wrong server? We are initiating the backup from one of the SOFS nodes, as opposed to from one of the HV nodes. This made more sense. Can you see any issue with doing it this way?

Posted

Well I don't really get why you're using a script etc anyway :p But surely you're running the backups on the HyperV nodes (as if you consider the SAN scenario like mine you can't run services on the SAN).

 

But in terms of where is your Veeam server installed? As we have our Veeam server in DR, that fires off the commands as such, and then the HyperV Hosts doing the backup on themselves.

 

Steve

Posted

The script is utilised, as until update 2b was released for B&R, no scheduling functions were included with it out-of-the-box. Update 2b provided support for PowerShell, which could then be utlised to perform scheduled backups.

Veeam Backup Free Edition: Now with PowerShell!

 

Veeam B&R is running on one of our scale-out nodes. We can do this as the SOFS nodes aren't connected to a SAN, but to a JBOD.

Posted (edited)

I may be missing the point entirely, but you realise that's the free edition you linked?

 

Not B&R for Virtual which has full scheduling etc on it :confused: or aren't you using the paid for version?

 

Steve

Edited by Steve21
Posted
It is the free version of Backup and Replication.

 

Ah sorry ignore me then haha :D Makes more sense.

 

But aye certainly would suggest the updates/rollups even if it is just the ones Veeam list as there's so many silly ones that aren't got by default on Windows updates.

 

Steve

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...