Gongalong Posted October 5, 2012 Posted October 5, 2012 Hi folks, I'm getting frequent backup failures with DPM. We use it to backup shares, system states, and complete VMs. The DPM server is an HP ProLiant DL120 G6 running 2008 R2 (fully updated) and DPM 2012. It's connected via iSCSI to a Buffalo Terastation NAS. The daily failures are typically only the complete VMs (often the same VMs), and the errors are also typically the same (example below), although there is a little variation. Affected area: \Backup Using Child Partition Snapshot\VC-DC1 Occurred since: 05/10/2012 12:28:04 Description: Recovery point creation jobs for Microsoft Hyper-V \Backup Using Child Partition Snapshot\VC-DC1 on SCVMM VC-DC1 Resources.HyperV.domain.local have been failing. The number of failed recovery point creation jobs = 1. If the data source protected has some dependent data sources (like a SharePoint Farm), then click on the Error Details to view the list of dependent data sources for which recovery point creation failed. (ID 3114) DPM failed to access the volume \\?\GLOBALROOT\Device\HarddiskVolumeShadowCopy473\ on HyperV-2.domain.local. This could be due to 1) Cluster failover during backup or 2) Inadequate disk space on the volume. (ID 2040 Details: The device is not ready (0x80070015)) There aren't space issues on the SAN or NAS. Both the Hyper-V hosts (ProLiant DL385 G7's) have the latest Support Packs installed and have the latest Windows Updates. The DPM server has had its network drivers updated, but doesn't seem to have a Support Pack. If I keep manually retrying the backup jobs they eventually complete. Anyone got any ideas on where to continue troubleshooting this? TIA
Bruce123 Posted October 29, 2012 Posted October 29, 2012 Hi folks, I'm getting frequent backup failures with DPM. We use it to backup shares, system states, and complete VMs. The DPM server is an HP ProLiant DL120 G6 running 2008 R2 (fully updated) and DPM 2012. It's connected via iSCSI to a Buffalo Terastation NAS. The daily failures are typically only the complete VMs (often the same VMs), and the errors are also typically the same (example below), although there is a little variation. Affected area: \Backup Using Child Partition Snapshot\VC-DC1 Occurred since: 05/10/2012 12:28:04 Description: Recovery point creation jobs for Microsoft Hyper-V \Backup Using Child Partition Snapshot\VC-DC1 on SCVMM VC-DC1 Resources.HyperV.domain.local have been failing. The number of failed recovery point creation jobs = 1. If the data source protected has some dependent data sources (like a SharePoint Farm), then click on the Error Details to view the list of dependent data sources for which recovery point creation failed. (ID 3114) DPM failed to access the volume \\?\GLOBALROOT\Device\HarddiskVolumeShadowCopy473\ on HyperV-2.domain.local. This could be due to 1) Cluster failover during backup or 2) Inadequate disk space on the volume. (ID 2040 Details: The device is not ready (0x80070015)) There aren't space issues on the SAN or NAS. Both the Hyper-V hosts (ProLiant DL385 G7's) have the latest Support Packs installed and have the latest Windows Updates. The DPM server has had its network drivers updated, but doesn't seem to have a Support Pack. If I keep manually retrying the backup jobs they eventually complete. Anyone got any ideas on where to continue troubleshooting this? TIA DPM can be prone to this sort of thing. A few things. You say that your DPM server is connected via iSCSI to a Buffalo NAS, which I assume is used as the Storage Pool? Is this via a dedicated NIC on the server? Or does it share the same NIC as what is used for pulling the data from the protected servers? If you are going to use iSCSI to a NAS, I would strongly recommend you dedicate a NIC for it (and use the other NIC to connect to the network to pull backup data from protected servers). In order to work affectively, iSCSI really requires a dediciated 10Gbit connection (or 1Gbit minimum). I could be wrong, but the error message appears to show the DPM server is having trouble connecting to it's Storage Pool, which could be due to the NIC being overloaded by the traffic of data it is pulling from the protected servers. Alternatively, if the error is referring to problems connecting to the protected server, then the issue could be the flipside of this. That the iSCSI traffic (perhaps from a concurrent backup job), could be causing the DPM server to lose connectivity with the protected server. I notice that you are backing up System States, these (and BMR) can be very network intensive (as I don't believe they are Incremental) and could potentially interfer with other DPM concurrent backups. Looking at the logs, does it looks like the failed VM backups could coincide with BMR/System State Backups? What happens if you manually re-start a failed VM backup when you know that there are no other backups happening? Could you schedule the BMR/System State backups so they don't run concurrently the VM backups? Finally, for each DPM protected source DPM sets up two volumes in the Storage Pool (one for the main replica and one to store any changes), check that they are both big enough (right click on the protected source). Good luck, Bruce. 1
Bruce123 Posted October 29, 2012 Posted October 29, 2012 Also, I would check the Event logs on both, the DPM server and probably more importantly, the Hyper-V host of the VM that it failed on. Thanks, Bruce. 1
Gongalong Posted October 31, 2012 Author Posted October 31, 2012 You say that your DPM server is connected via iSCSI to a Buffalo NAS, which I assume is used as the Storage Pool? Yes. Is this via a dedicated NIC on the server? Dedicated. 1Gb/s. Looking at the logs, does it looks like the failed VM backups could coincide with BMR/System State Backups? Complete VM backups start at 6pm everyday. There are 12 of those, and it's typically 2 or 3 of these that fail. File shares are done twice a day at 2:30pm and 10:30pm. These have yet to fail, as far as I can recall. VM and DC system state recovery points are taken at 5am, and there's a full backup everyday at 8pm. I don't remember these failing either. What happens if you manually re-start a failed VM backup when you know that there are no other backups happening? Often it fails again, but if I persist it will eventually work. Could you schedule the BMR/System State backups so they don't run concurrently the VM backups? I think that's the case currently. This was one of the things I checked, and tried to get the various backups not to coincide with each other. Finally, for each DPM protected source DPM sets up two volumes in the Storage Pool (one for the main replica and one to store any changes), check that they are both big enough (right click on the protected source). Would DPM have done that already? The only other oddity is that if I check the Management section it shows some (but not all) protected servers. One of the servers it says is unprotected, but has the protection agent. Except this server is in the protection group. Coincidentally or not, this is the one that typically fails frequently, but will eventually backup if I persist with manual retries.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now