Jump to content

Recommended Posts

Posted
Hey Guys.

 

What mirroring / redundancy options are available on a pair of 7110 2tb SANs?

 

I originally bought a pair of these SAN's based on a project that was to create a new building with a second server room and have a SAN in server room building 1 and a san in server room building 2 and both of them mirroring - but the entire new building plan got shelved, and so I am left with both SAN's in the same rack of my only server room. *sigh*

 

Not ideal!!! But better than nothing.

 

This storage is only being used for VM's over NFS. (no CIFS or iSCSI). So what do you knowledgable guys thing is the best way for me to configure it so that my VMs are mirrored across both SANS so that if one san does go down I will be able to access the VM's off the second SAN straight away. I see in the interface there is "Replication" and also on the filesystem there is "Snapshots" but I don't know what does what. This is what I am aiming for. Can it be done?

 

I have not had time to do proper research into getting this configured on 7110, I have 6 big projects on at the moment as well as the day to day stuff, and Offstead in in January. As you can imagine I am struggling somewhat!!!! If anyone can help me my giving me the lowdown or idiots guide on this - I will be most greateful.

 

Cheers!

 

Butuz

 

 

I'll be interested to know how you get on with this. I'm hoping that we'll be able to do this with a currrent 2Tb version to a 4Tb version (which we'll buy soon). I'd like to have all our VM's on the current 2Tb box and mirror this to the 4Tb box, we would then use the extra 2Tb on the 4Tb unit for data. Basically the aim is to provide redundancy for the VM's worst case we can move the data to another server temprarily.

Posted (edited)

Excellent!!!!

 

Thanks for the quick replies!!

 

Sounds like replication is what I need. Yes I to am very interested in how replication will work on live VM's? Will I have to shut down all my VM's in order to perform a "replication"? How does it work on VM's?! I think I'll start off with a schedule of replication each night and see how that goes - if it works well I may try continuous replication!

 

To be honest I didn't expect automatic failover and will be happy to manually fail over to the other san if one should god forbid fail. Even with manual failovers - this is going to give us a level of data protection waaaaay beyond what we had previously as I will be combing the SAN replication with anormal disk to tape backup of all the VM's.

 

Just updating the firmware on both SANS!

 

Cheers

 

Butuz

Edited by Butuz
Posted
I plan to cluster two 7410's across the site next summer to do what you're talking about.

 

You are unable to cluster a pair 7310/7410 between two computer rooms. There is a number of limitations for S7000 clustering.

 

  • Clustron card requires serial connections
  • SAS cables don't extend well over 10meters.

Replication is the only way forward for multiple computer rooms with the S7000. As with most replication technologies and ESX, VM's have to recovery manually with scripts or use SRM.

 

Andy

  • Thanks 1
Posted

Just found a "The service processor needs to be reset to ensure proper functioning" sat in the "Problems" tab on our 7110.

 

Also discovered that the root password isn't being accepted in the bui or gui.

 

a) are the two related in any way? I suspect not, but y'know.

b) is there any way to reset the root password that doesn't involve swapping the jumper over or otherwise redoing the config?

 

I can still log in as root via the ILO, but I suspect that won't let me change the other root password, right?

Posted
You are unable to cluster a pair 7310/7410 between two computer rooms. There is a number of limitations for S7000 clustering.

 

Well... nuts. I assumed clustering would be bandwidth intensive but I figured a 10Gb feed would cover it. I hadn't looked into it that closely yet and didn't realise it was over SAS rather than Ethernet/Fibre.

 

Sooo... two clustered heads in each site with replication between the two sites. Four heads in total... time to start playing the lottery maybe? :rolleyes:

 

Pete - Presumably you had a different account on the 7110 to allow you to get to the problems page? I had that same error once, clicked the 'fix' button and it just resolved itself with no downtime. Not sure if it would affect the root password but when I had that problem I could still log into the BUI as root. You could probably reset the password from the low-level shell but it's unsupported so Sun would need to do it. Best option would probably be to submit a case - 0870 600 3222.

 

Cheers,

Chris

Posted
Sooo... two clustered heads in each site with replication between the two sites. Four heads in total... time to start playing the lottery maybe? :rolleyes:

 

Pete - Presumably you had a different account on the 7110 to allow you to get to the problems page? I had that same error once, clicked the 'fix' button and it just resolved itself with no downtime. Not sure if it would affect the root password but when I had that problem I could still log into the BUI as root. You could probably reset the password from the low-level shell but it's unsupported so Sun would need to do it. Best option would probably be to submit a case - 0870 600 3222.

 

Yeah, the account that can see the problem doesn't have the rights to fix it. :(

 

I'm pretty damn sure (95%) that I'm entering the right root password (X or a slight varitation on X) and I've tried it with different browsers and workstations, just in case. ILO has a separate root account to the main root account as far as I can see.

Posted
Yeah, the account that can see the problem doesn't have the rights to fix it. :(

 

I'm pretty damn sure (95%) that I'm entering the right root password (X or a slight varitation on X) and I've tried it with different browsers and workstations, just in case. ILO has a separate root account to the main root account as far as I can see.

 

Have you tried a console connection login.

Login to iLOM as root

sp> start /SP/console

Andy

Posted
Have you tried a console connection login.

Login to iLOM as root

sp> start /SP/console

Andy

 

IIRC (was last night when I checked), I get prompted for a root password that isn't the same as the ILO one.

Posted
Well... nuts. I assumed clustering would be bandwidth intensive but I figured a 10Gb feed would cover it. I hadn't looked into it that closely yet and didn't realise it was over SAS rather than Ethernet/Fibre.

 

Sooo... two clustered heads in each site with replication between the two sites. Four heads in total... time to start playing the lottery maybe? :rolleyes:

 

I suspected that clustering could only happen over very short distances but apaton's confirmed it for me now :(. I supose you could always get away with "only" 3 7410/7310 heads...or 2 7410 heads/1 7310. Have the main selection clustered and with enough JBOD's to give you no single point of failure, replication to a 7410/7310 in the backup server room.

 

Still talking lots of £££ though :(

Posted
I suspected that clustering could only happen over very short distances but apaton's confirmed it for me now :(. I supose you could always get away with "only" 3 7410/7310 heads...or 2 7410 heads/1 7310. Have the main selection clustered and with enough JBOD's to give you no single point of failure, replication to a 7410/7310 in the backup server room.

 

Still talking lots of £££ though :(

 

Indeed, I mean it's still a lot better than having one Windows server holding everything, but proper real-time clustering between two heads and sets of JBODs in different areas of the site would be ideal. Doesn't look like I can do it on the cheap though!

 

I go back to my question on the Cutter forums then: How do the big players handle this? Google use GFS and are different to most, but what about Amazon, Microsoft, BBC, etc? They must be using SANs in some form, and probably a lot of them. One's got to fail eventually, how do they handle that with no downtime or dataloss?

 

Chris

Posted
Indeed, I mean it's still a lot better than having one Windows server holding everything, but proper real-time clustering between two heads and sets of JBODs in different areas of the site would be ideal. Doesn't look like I can do it on the cheap though!

 

I go back to my question on the Cutter forums then: How do the big players handle this? Google use GFS and are different to most, but what about Amazon, Microsoft, BBC, etc? They must be using SANs in some form, and probably a lot of them. One's got to fail eventually, how do they handle that with no downtime or dataloss?

 

Chris

 

Gonna call

Posted
I'd be interested to hear how well VMs handle a fallover in a cluster to a different head unit as well. Looking at the 7310/7410 with clustering in the next 12-24 months and being able to survive a head unit dying is one of the questions I had :)
Posted

In theory that should be fine. Clustering in the most general sense suggests the two heads will handle any failures with no downtime as they're kept in sync, so you shouldn't notice any issues with half of a cluster going down as that would defeat the purpose.

 

I had a chat with Andy and I've sent a couple of questions over to Phil at Sun regarding clustering and replication, I'll let you know what he says when I get a reply. ;)

 

Chris

  • Thanks 1
Posted
In theory that should be fine. Clustering in the most general sense suggests the two heads will handle any failures with no downtime as they're kept in sync, so you shouldn't notice any issues with half of a cluster going down as that would defeat the purpose.

 

I had a chat with Andy and I've sent a couple of questions over to Phil at Sun regarding clustering and replication, I'll let you know what he says when I get a reply. ;)

 

Chris

 

That's what I suspected, but if someones actually run some tests on it that would be great to know :D

Posted (edited)
I go back to my question on the Cutter forums then: How do the big players handle this? Google use GFS and are different to most, but what about Amazon, Microsoft, BBC, etc? They must be using SANs in some form, and probably a lot of them. One's got to fail eventually, how do they handle that with no downtime or dataloss?

 

Not a simple problem to solve, even with two well connected computer rooms. Issues with split-brain services never mind storage availability and the need for quorum devices.

 

The best answer is typically mirroring between multiple storage units. VMWARE ESX doesn't have mirroring which is most frustrating. So relies on dual controllers in the storage or network/SAN based mirroring. (FalconStor's IPstor, IBM's SVC)

 

OS's like Windows, Solaris, AIX and Linux have mirroring between devices (VxVM,SVM,LVM,mdtool), so if one mirrors fails the other is instantly available. This give you many more options in the design of storage, including mirroring between sites over dark fibre.

 

Another question to ask, "VMWARE will cope quite happily with failovers, but does the VM/Application ?". Test Test Testing........

 

Andy

Edited by apaton
  • Thanks 1
Posted
I'd be interested to hear how well VMs handle a fallover in a cluster to a different head unit as well. Looking at the 7310/7410 with clustering in the next 12-24 months and being able to survive a head unit dying is one of the questions I had :)

 

I think we need to find someone with a couple of S7000's they can set up and test with VMs for both clustering and replication. It doesn't look like anyone has done any extensive testing out there yet and Phil at Sun was unable to 100% confirm my question able replication.

 

Is there anyone out there with spare hardware that can test this for us?

 

Clustering

Two clustered heads with JBOD storage

VMs running from the storage

Kill one of the heads, is there any downtime or instability in the running VMs?

Main interest is with NFS and ESX?

 

Replication

Two separate heads and JBODs

VMs running from one box replicated to the other

Synchronous and scheduled replication configured between the two SANs

If primary storage unit is killed and replication target brought online, what state are the VMs in?

Main interest is with NFS and ESX again?

 

Cheers,

Chris

Posted
I can test replication, that is the setup I am looking at doing, so its on my list but may take me a couple of weeks as I'm a bit snowed under!
  • Thanks 1
Posted
I can test replication, that is the setup I am looking at doing, so its on my list but may take me a couple of weeks as I'm a bit snowed under!

 

No rush, but I that'd be really appreciated if you could! :)

 

I know the S7000 stuff is VMware certified, if anyone who has any contacts at Sun is reading this, could clustering/replication be tested and confirmed so we all know where we stand and what the best practices are?

 

Chris

Posted

Reply from one of the Sun guys regarding running VMs on a clustered head:

 

In the case of the 7310 failing, yes, the NFS service will resume on the surviving node after the appropriate takeover delay (depends on # of disk sets, shares, services configured on the 7000). The VM will experience a pause as the service is transferred over. Most VMs can be configured to handle a 60-120 second delay without too much issue (say 2 JBODs on any one 7310 head in an active/active config to be comfortable with that 60-120 sec takeover), however, applications running *inside* the VM may or may not be able to handle that sort of pause, it would depend on the apps.

 

I'm not sure this would provide 'faultless' service? As mentioned, it would depend on the specific setup so this probably needs to be tested further I guess.

 

Cheers,

Chris

Posted
Small note to other users, if one is having a blonde moment (or senior moment however you prefer to term it), and assign your SANs static IP to another device by accident, said S7110 has a BIG sulk, and responds to ping's and web-interface access but refuses to serve files until its had a hard reboot and then decides it wishes to serve its gracious masters with files :o
  • Thanks 1
Posted

I apologize for being late to the game here.

 

I'm interested in a 7110 for our company SAN. I'd like to use it for iSCSI storage for my VMWare VM's as well as CIFS shares. To top it off, I'd like to be able to back up the whole SAN to something like a Quantum Superloader (Something like D2D2T).

 

Is this possible with the 7110? What other additional hardware considerations would I need to visit? We have about 1TB of data right now, so I'd like room to grow. I'd imagine we have a backup window of something like 8-9 hours to run the tape and with the "small" amount of data we are storing now I can get that done in that window, right?

 

If I wanted to back up the SAN to tape, would I need an external SAS/SCSI card for the 7110 to hook it to the SuperLoader?

Posted
I apologize for being late to the game here.

 

I'm interested in a 7110 for our company SAN. I'd like to use it for iSCSI storage for my VMWare VM's as well as CIFS shares. To top it off, I'd like to be able to back up the whole SAN to something like a Quantum Superloader (Something like D2D2T).

 

Is this possible with the 7110? What other additional hardware considerations would I need to visit? We have about 1TB of data right now, so I'd like room to grow. I'd imagine we have a backup window of something like 8-9 hours to run the tape and with the "small" amount of data we are storing now I can get that done in that window, right?

 

If I wanted to back up the SAN to tape, would I need an external SAS/SCSI card for the 7110 to hook it to the SuperLoader?

 

I would not use iSCSI for the VMs, you will get much better performance over NFS.

 

How many VM servers and how many uses is the next question.

 

Back up cards available are

 

* Dual Channel 4Gb FC HBA

* Dual Channel Ultra320 SCSI HBA

 

You would need something NDMP v3/v4 capable to do the backups

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...