Friday, 25 May 2007

Mine's Bigger than yours - Part Deux

Old Barry's at it again with the DMX pitch, this time for VTL.

Here's a quick comparison with Copan's Revolution 220TX system and the specification of the DL6100 from EMC's website.

Copan Revolution 220TX

  • Single Cabinet
  • 896 drives (500GB SATA)
  • Power (max): 6368 watts
  • Throughput: 5.2TB/hour
  • Capacity: 448TB max
  • Emulation: 56 libraries, 56 drives, 8192 virtual cartridges

EMC DL6000

  • 7 cabinets (max configuration)
  • 1440 drives (500GB LCFC)
  • Power (max): 49100 watts
  • Throughput: 6.4TB/hour
  • Capacity: 615TB max
  • Emulation: 256 libraries, 2048 drives, 128000 virtual cartridges

So, you'd need 1.5 Copan devices to match the EMC kit and yes, it doesn't scale as well in terms of virtual components, but there are some big issues here. For instance, EMC's device isn't green - the power demands are huge and not surprising, as the drives are all spinning all the time (Copan have only 25% max if theirs in use at any time). Floor density is not good in the DMX - why? Because the drives are in use all the time and it is an Enterprise array, so timely replacement of disk failures is important - but less so for a virtual tape system which can tolerate downtime.

So what would you go for? My money would be on a DMX for Enterprise work, which it is great at - and not as a "one size fits all" system, which it quite plainly is not.

Tuesday, 22 May 2007

Trusting TrueCopy

For those people who use TrueCopy on a daily basis, you'll know that the assignment of a TrueCopy pair is based on a source and target storage port, host storage domain and LUN. This means a LUN has to be assigned out to replicate it.

The part that has always worried me is the fact that the target LUN does not need to be in a read-only status and can be read-write.

Please, HDS to save my sweaty palms, change the requirement to make the target volume read-only before it can be a TrueCopy target....

Thin Provisioning - Cisco Style

There have been so many discussions on thin provisioning since the Hitachi USP-V announcement. When a major player takes on a specific technology, all of a sudden we realise that everyone else has already been doing it. Tony Asaro's post probably provides the best summary of those who do it today.

One vendor not appearing on the list is Cisco. Not surprising as they don't produce storage systems, however they do produce fibre channel switches which also implement thin provisioning.

I talked about the issue not that long ago here. Now I'm laying out ports for real and its not as clear as it seems. The low down is if you put ports in "dedicated" mode, you get the port speed reserved (or 4Gbps if you set to auto, regardless of the speed negotiated) and in "shared" mode you have a minimum requirement, 4Gbps ports need 0.8Gbps reserved for instance, and the figure reduces in proportion for 2 and 1Gbps ports. More details can be found here. This means not all combinations are possible and you get "Bandwidth Not Available" messages when you don't expect it. As this was confusing me, I've put together a port speed calculator, you can pick it up at Cisco Rate Calculator.

Thursday, 17 May 2007

New Product Announcement,,,

There's another new product out this week from HP - the XP24000... It seems to have some similarities to another product launch, 224 ports, thin provisioning, partitioning... :-)

Wednesday, 16 May 2007

Poll Results

Thanks to everyone who voted in the poll on Optimiser/Cruise Control; seems like most people like the recommendations but don't trust it to make automatic changes.

I've got a new poll running for the next month - tell me what you think about USP-V. Now I'm *sure* there will be lots of opinion on this....

Port Indexes on Cisco Switches

I spent today building some Cisco MDS9513 switches. It's good to get into the nuts and bolts of technology once in a while. They are big beasts - 2nd generation technology and with 11 usable slots, the biggest available line cards (48 ports), 528 ports in a single chassis. However the build had one fly in the ointment. As part of the (inherited) design, the chassis included a "generation 1" eight ethernet port IP line card for implementing FCIP.

The issue revolves around a feature called port indexes, which are used to track the ports installed in a MDS chassis. Generation 1 chassis have a maximum of 252 port indexes, generation 2 technology supports up to 1020 port indexes. However, a generation 1 line card inserted into a generation 2 chassis dumbs down the whole switch to generation 1 port indexes. So, with 252 indexes, 32 taken up by the FCIP card, only 240 port indexes are left which directly translates to 240 ports - less than half the switch capacity! If line cards are installed that take the port count above 252 then these additional cards won't come online - they will initially power up then power down.

In this instance the solution will be to move IP services to another (smaller) switch and hopefully Cisco will bring out a generation 2 version of the IP blade soon. There is a bigger problem though, and that is for any customers looking to take 9513s and use them with the SSM module, which supports products such as EMC Invista and Kashya (now EMC RecoverPoint). As far as I am aware, there's no plans for a generation 2 SSM module any time soon, so using the SSM module in the 9513 chassis will create the same port restriction issues.

I don't see why Cisco couldn't simply produce a gen 2 version of the old gen 1 blades which did nothing more than re-jig the port indexing. Come to think of it, surely they could patch the firmware on the gen1 line-cards to fix the problem. Obviously it is not that simple, or perhaps not that many people are using Invista and RecoverPoint to make it worthwhile.

USP-V...and another thing

I hate it when I write a post then think of other things afterwards. One more USP-V thought (I'm sure there will be more). One of the drawbacks of the current hardware with virtualisation is the effort to remove/upgrade the USP. Was there anything in the announcements on Monday to cater for this? I was hoping HDS would announce USP "clustering". Although they don't think it is necessary from a resiliency perspective, it certainly is if you want to upgrade and haven't done a 1 for 1 passthrough on LUNs (i.e. presented "big" LUNs to the array and then carved them up in the USP).

So HDS, did you do it?

Oh, and another one...for customers who've just purchased a standard USP, will there be a field upgrade to USP-V?

Tuesday, 15 May 2007

USP-V - bit of a let down?

It seems from the posts seen so far on the blogosphere that the USP release is causing a bit of a stir (10 points for stating the obvious I think). So, here’s my take on the announcements so far.

First of all, it’s called USP-V – presumably because of the “Massive 500-Percent Increase in Virtualized Storage Port Performance for External Storage”. I'm not sure what that means - possibly more on that later.

As previously pointed out, the USP-V doesn’t increase the number of disks it supports. It stays at 1152 and disappointingly the largest drive size is still 300GB and only 146GB for 15K drives. I assume HDS intends to suggest that customers should be using virtualisation to connect to lower cost, higher capacity storage. That’s a laudible suggestion and only works if the Universal Volume Manager licence is attractive to make it work. In my experience this is an expensive feature and unless you’re virtualising a shed-load of storage, then it probably isn’t cost effective.

There have been some capacity increases; the number of logical LUNs increases 4-fold. I think this has been needed for some time, especially if using virtualisation. 332TB with 16384 virtual LUNs meant an average of 20GB per LUN, obviously now it is only 4GB. Incidentally, the HDS website originally showed the wrong internal capacity here: http://www.hds.com/products/storage-systems/capacity.html, showing the USP-V figures the same as the base USP100. It’s now been corrected!

Front-end ports and back-end directors have been increased. For fibre-channel the increase is from 192 to 224 ports (presumably 12 to 14 boards) and back-end directors increase from a maximum of 4 to 8. I’m not sure why this is if the number of supportable drives hasn’t been increased (do HDS think 4 was insufficient or will be see a USP-V MKII with support for more drives?). Although these are theoretical maxima, the figures need to be taken with a pinch of salt. For example, the front-end ports are 16-port cards in the USP and there are 6 slots. This provides 96 ports, the next 96 are provided by stealing back-end directors (this is similar to DMX-3 – 64 ports maximum which can be increased to 80 by removing disk director cards). Surprisingly, throughput hasn’t been increased. Control bandwidth has, but not cache bandwidth. Does the control bandwidth increase provide the 500% increase in virtualisation throughput for external storage?

What about the “good stuff”? So, far, all I can see is Dynamic (thin) Provisioning and some enhancements to virtual partitions. The thin provisioning claims to create virtual LUNs and spread data across a wide number of array groups. I suspect this is simply an extension of the existing copy-on-write technology, which if it is, makes it hardly revolutionary.

I’d say the USP-V is an incremental change to the existing USP and not quite worthy of the fanfare and secrecy of the last 6 months. I’d like to see some more technical detail (HDS feel free to forward me whatever you like to contradict my opinion).

One other thought. I don’t think the DMX-3 was any less or more radical when it was released....

Monday, 14 May 2007

100 Up


I've reached 100! No, not 100 years old, but my 100th post. I've surprised myself that I've managed to keep going this long.


So a bit of humour - I'll save the USP-V discussions until later - who remembers Storage Navigator for the 9900 series?


Well, the left-hand toolbar has various icons on it. One of them I've reproduced here;

it looks like a fried egg and a spanner. But what is it meant to represent (btw, it's the icon for LUSE/VLL).

Answers on a postcard please....

Friday, 11 May 2007

Mine's bigger than yours - do we care?

Our resident storage anarchist has been vigorously defending DMX - here. It's all in response to discussions with Hu regarding whether USP is better than DMX. Or, should I say DMX-3 and Broadway (whoops, I mentioned the unmentionable, you'll have to shoot me now).

I have to say I enjoyed the technical detail of the exchanges and I hope there will be a lot more to come. Any insight into how to make what are very expensive storage subsystems work more effectively has to be a good thing.

But here's the rub. Do we care about how much faster DMX-3 is over USP? I doubt the differences are more than incremental and as I've both installed, configured and provisioned storage on 9980V/USP/8730/8830/DMX/DMX2/DMX3, I think I've enough practical experience to qualify it. (By the way, I loved StorArch's comment about how flexible BIN file changes are now. Well, they may be, but in reality I've found EMC cumbersome to release configuration changes).

Finally I'll get round to the point of this post; most large enterprise subsystems are of the same order of magnitude of performance. However I've yet to see any deployment where performance management is executed to such a degree that, hand on heart, the storage admins there can claim they sequeeze 100% efficient throughput. I'd make an estimate that things probably run 80% efficient, with the major bottlenecks being port layout, backend array layout and host configuration.

So the theoretical bantering on who is more performant than the other is moot; now, EMC, HDS or IBM, come up with a *self tuning* array then you've got a winner...

Wednesday, 9 May 2007

Simulator Update

I managed to get a copy of the Celerra Simulator last week and I've just managed to get it installed. Although it is simple, it is quite specific on requirements - it runs under VMware ACE as a Linux program. It needs an Intel processor and can't run on a machine with VMware already installed. Fortunately my test server fits the bill (once I uninstalled VMware Server). Once up, you administer through a browser.

At this stage that's as far as I've got - however it looks good. More soon.

Port Oversubscription

Following on from snig’s post, I promised a blog on FC switch oversubscription. It’s been on my list for some time and I have discussed it before, however it has also a subject I’ve discussed with clients from a financial perspective and here’s why; most people look at the cost of a fibre channel switch on a per port basis, regardless of the underlying feature/functionality of that port.

Not that long ago, switches from companies such as McDATA (remember them? :-) ) provided full non-blocking architecture. That is, they allowed full 2Gb/s for any and all ports, point to point. As we moved to 4Gb/s, it was clear that Cisco, McDATA and Brocade couldn’t manage (or didn’t want) to deliver full port speed as the port density of blades increased. I suspect there were be issues with ASIC cost and fitting the hardware onto blades and cooling it (although Brocade have just about managed it).

For example, on Generation 2 Cisco 9513 switches, the bandwidth per port module is an aggregate 48Gb/s. This is regardless of the port count (12, 24 or 48), so, although a 48 port blade can (theoretically) have all ports set to 4Gb/s, the ports on average only have 1Gb/s of bandwidth.

However the configuration is more complex; ports are grouped into port groups, 4 groups per blade of 12Gb/s each, putting even more restriction on the ability to use available bandwidth across all ports. Ports can be dedicated or use shared bandwidth within a port group. In a port group of 12 ports, set three to dedicated bandwidth of 4Gb/s and the rest are (literally) unusable. Whilst I worked at a recent client, we challenged this option with Cisco. As a consequence, 3.1(1) of SAN-OS allows the disabling of all restrictions so you can take the risk and set all ports to 4Gb/s and then you’re on your own.

How much should you pay for these ports? What are they actually worth? Should a 48-port line card port cost the same as a 24-port line card port? Or, should they be rated on bandwidth? Some customers choose to use 24-port line cards for storage connections or even the 4-port 10Gb/s cards for ISLs. I think they are pointless. Cisco 9513’s are big beasts; they eat a lot of power and need a lot of cooling. Why wouldn’t you want to cram as many ports into a chassis as possible?

The answer is to look at a new model for port allocation. Move away from the concepts of core-edge and edge-core-edge and mix storage ports and host ports on the same switch and where possible within the same port group. This would minimise the impact of moving off-blade, off-switch or even out of port group.

How much should you pay for these ports? I’d prefer to work out a price per Gb/s. From the prices I’ve seen, that makes Brocade way cheaper than Cisco.

Tuesday, 1 May 2007

Tuning Manager CLI

I've been working with the HiCommand Tuning Manager CLI over the last few days in order to get more performance information on 9900 arrays. Tuning Manager (5.1 in my case) just doesn't let me present data in a format I find useful, and I suppose that's not really surprising as, unless you're going to add a complete reporting engine into the product, then you'll be wanting to get the data out of the HTnM database and build your own bespoke reports.

So I had high hopes for the HTnM CLI, but I was unfortunately disappointed. Yes, I can drag out port, LDEV, subsystem (cache etc) and array group details, however I can only extract one time period of records at a time. I can display all the LDEVs for a specific hour, or a day (if I've been aggregating the data) but I can't specify a date or time range. This means I've had to script extracting and merging the data and the result is, it is sloooow. Really slow. One other really annoying feature - fields that report byte throughput sometimes report as "4.2 KB" sometimes as "5 MB" - which programmer thought a comma delimited output would want a unit suffix?

I'm expecting delivery of HTnM 5.5 (I think 5.5.3 to be specific) this week and here's what I'm hoping to find; (a) the ability to report over date/time range (b) the database schema to be exposed for me to extract data directly. I'm not asking much - nothing much more than other products offer. Oh, and hopefully something considerably faster than now.

Wednesday, 25 April 2007

What's your favorite fruit? EMC versus HDS

Nigel has posted the age old question, which is best EMC or HDS? For those who watch Harry Hill - there's only one way to sort it out - fiiiiight!

But seriously, I have been working with both EMC and HDS for the last 6 years on large scale deployments and you can bet I have my opinion - Nigel, more opinion than I can put on a comment on your site, so forgive me for hijacking your post.

Firstly, the USP and DMX have fundamentally the same architecture. Front end adaptor ports and processors, centralised and replicated cache and disks on back-end directors. All components are connected to each other providing a "shared everything" configuration. Both arrays use hard disk drives from the same manufacturers which have similar performance characteristics. Both offer multiple drive types, the DMX3 including 500GB drives.

From a scalability perspective the (current) USP scales to more front-end ports but can't scale to the capacity of the DMX3. Personally, I think the DMX3 scaling is irrelevant. Who in their right mind would put 2400 drives into a single array (especially with only 64 FC ports)? The USP offers 4GB FC ports, I'm not sure if DMX3 offers that too. The USP scales to 192 ports, the DMX3 only 64 (or 80 if you lose some back-end directors).

The way DMX3 and USP disks are laid out is different. The USP groups disks into array groups depending on the RAID type - for instance a 6+2 RAID group has 8 drives. It's then up to you how you carve out the LUNs - they're completely customisable to your choice of size. Although a configuration file can be loaded (like an EMC binfile) its usually never used and LUNs are user-created through a web interface to the USP SVP called Storage Navigator. LUN numbering is also user configured, so it's possible to carve all LUNs consecutively from the same RAID group - not desirable if you assign LUNs sequentially and put them on the same host. EMC split physical drives into hypers. Hypers are then recombined to create LUNs - two hypers for a RAID1 LUN, 4 hypers for a RAID5 LUN. The hypers are selected from different (and usually opposing) back-end FC loops to provide resiliency and performance. It is possible for users to create LUNs on EMC arrays (using Solutions Enabler), but usually not done. Customers tend to get EMC to create new LUNs via a binfile change which replaces the mapping of LUNs with a new configuration. This can be a pain as it has to go though EMC validation and the configuration has to be locked for new configurations until EMC implement the binfile.

For me, the main difference is how features such as synchronous replication are managed. With EMC, each LUN has a personality even before it is assigned to a host or storage port. This may be a source LUN for SRDF (an R1) or a target LUN (an R2). Replication is defined from LUN to LUN, irrespective of how the LUNs are then assigned out. HDS on the other hand, only allow replication for LUNs to be established once they are presented on a storage port and the pairing is based on the position of the LUN on the port. This isn't easy to manage and I think prone to error.

Now we come to software. EMC wipe the floor with HDS at this point. Solutions Enabler, the tool used to interact with the DMX is slick, simple to operate and (usually) works with a consistent syntax. The logic to ensure replication and point-in-time commands don't conflict or lose data is very good and it takes a certain amount of effort to screw data up. Solutions Enabler is a CLI and so quick to install and a "lite" application. There's a GUI version (SMC) and then the full blown ECC.

HDS's software still leaves a lot to be desired. Tools such as Tuning Manager and Device Manager are still cumbersome. There is CLIEX, which provides some functionality via the command line, but none of it is as slick as EMC. Anyone who uses CCI (especially earlier versions) will know how fraught with danger using CCI commands can be.

For reliability, I can only comment on my experiences. I've found HDS marginally more reliable than EMC, but that's not to say DMX isn't reliable.

Overall, I'd choose HDS for hardware. I can configure it more easily, it scales better, and - as Hu mentions almost weekly, it supports virtualisation (more on that in a moment). If I was dependent on a complex replication configuration, then I'd choose EMC.

One feature I've not mentioned earlier is virtualisation. HDS USP and NSC55 offer the ability to present externally connected arrays and present them as HDS storage. There are lots of benefits for this - migration, cost saving etc. I don't need to list them all. It's true that virtualisation is a great feature but it is *not* free and you have to look at the cost benefit of using it - or beat your HDS salesman up to give you it for free. Another useful HDS feature is partitioning. An array can be partitioned to look like up to 32 separate arrays. Great if you want to segment cache, ports and array groups to isolate for performance or security.

There are lots of other things I could talk about but I think if I go on much further I will start rambling...

Tuesday, 24 April 2007

Optimisation tools

Large disk arrays can suffer from an imbalance of data across their RAID/parity groups. This is inevitable even if you plan your LUN allocation as data profiles change over time and storage is allocated and de-allocated.

So, tools are available. Think of EMC Optimizer, HDS Cruise Control and Volume Migrator.

I've put a poll up on the blog to see what people think - I have my own views and I'll save them until after the vote closes next week.

Goodbye ASNP

It's all over. ASNP is no more. Not really a surprise as it stood for nothing useful. With 2500 members, it could have been so much more, however I think it won't be missed.

Kryptonite Discovered!

Totally off post but fantastic non the less!! Kryptonite Discovered

Monday, 23 April 2007

Hurrah for EMC

Hurrah! EMC has implemented SMI-S v1.2 in ControlCenter and DMX/Clariion (although the reference on the SNIA website seems to relate to SMI-S v1.1). Actually it seems that you need ECC v6.0 (not out yet and likely to be a mother of an upgrade from the current version) and I'd imagine the array support has been achieved using Solutions Enabler.

So quick poll, how many of you out there are using ECC to manage IBM DSxxx or HDS USP arrays? How many of you are using HSSM to manage ECC arrays? How many of you are using IBM TPC to manage anything other than DSxxx arrays??

Simulator Update

Following a few comments on the previous simulator post, it doesn't look like there are any more simulators out there for general use.

If anyone does know - feel free to comment!

Simulator Update

Following a few comments on the previous simulator post, it doesn't look like there are any more simulators out there for general use.

If anyone does know - feel free to comment!