Thursday, 28 September 2006

To Copy or not to Copy?

Sony have just announced the availability of their next generation of AIT; version 5. More on AIT here http://www.storagewiki.com/ow.asp?Advanced%5FIntelligent%5FTape. Speeds and feeds; 400GB native capacity, 1TB compressed and 24MB/s write speed. The write speed seems a little slow for me but the thing that scares me more are the capacity figures. 1TB - think about it 1TB! That's a shedload of data. Whilst that's great for density in the datacentre, it isn't good if you have a tape error. Imagine a backup failing after writing 95% of the data of a 700GB backup - or more of an issue, finding multiple read errors on a TB of data on a single cartridge.

No-one in their right mind would put 1TB of data onto a disk and hope the disk would never fail. So why do we do it with tape? Well probably because tape was traditionally used as a method of simply recovering data to a point in time. If one backup wasn't usable, you went back to the previous one. However, the world is a different place today. Increased regulation means backups are being used to provide data archiving facilities in the absence of proper application based archival. This means that every backup is essential as it indicates the state of data at a point in time. Data on tape is therefore so much more valuable than it used to be.

So, I would always create duplicate backups of those (probably production) applications which are most valuable and can justify the additional expense. That means talking to your customers and establishing the value of backup data.

Incidentally, you should be looking at your backup software. It should allow restarting backups after hardware failure. It should also allow you to easily recover the data in a backup from a partially readable tape. I mean *easily*, not oh, it can be done, but its hassle and you have to do a lot of work. Alternatively, look at D2D2T (http://www.storagewiki.com/ow.asp?Disk%5FTo%5FDisk%5FTo%5FTape).....

Monday, 25 September 2006

It's not easy being green

OK, not a reference to a song by Kermit the Frog, but a comment relating to an article I read on The Register recently (http://www.theregister.co.uk/2006/09/22/cisco_goes_green/). Cisco are attempting to cut carbon emissions by cutting back on travel, etc. Whilst this is laudible, Cisco would get more green kudos if they simply made their equipment more efficient. A 10% saving in power/cooling for all the Cisco switches in the world would make the reduction on corporate travel look like a drop in the ocean.

Sunday, 24 September 2006

Standards for Shelving

Just found this the other day: http://www.sbbwg.org/home/ - a group of vendors working to get a common standard for disk shelves. Will we see a Clariion shelf attached to a Netapp filer head?

Common Agent Standards

The deployment of multiple tools into a large storage environment does present problems. For example, EMCs ECC product claims to support HDS hardware and it does. However it didn't support the NSC product correctly until version 5.2 SP4. Keeping the agents up to date for each management product to get it to support all hardware is a nightmare. I haven't even discussed host issues. Simply having to deploy multiple agents to lots of hosts presents a series of problems; will they be compatible, how much resource will they all demand from the server, how often will they need upgrading, what level of access will be required?

Now the answer would be to have a common set of agents which all products could use. I thought that's what CIM/SMI was supposed to provide us, but at least 4 years after I read articles from the industry saying the next version of their products would be CIM compatible I still don't see it. For instance, looking at the ECC example I mentioned above, ECC has Symm, SDM, HDS and NAS agents to manage each of the different components. Why can't a single agent collect for each and any subsystem?

Hopefully, someone will correct me and point out how wrong I am, however in the meantime, I'd like to say what I want:

  1. A single set of agents for every product. This would also be a single set of hosts agents. None of this proxy agent nonsense where another agent has to sit in the way and manage systems on behalf of the system itself.
  2. A consistent upgrade and support model. All agents should work with all software, however if the wrong version of an agent is installed, then it simply reports back on the data available.
  3. The ability to upgrade any agent to introduce new features without direct dependence on software upgrades.

Thursday, 21 September 2006

Pay Attention 007...

Sometimes things you should spot just pass you by when you're not paying attention. So it is for me with N-Port ID Virtualisation. Taking a step back; mainframes had virtualisation in the early 90s. EMIF (ESCON Multi-Image Facility) allowed a single physical connection to be virtualised and shared across multiple LPARs (domains). When I was working on Sun multi-domain machines a few years ago, I was disappointed to see I/O boards couldn't share their devices between domains, so each domain needed dedicated physical HBAs. More recently I've been looking at having large servers with lots of storage using a data and tape connection - hopefully through the same physical connection, but other than using dual port HBAs, it wasn't easily possible without compromising quality of service. Dual port HBAs don't really solve the problems because I still have to pay for extra ports on the SAN.

Now Emulex have announced their support of N-Port ID Virtualisation (NPIV). See the detail here http://www.emulex.com/press/2006/0918-01.html. So what does it mean? For the uninitiated, when a fibre channel device logs into a SAN, it registers its physical address (WWN, World Wide Name) and requests a node ID. This ID is used to zone HBAs to target storage devices. NPIV allows a single HBA to request more than one node ID. This means if the server supports virtual domains (like VMware, XEN, or MS Virtual Server) then each domain can have a unique WWN and be zoned (protected separately). Also this potentially solves my disk/tape issue allowing me to have multiple data types through the same physical interface. Tie this with virtual SANs (like Cisco VSANs) and I can put quality of service onto each traffic type at the VSAN level. Voila!

I can't wait to see the first implementation; I really hope it does what it says on the tin.

Tuesday, 19 September 2006

RSS Update

I've always gone on and on about RSS and XML and how good technologies they are; XML as a data exchange technology and RSS to provide information feeds in a consistent format. My ideal world is to have all vendor information published via RSS. By that I mean not just the nice press releases and how they've sold their 50 millionth hard drive this week to help orphanages, but useful stuff like product news, security advisories and patch information.

Finally vendors are starting to realise this is useful. Cisco so far seems the best although IBM look good as do HP. McDATA and Brocade are nowhere, providing their information in POHF (Plain Old HTML Format). Hopefully all vendors of note will catch up, meanwhile I've started a list on my Wiki (http://www.storagewiki.com/ow.asp?Vendor%5FNews%5FFeeds), feel free to let me know if I've missed anyone.

Friday, 15 September 2006

A Mini/McDATA Adventure?

Apparently one in 6 cars sold by BMW is a mini. It is amazing to see how they have taken a classic but fading brand and completely re-invented it into a cool and desirable product. McDATA have just released an upgraded version of their i10K product, branded the i10K Xtreme. http://www.mcdata.com/about/news/releases/2006/0913.html This finally brings 4Gb/s speed and VSAN/LPAR technology allowing resources in a single chassis to be segmented for improved managability and security. I understand that the Brocade acquisition of McDATA is proceeding at a pace. It will be interesting to see how McDATAs new owners will treat their own McDATA "mini" going forward.

Wednesday, 30 August 2006

More on solid state disks

I'm working on a DMX-3 installation at the moment. For those who aren't aware, EMC moved from the multi-cabinet Symmetrix-5 hardware (8830/8730 type stuff) to the DMX which was fixed size - DMX1000, 2000 and 3000 models of 1, 2 and 3 cabinet installations respectively. I never really understood this move; yes it may have made life easier for EMC Engineering and CEs as the equipment shipped with its final footprint configuration, only cache/cards and disk could be added; but you had to think and commit upfront to the size of the array you wanted, which might not be financially attractive. Perhaps the idea was to create a "storage network" where every bit of data could be moved around with the minimum of effort; unfortunately that never happened and is unlikely to in the short term.

Anyway, I digress; back to DMX-3. So, doing the binfile (configuration) for the new arrays, I notice we lose a small amount of data on each of the first set of disks installed. This is for vaulting. Track back to previous models; if power was lost or likely to be lost, an array would destage all uncommitted tracks to disk after blocking further I/O on FAs. Unfortunately in a maximum configuration of 10 cabinets of 240 disks, it simply wouldn't be possible to provide battery backup for all the hard drives to destage the data. A quick calculation shows a standard HDD consumes 12.5 watts of power (300GB model), so that's 3000W per cabinet or 30,000W for a full configuration. Imagine the batteries needed for this, just on that rare off chance that power is lost. Vaulting simplifies the battery requirements by creating a save area on disk to which the pending tracks are written. When the array powers back up, the contents of cache, the vault and disk are compared to return the array back to the pre-power loss position.

This is a much better solution than simply shoving more batteries in.

So moving on from this, I thought more about exactly what 12.5W per drive means. Imagine a full configuration of 10 cabinets of 240 drives (an unlikely prospect, I know) which requires 30,000W of power. In fact the power requirements are almost double that. This is a significant amount of energy and cooling.

Going back to solid state disks I mentioned some time ago, I'd expect that the power usage would come down considerably depending on the percentage of writes that go to NAND cache. NAND devices are already up to 16GB (so could be 5% of a current 300GB drive) and growing. If power savings of 50% can be achieved, then this could have a dramatic effect on datacentre design. Come on Samsung, give us some more details of your flashon drives and we can see what hybrid drives can do......

Monday, 28 August 2006

EMC and storage growth

I had the pleasure of a tour around the EMC manufacturing plant in Cork last week. Other than the obvious interest in attention to quality I experienced, the thing that struck me most was the sheer volume of equipment being shipped out the door. Now I know the plant ships to most of the world bar North America, however seeing all those DMX and Clariion units waiting to be shipped out was amazing.

It was assuring to see the volume of storage being shipped all over the world - especially from my position as a Storage Consultant. However it brings home more than ever the challenges going forward of managing ever increasing volumes of storage. I also spent some time to see the latest versions of EMC products such as ECC, SAN Advisor and SMARTS.

ECC has moved on. It used to be a monolithic ineffectual product in the original releases at version 5, however the product seems wholly more usable. I went back after the visit and looked at an ECC 5.2 SP4 installation. There were features I should have been using; I will be using ECC more.

SAN Advisor looks potentially good - it matches your installation to the EMC Support Matrix (which used to be a PDF and now is a tool) and highlights any issues of non-conformance. In a large environment SAN Advisor would be extremely useful, however not all environments are EMC only and multi-vendor support for the tool will be essential. Secondly, the interface seems a bit clunky, work needed there. Lastly, I'd want to add my own rules and want to make them complex - so for instance where I was migrating to new storage, I'd want to validate my existing environment against it and to highlight devices not conforming to my own internal support matrix.

SMARTS uses clever technology to perform root cause analysis of faults in an IP or Storage Network. The IP functionality looked good however at this stage I could see limited appeal on the SAN side.

All in all, food for thought and a great trip!

Monday, 14 August 2006

Convergence of IP and Fibre Channel

I'm working on Cisco fibre channel equipment at the moment. As part of a requirement to provide data mobility (data replication over distance) and tape backup over distance, I've deployed a Cisco fabric between a number of sites. Previously I've worked with McDATA equipment and Brocade (although admittedly not the latest Brocade kit) and there was always a clear distinction between IP technologies and FC.

Working with Cisco, the boundaries are blurred. There are a lot of integrated features which start to blur the distinctions between the two technologies. This I guess should be no surprise with the history of Cisco as the kings of IP networking. What is interesting is to see how this expertise is being applied to fibre channel. OK, so some things are certainly non-standard. For example, the Cisco implementation of VSANs is certainly not a ratified protocol. Connect a Cisco switch to another vendor's technology and VSANs will not work. However VSANs are useful for segmenting both resources and traffic.

As things move forward following the Brocade/McDATA takeover (is it BrocDATA or McCade?) the FC world is set to get more interesting. McDATA products were firmly rooted in the FC world (the implementation of IP was restricted to a separate box and not integrated into directors) - Brocade seem a bit more open to embracing integrated infrastructure. Keep watching, it's going to be fun....

Tuesday, 8 August 2006

Tuning Manager continued

OK

I've got Tuning Manager up and running now (this is version 5). I've configured the software to pick up data from 5 arrays, which are a mixture of USP, NSC and 9980. The data is being collected via two servers - each running the RAID agent.

After 5 days of collecting, here are my initial thoughts;

  1. The interface looks good; a drastic improvement on the previous version.
  2. The interface is quick; the graphs are good and the operation is intuitive.
  3. The presented data is logically arranged and easy to follow.

However there are some negatives;

  1. Installation is still a major hassle; you've got to be 100% sure the previous version is totally uninstalled or it doesn't work.
  2. RAID agent configuration (or rather I should call it creation) is cumbersome and has to be done at the command line; the installation doesn't install the tools directory onto your command line path, so you have to trawl to the directory (or install something like "cmdline here").
  3. The database limit is way too small; 16,000 objects simply isn't enough (an object is a LUN, path etc). I don't want to install lots of instances, especially when the RAID agent isn't supported under VMware.

Overall, so far this version is a huge improvement and makes the product totally usable. I've managed to set the graph ranges to show data over a number of days and I'm now spending more time digging into the detail; Tuning Manager aggregates data up to an hourly view. The next step is to use Performance Reporter to look at some realtime detail....

Brocade and McDATA; a marriage made in heaven?

So as everyone is probably aware, Brocade are buying McDATA for a shedload of shares; around $713m for the entire company. That values each share at $4.61, so no surprise McDATA shares are up and Brocade's are down. It looks like McDATA is the poorer partner and Brocade have just bought themselves a set of new customers.

I liked McDATA; the earlier products up to the 6140 product line were great. Unfortunately the i10K for me set their demise. It created a brand new product line with no backward compatibility. It was slow to market, OEM vendors took forever to certify it. Then there was the design - physical blade swap out for 4GB; no paddle replacement; a separate box for routing; all the product operating systems are different between the routing, legacy and i10K products. Moving forward, where's the smaller i10K model like a 128 or 64 port version?

The purchase of Sanera and CNT didn't appear to go well; the product roadmap lacked strategy.

So, time to talk to my new friends at Brocade, I bet they've got a smile on their face today...

Wednesday, 2 August 2006

Holiday is over

I've been off on holiday for a week (Spain as it happens), now I'm back. I'm pleased to say I didn't think of storage once. Not true actually; the cars we'd hired didn't have enough storage space for the luggage - why does that always happen?

So back to work and continuing on virtualisation. For those who haven't read the previous posts, I'm presenting AMS storage through a USP. I finished the presentation of the storage; the USP can see the AMS disks. These are all about 400GB each, to ensure I can present all the data I need.

Now, I'm presenting three AMS systems through one USP. A big issue is how cache should be managed. Imagine the write process; a write operation is received by the USP first and confirmed to the host after writing to USP cache. The USP then destages the write afterwards down to the AMS - which also has cache - and then to physical disk. Aside from the obvious question of "where is my data?" if there is ever a hardware or component failure, my current concern is managing cache correctly. The AMS has only 16GB of cache, that's a total of 48GB in all three systems. The USP has 128GB of cache, over twice the total of the AMS systems. It's therefore possible for the USP to accept a significant amount of data - especially if it is the target of TrueCopy, ShadowImage or HUR operations. When this is destaged, the AMS systems run the risk of being overwhelmed.

This is a significant concern. I will be using Tuning Manager to keep on top of the performance stats. In the meantime, I will configure some test LUNs and see what performance is like. The USP also has a single array group (co-incidentally also about 400GB) which I will use to perform a comparison test.

This is starting to get interesting...

Wednesday, 19 July 2006

HDS Virtualisation

I may have mentioned before that I'm working on deploying HDS virtualisation. I'm deploying a USP100 with no disk (well 4 drives, the bare minimum) virtualising 3x AMS1000 with 65TB of storage each. So now the tricky part; how to configure the storage and present it through the USP.

The trouble is; with the LUN size that the customer requires (16GB), the AMS units can't present all of their storage. The limit is 2048 devices per AMS (whilst retaining dual pathing), so that means having either only 32TB of usable storage per AMS or increasing the LUN size to 32GB. Now that presents a dilemma; one of the selling points of the HDS solution is the ability to remove the USP and talk directly to the AMS if I so chose to remove virtualisation (unlikely in this instance but as Sean Connery learned, you should Never Say Never). I can't present the final LUN size from the AMS of 16GB, I'll have to present larger LUNs, carve them up using the USP and forego the ability remove the USP in the future. In this instance this may not be a big deal, but bear it in mind, for some customers is may be.

So, presentation will be a 6+2 array group; 6x 300GB which actually results in 1607GB of usable storage. This is obviously salesman sized disk allocations; my 300GB disk actually gives me 267.83GB... I'll then carve up this 1607GB of storage using the USP. At this point it is very important to consider dispersal groups. A little lesson for the HDS uninitiated here; the USP (and NSC and 99xx before it) divides up disks into array groups (also called RAID groups), which with 6+2 RAID, is 8 drives. It is possible to create LUNs from the storage in an array group in a sequential manner, i.e. LUN 00:00 then 00:01, 00:02 and so on. This is a bad idea as the storage for a single host will probably be allocated out sequentially by the Storage Administrator and then all the I/O for a single host will be hitting a small number of physical spindles. More sensible is to disperse the LUNs across a number of array groups (say 6 or 12) where the 1st LUN comes from the first array group, the second from the second and so on until the series repeats at the 7th (or 13th using our examples) LUN. This way, sequentially allocated LUNs will be dispersed across a number of array groups.

Good, so lesson over; using external storage as presented, it will be even more important to ensure LUNs are dispersed across what are effectively externally presented array groups. If not, performance will be terrible.

Having thought it over, what I'll probably do is divide the AMS RAID group into four and present four LUNs of about 400GB each. This will be equivalent to having a single disk on a disk loop behind the USP, as internal storage would be configured. This will be better than a single 1.6TB LUN. I hope to have some storage configured by the end of the week - and an idea of how it performs; watch this space!

Monday, 17 July 2006

EMC Direction

EMC posted their latest figures. So they're quoting double digit revenue growth again (although I couldn't get the figures to show that). My question is; where is EMC going? The latest DMX and Clariion improvements are just that - performance improvements over the existing systems. I don't see anything new here. The software strategy seems to be to purchase lots of technology, but where's the integration piece, where's the consolidated product line? ECC still looks as poor as ever.

So what is the overall strategy? I think it's to be IBM. Shame IBM couldn't hold on to their dominant position in the market. I can see EMC going the same way as people overtake and improve on their core technologies.

Data Migration Strategies

Everyone loves the idea of a brand-new shiny SAN or NAS infrastructure. However over time this new infrastructure needs to be maintained. Not just at the driver level, but eventually arrays, fabrics and so on. So, data migration will become a continous BAU process we'll all have to adopt. Whilst I think further, I've distilled the requirements into some simple criteria:

Migration Scenarios

Within the same array
Between arrays from the same vendor
Between arrays from different vendors
Between similar protocol types
Between different protocol types

Migration Methods

Array based
Host Based
3rd Party Host based

Migration Issues

Performance – time to replicate, ability to keep replica refreshed
Protection – ensuring safety of the primary, secondary and any replicated copies
Platform Support – migration between storage technologies
Application Downtime – ensuring migration has minimum application downtime

Migration Reasons

Removal of old technology
Load balancing
Capacity balancing
Performance improvements (hotspot elimination)

Time to do some more thinking about filling out the detail. Incidentally, when new systems are deployed, a lot of though goes into how to deploy the new infrastructure, how it will work wth existing technology - how much time goes into planning how that system will be removed in the future?

Tuesday, 11 July 2006

More on Virtual Tape

Last week I talked about virtual tape solutions. On Monday HDS released the news that they're reselling Diligent's VTL solution. From memory, when I last saw this about 12 months ago, it was a software only solution for emulating tape, sure enough that's still the same. Probably the more useful thing though is ProtecTIER, which scans all incoming data and proactively compresses the incoming stream by de-duplicating the data. The algorithm relies on a cache map of previously scanned data held in memory on the product's server, with a claimed supported capacity of 1PB raw, or 25PB at the alleged 25:1 compression ratio.

As Hu Yoshida referenced the release on his blog, I asked the question on compression. 25:1 is a claimed average from customer experience, some customers have seen even greater savings.

So, my questions; what is performance like? What happens if the ProtecTIER server goes titsup (technical expression)? I seem to remember from my previous presentation on the product that the data was stored in a self referential form, allowing the archive index to be easily recreated.

If the compression ratio is true and performance is not compromised, then this could truly be a superb product for bringing disk based backup into the mainstream. Forget the VTL solution; just do straight disk based backups using something like Netbackup disk storage groups and relish in the benefits!!!

Monday, 10 July 2006

Performance End to End

Performance Management is a recurring theme in the storage world. As fibre channel SANs grow and become more complex, the very nature of a shared infrastructure becomes prone to performance bottlenecks. Worse still, without sensible design (e.g. things like not mixing development data in with production) production performance can be unnecessarily compromised.

The problem is, there aren't really the tools to manage and monitor performance to the degree I'd like. Here's why. Back in the "old days" of the mainframe, we could do end to end performance management - issue an I/O and you could break down the I/O transaction into the constituent parts; you could see the connect time (time the data was being transferred), disconnect time (time waiting for the data to be ready so it can be returned to the host) and other things which aren't quite as relevant now like seek time and rotational delay. This was all possible because the mainframe entity was a single infrastructure; there was a single time clock against which the I/O transaction could be measured - also the I/O protocol catered for collecting the I/O data response times.

SANs are somewhat different. Firstly the protocol doesn't cater for collecting in-flight performance statistics, so all performance measurements are based on observations from tracing the entire environment. The vendors will tell you they can do performance measurements and it is true, they can collect whatever the host, storage and SAN components offer - the trouble is, those figures are likely to be averages and either not possible or not easy to relate the figures to specific LUNs on specific hosts.

For you storage and fabric vendors out there, here's what I'd like; first I want to trace an entire I/O from host to storage and back; I want to know at each point in the I/O exchange what the response time was. I still want to total and average those figures. Don't forget replication - TrueCopy/SRDF/PPRC, I still want to know that part of the I/O.

One thought, I have a feeling fabric virtualisation products might be able to produce some of this information. After all, if they are receiving an I/O request for a virtual device and returning it to the host, the environment is there to map the I/O response to the LUN. Perhaps that exists today?

Wednesday, 5 July 2006

Virtual tape libraries

I previously mentioned virtual tape libraries. Two examples of products I've been looking at are Netapp's Nearstore Virtual Tape Library and ADIC's Pathlight products. Both effectively simulate tape drives and allow a virtual tape to be exported to a physical tape. Here are some of the issues as I see it:

1. How many virtual devices can I write to? Products such as Netbackup do well having lots of drives to write to; a separate drive is needed (for instance) for each retention period (monthly/weekly/daily) and for different storage pools. This can cause issues with the ability to make most effective use of drives, especially when multiplexing. So, the more drives, the better.

2. How much data can I stream? OK, it's great having lots of virtual drives, but how much data can I actually write? The Netapp product for instance can have up to 3000 virtual drives but can only sustain 1000MB/s throughput (equivalent to about 30 LTO2 drives).

3. How is compression handled? Data written to tape will usually be compressed by the drive, resulting in variable capacity on each tape, some data will compress well, some won't. A virtual tape system that writes to physical tape must ensure that the level of compression doesn't prevent a virtual tape being written to the physical tape (imagine getting such poor compression that only 80% of data could be written to the tape, what use is that). However the flip side to this is ensuring that all the tape capacity can be utilised - it's easy to simply write 1/2 full tapes to get around the compression issue.

4. How secure is my data? So you now have multiple terabytes of data on a single virtual tape unit. How is that protected? What RAID protection is there? How is the index of data on the VTL protected? Can the index be backed up externally? One of the great benefits of tape is the portability. Can I replicate my VTL to another VTL? If so, how is this managed?

5. What is the TCO? Probably one of the most important questions. Why should I buy a VTL when I could simply buy more tape drives and create an effective media ejection policy? The VTL must be cost effective. I'll touch on TCO another time.

6. The Freedom Factor. How tied will I be to this technology? The solution may not be appropriate, the vendor may go out of business. How quickly can I extricate myself from the product.

Tuesday, 4 July 2006

The Incredible Shrinking Storage Arrays

For those who don't relate to the title, check out IMDB....

You know how it is, you go to buy some more storage. You need, say 10TB. You get the vendor to quote - but how much to do you actually end up with? First of all, disk drives never have the capacity they purport to have even if you take into consideration the "binary" TB versus the "decimal" TB. Next there's the effect of RAID. That's a known quantity and expected, so I guess we can't complain about that one. But then there's the effect of carving up the physical storage into logical LUNs; this can easily result in 10% wastage. Plus there's more: EMC on the DMX-3 now uses the first 30 or so disks installed into the box to store a copy of cache memory in case of a power outage; OK, good feature but it carves into the host available space. Apologies to EMC there - you were the first vendor I thought to have a dig at.

Enterprise arrays are not alone in this attritious behaviour - Netapp filers for instance will "lose" 10% of their storage to inodes (the part which keeps track of your files) and another 20% reserved for snapshots - that's after RAID has been accounted for.

What we need is clear unambiguous statements from our vendors - "you want to buy 10TB of storage? Then it will cost you 20....."