Tuesday, 16 January 2007

Slow Provisioning

Poor provisioning tools annoy me. I've been annoyed today. I've been changing some VMware metas from 100GB to 200GB on a DMX. Unfortunately they were already presented (but not used) and replicated with SRDF. So I had to:

  1. "Not Ready" the R1 and R2 drives
  2. Unmask the LUNs from the FA
  3. Split the SRDF relationship
  4. Break the SRDF relationship
  5. Unmap the LUNs from their FAs
  6. Dissolve the metas
  7. Create the metas
  8. Re-establish and resync SRDF
  9. Map the LUNs to the FA
  10. Mask the LUNs to the hosts

10 steps which take some considerable time to write validate and execute. I don't do this stuff often enough to justify writing scripts to help me out; but I think this should be a vendor thing - a software tool with various configure options that creates the symconfigure and associated commands for you and indicates the steps you will have to perform. ECC is *supposed* to do it but it doesn't. Roll on some good software.

Sunday, 14 January 2007

more about iSCSI

I mentioned as a "Storage Resolution" to look more in-depth at iSCSI. Well I've started doing just that today.

The first thing I thought I needed was a working environment. I'm not keen on investing in an entire storage array (at this stage) to do the testing (unless some *very* generous vendor out there wants to let me "loan" one) so I've build a virtual environment based on a number of free components.

I've a dedicated VMware testing machine recently built which has a dual Core Intel processor, 2GB of RAM and a SATA drive. Nice and simple. It runs Win2K3 with the free VMware server, onto which I've created another Win2K3 R2 partition and a Linux partition running Fedora Core 6. This is where my iSCSI "target" will sit.

For those unfamiliar with SCSI terminology, the source disk or disk system presents LUNs which are referred to as targets. The host accessing those LUNs is the initiator; simply put the host initiates a connection to a target device, hence the names. My iSCSI target in this instance is a copy of the Netapp simulator running on Linux.

Most people are probably aware of the simulator. If not, Dave Hitz talks about it here. I've created a number of disks into an OnTAP volume and out of that created a LUN. LUNs can be presented out as FC or iSCSI, in this instance I've presented it out as iSCSI.

By default the simulator doesn't enable iSCSI so I enabled it with the standard settings. This means my target's iSCSI address is all based on Netapp defaults. I'm going to work on what the best practices should be for these settings over the coming days. Anyway, I've presented 2 LUNs and numbered them LUN 4 and LUN 9.

At the initiator (host) end, I've used my Win2K3 Server and installed the iSCSI initiator software from Microsoft. This gave me a desktop icon to configure the settings. Again, I've ended up with the default names for my iSCSI initiator, but that doesn't matter; all I had to do was specify in the iSCSI initiator settings the IP address of my target, log on and it finds the LUNs (oh, one small point, I had to authorise the initiator on the simulator). Voila, I now have 2 disks configured to my Windows host which can be formatted as standard LUNs.

As a performance test, I ran HdTach against the iSCSI LUNs on Win2k3. I got a respectable 45MB/s throughput, which isn't bad bearing in mind this environment is all virtual on the same physical machine.

All the above sounds a bit complicated, so I'll break it down over the coming days as to what I had to do; I'll also explain the iSCSI settings I needed to make and my experiments with dual pathing and taking the iSCSI devices away from Windows in mid-operation.

Thursday, 11 January 2007

WWN Decoding

The WWN Decoder page on my main site is updated. It's at http://www.brookend.com/html/main/wwndecoder.asp

If I've done it correctly, it now handles EMC up to DMX3 and HDS USP/NSC/AMS and also indicates the model type of the discovered device.

If anyone has samples from their arrays which they are willing to share, let me know and I'll validate them to see if it helps amend the decoder.

Friday, 5 January 2007

Manic Miner and Storage Resource Management

I tried a bit of nostalgia the other day. From a "freebie" CD-ROM I installed a games emulator for the ZX Spectrum, a personal computer that was hugely popular in the '80s. The game I installed was called Manic Miner, one of the original platform games. At the time (1983) it was a classic and (shamefully) I even hacked the copy I had to remove the protection (you had to load a 4 or 6 digit code from a sheet of blue paper, which couldn't be photocopied). When my children saw the game, they fell about laughing, not surprising when you compare it to their latest play, Star Wars Battlefront.

It made me think how things have changed in 20 years; from 32x24 graphics to 1280x1024 with advanced polygon shading etc. What has this to do with storage? Well, I ponder on what will happen to Storage Resource Management in the next 20 years.

I think what we'll see is artificial intelligence-based software managing our data. The software will proactively fix hardware faults, relocate data based on our usage/value policies, provide CDP and CDR, deliver optimum performance and make all storage administrators obsolete.

Er, well all except the last one; yes I do think the worries we have about SRM tools will be resolved, however I think with the growth in capacity, complexity and features of todays storage, that Storage Administrators will be needed for a long time to come.

Tower of Tera

Lots of talk today about the 1 terabyte drive from Hitachi. In fact the drive is more likely to be about 931GB based on the dubious practice of using decimal 1000's rather than binary (whilst we're on that subject, the concept of decimal versus binary does annoy me - what with that and overhead, on some of the AMS's I've installed, a 300GB drive comes out as 267GB).

So, yes, I want a 1TB drive - no idea what I want to put on it, or how I'll back it up - but I want one.

Thursday, 4 January 2007

Hybrid Storage Alliance




Fujitsu, Hitachi, Toshiba, Samsung and Seagate are getting together as the Hybrid Storage Alliance to promote the use of HDDs with large amounts of additional cache. They've set up a (not so snazzy) website at www.hybridstorage.org (check out the builder on the features page - when did you last see a builder using a laptop, never mind understanding what a hard disk is).

It's good news. I love the idea and have mentioned my thoughts before. I also think that onboard cache provides more options to develop the successor to RAID, although I'm still thinking how it could be done.

Sandisk also announced the device shown on the right - a solid state HDD. After 50 years, the hard disk is seeing some exciting changes.

Wednesday, 3 January 2007

My Favourite De-Duplication Technology

Here's one of my favourite websites; www.shazam.com. In fact it isn't the website that is the favourite thing, it's what Shazam do. In the UK (apologies for US readers, I don't know your number), dialling 2580 down the middle of your 'phone and holding your mobile up to a music source for 30 seconds will give you a text back with the track title and the artist. Seeing this for the first time is amazing; as long as the track is reasonably clear, any 30 second clip will usually work. I've astounded (and bored) dozens of friends and it only costs me 50p each time.

So, Shazam got me thinking. How can they track the almost millions of music tracks in existence today and match this to a random clip of music I provide over a tinny link from my mobile phone? The most obvious issues are those of quality; I've almost only ever used the service in a bar with a lot of background noise (mostly druken colleagues). That aside, I tried to see how they could have indexed all the tracks and still allowed me to provide a random piece of the track against which they match.

I started thinking about pattern matching and data de-duplication as it exists today. Most de-dupe technology seems to rely on identifying common patterns within data and indexing those against a generated "hash" code which (hopefully) uniquely references that piece of data. With suitable data containing lots of similar content then a lot of duplication can be removed from storage and referenced as pointers. Good examples would be backups of email data (where either the user or a group of users share the same content) and database backups where only a small percentage of the database has changed. The clever de-dupe technology would be able to identify variable length patterns and determine variable start positions (i.e. byte level granularity) when indexing content. This would be extremely important where database compression re-aligns data on non-uniform boundaries.

Now this is where I failed to understand how Shazam could work; OK, so the source content could be de-duped and indexed, but how could they determine where my sample occurred in the music track? A simple internet search located the following presentation from the guy (Avery Wang) who developed the technology. The detail is here http://ismir2003.ismir.net/presentations/Wang.PDF. The de-dupe process actually generates a fingerprint for each track, highlighting specific unique spectogram peaks in the sounds of the music, then uses a number of these to generate hash tokens via a technique called "combinatorial hashing". This uniquely identifies a track, but also provides the answer as to how any clip can be used to identify a track; the relative offsets of each hash token is used to identify the track, so the absolute offset of the sample isn't important.

Anyway, enough of the techie talk, try Shazam - amaze your friends!

Tuesday, 2 January 2007

Storage Resolutions

The new year is here. Everyone loves to make resolutions to say how they are going to improve their lives. Personally I think it is nonsense; if you want to change, then you can do it any time rather than the abitrary time of new year.

Anyway, enough of my humbug. Here's a few storage resolutions I hope to maintain:


  • iSCSI - I haven't paid enough attention to this. I think iSCSI is due to hit its tipping point this year and get much more widespread adoption.
  • WAFS - Wide Area File Systems interest me. Any opportunity to reduce the volume of data being moved across networks while centralising the gold copy strikes me as a sensible idea.
  • CDP - There are some interesting products around providing continuous data protection. They aren't scalable yet, but when they are I can see them being big.
  • NAS Virtualisation - OK, I know how it works and what the products are; I just need to get into more detail.
  • CAS - I've always seen CAS as pointless. It's time to give it a second chance.

So that's the technology side covered. What about process?

  • ILM - I think it is time to harp on about proper ILM - i.e. that which is integrated into the application rather than the poor efforts we've seen to date. I think application development needs to be addressed to cover this.
  • SRM - how about some proper tools which actually do the job of managing the (large scale) process of storage deployment? More thought required here.
  • Cost Management - I believe there are lots of options for managing and reducing cost, I should expound on them more.
  • Technology Refresh - Always a problem and certainly needs more thought for mature datacentres.

Hmm, funny how each year's resolutions end up sounding just like the ones the year before?

Netapp 1 EMC 0

My RSS reader just picked up a lovely report from Netapp countering an EMC report showing that with MS-Exchange workloads, EMC was better than Netapp (CX3-40 v 3050). It just shows how when vendors are challenged to defend their products, they can make them work much better than the "standard" configuration. I can't help thinking it would be better if these products did this without needing the vendor to do a lot of configuration work.

Original report here.... http://www.netapp.com/library/tr/3521.pdf

Thursday, 21 December 2006

Understanding Statistics

I've been reading a few IDC press releases today. The most interesting (if any ever are) was that relating to Q3 2006 revenue figures for the top vendors. It goes like this:

Top 5 Vendors, Worldwide External Disk Storage Systems Factory Revenue 3Q2006 (millions)

EMC: $927 (21.4%)
HP: $760 (17.6%)
IBM: $591 (13.7%)
Dell: $347 (8.0%)
Hitachi: $340 (7.9%)

So EMC comes out on top, followed by the other usual suspects. EMC gained market share from all other vendors including "The Others" (makes me think of Lost - who are those "others"?). However, IDC also quote the following:

Top 5 Vendors, Worldwide Total Disk Storage Systems Factory Revenue 3Q2006 (millions)

HP: $1406 (22.7%)
IBM: $1250 (20.2%)
EMC: $927 (15.0%)
Dell: $507 (8.2%)
Hitachi: $348 (5.6%)

So what does this mean? IDC defines a Disk Storage System as at least 3 disk drives and the associated cables etc, to connect them to a server. This could mean 3 disks with a RAID controller in a server. Clearly EMC don't ship anything other than external disks as their figures are the same in each list. HP make only 50% of their disk revenue from external systems, the rest presumably are disks shipped with servers, IBM even less as a percentage, Dell about $160m. The intruiging one was Hitachi - what Disk Storage Systems to they sell (other than external) which made $8m of revenue? The source of my data can be found here: http://www.idc.com/getdoc.jsp?containerId=prUS20457106

What does this tell me? It says there's a hell of a lot of DAS/JBOD stuff still being shipped out there - about 30% of total revenue.

Now, if EMC were to buy Dell or the other way around, between them they could (just) pip HP to the post. Are EMC and Dell merging? I don't know, but I don't mind starting a rumour...

Oh, another intesting IDC article I found referred to how big storage virtualisation is about to come. I've been saying the word "virtualisation" for about 12 months in meetings just to get people to even listen to the concept, even when the meeting has nothing to do with the subject. I'm starting to feel vindicated.

Wednesday, 20 December 2006

Modular Storage Products

I’ve read a lot of posts recently on various storage related websites asking for comparisons of modular storage products. By that I’m referring to “dual controller architecture” products such as the HDS AMS, EMC Clariion and HP EVA. The questions come up time and time again, usually comparing IBM to EMC or HP and little comparison to HDS, but lots of people recommending HDS.

So, to be more objective, I’ve started compiling a features comparison of the various models from HDS, IBM, EMC and HP. Before anyone starts, I know there are other vendors out there – 3PAR, Pillar and others come to mind. At some stage, I’ll drag in some comparisons to them too, but to begin with this is simply the “big boys”. The spreadsheet attached is my first attempt. It has a few gaps where I couldn’t determine the comparable data, mainly on whether iSCSI or NAS is a supported option and the obvious problem of performance throughput.

So, from a simple physical perspective, these arrays are pretty simple to compare. EMC and HDS give the highest disk capacity options, EMC, HDS and IBM offer the same maximum levels of cache. Only HDS offers RAID6 (at the moment), most vendors offer a range of disk drives and speeds. Most products offer 4Gb/s front-end connections and there are various options for 2/4Gb/s speeds at the back end.

Choosing a vendor on physical specifications alone is simple using the spreadsheet. However there are plenty of other factors not included here. First, there’s performance. Only IBM (from what I can find) offers their arrays to scrutiny by the Storage Performance Council. Without a consistent testing method, any other figures offered by vendors are completely subjective.

Next, there’s the thorny subject of feature sets. All vendors offer variable LUN sizes, some kind of failover (I think most are active/passive), multiple O/S support, replication and remote copy functionality and so on. Comparing these isn’t simple, though as the implementation of what should be common features can vary widely.

Lastly there’s reliability and the bugs and gotchas that all products have and which the manufacturers don’t document. I’ll pick an example or two; do FC front-end ports share a multiprocessor? If so, what impact does load on one port have on the other shared port? What downtime is required to do maintenance, such as code upgrades? What is the level of SNMP or other management/alerting software?

The last set of issues would prove more difficult to track so I’m working on a consistent set of requirements from a product. In the meantime, I hope the spreadsheet is useful and if anyone can fill the gaps or wants to suggest other comparable mid-range/modular products, let me know.

You can download the spreadsheet here: http://www.storagewiki.com/attachments/modular%20products.xls

Tuesday, 19 December 2006

New RAID

Lots of people are talking about how we need a new way to protect our data and that RAID has had it. Agreed, going RAID6 gives some benefits (i.e. puts off the inevitable failure by a factor again), however the single problem to my mind with RAID today is the need to read all the other disks when a real failure occurs. Dave over at Netapp once calculated the risk of re-reading all those disks in terms of the chance of a hard failure.

The problem is, the drive is not involved in the rebuild process - it dumbly responds to the request from the controller to re-read all the data. What we need are more intelligent drives combined with more intelligent controllers; for example; why not have multiple interfaces to a single HDD? Use a hybrid drive with more onboard memory to cache reads while the heads are moving to obtain real data requests. Store that data in memory on the drive to be used for drive rebuilds. Secondly, why do we need to store all the data for all rebuilds across all drives? Why with a disk array of 16 drives can't we run multiple instances of 6+2 RAID across different sections of the drive?

I'd love to be the person who patents the next version of RAID....

Tuesday, 5 December 2006

How low can they go!


I love this picture. Toshiba announced today that they are producing a new 1.8" disk drive using perpendicular recording techniques. This drive has a capacity of 100GB!

It will be used on portable devices such as music players; its only 54 x 71 x 8 mm in size, weighs 59g and can transfer data at 100MB/s using an ATA interface.

I thought I'd compare this to some technology I used to use many years ago - the 3380 disk drive.

The model shown on the right is the 3380 CJ2 with a massive 1.26GB per unit and an access time similar to the Toshiba device. However the transfer rate was only about 3MB/s.

I couldn't find any dimensions for the 3380, but from the picture of the lovely lady, I'd estimate it is 1700 x 850 x 500mm which means 23,500 of the Tosh drives could fit in the same space!

Where will we be in the next 20 years? I suspect we'll see more hybrid drives, with NAND memory used to increase HDD cache then more pure NAND drives (there are already some 32GB drives announced). Exciting times...

Monday, 4 December 2006

Is the Revolution Over?

I noticed over the weekend that http://www.storagerevolution.com's website was down. Well it's back up today, but the forums have disappeared. I wonder, is the revolution over for JWT and pals?

Thursday, 30 November 2006

A rival to iSCSI?

I've said before, somethings just pass you by, and so it has been with me and Coraid. This company uses ATA-over-Ethernet, a lightweight IP storage protocol to connect hosts to their storage products.

It took me a while before I realised how good this protocol could be. Currently iSCSI uses TCP/IP, which encapsulates SCSI commands within TCP packets. This creates an overhead but does allow the protocol to be routed over any distance. AoE doesn't use TCP, it uses its own protocol so doesn't suffer the TCP overhead, but can't be routed. This isn't that much of an issue as there are plenty of storage networks which are locally deployed.

Ethernet hardware is cheap. I found a 24 port GigE switch for less than £1000 (£40 a port). NICs can be had for less than £15. Ethernet switches can be stacked providing plenty of redundancy options. This kills fibre channel on cost. iSCSI can use the same hardware; AoE is more efficient.

AoE is already being bundled with Linux, drivers are available for other platforms. Does this mean the end of iSCSI? I don't think so; mainly because AoE is so different and also more importantly it isn't routable. However what AoE does offer is another alternative which can use standard cheap, readily available products. From what I've read, AoE is simple to configure too; certainly with less effort than iSCSI. I hope it has a chance.

Wednesday, 22 November 2006

Brocade on the up

Brocade shares were up 10.79% today. I know they posted good results but this is a big leap for one day. I wish I hadn't sold my shares now! To be fair, I bought at $4 and sold at $8 so I did OK - and it was some time ago.

So does this bode well for the McDATA merge? I hope so. I've been working on Cisco and the use of VSANs and I'm struggling to see what real benefit I can get out of them. Example: I could use VSANs to segregate by: Line of Business; Storage Tier; Host Type (Prod/UAT/DEV). Doing this immediately gives me 10's of combinations. Now the issue is I can only assign a single storage port to one VSAN - so I have to decide; do I want to segment my resources to the extent I can't use an almost unallocated UAT storage port when I'm desperate for connectivity for production? At the moment I see VSANs as likely to create more fragmentation than anything else. There's still more thinking to do.

I'd like to hear from anyone who has practical standards for VSANs. It would be good to see what best practice is out there.

Tuesday, 14 November 2006

Backup Trawling

I had a thought. As we back up more and more data (which we must be, as the amount of storage being deployed increases at 50-100% per year, depending on who you are) then, there must be more of an issue finding the data which needs to be restored. Are we going to need better tools to trawl the backup index to find what we want?

Virtually There

VMware looks like it is going to be a challenge; I'm starting to migrate standard storage tools into a VMware infrastructure. First of all, there's a standard VMware build for Windows (2003). Where should the swap file go? How should we configure the VMFS that it sits on? What about the performance issues in aligning the Windows view of the (EMC) storage with the physical device so we get best performance? What about Gatekeeper support? and so on.

It certainly will be a challenge. First thing that I need to solve, is if I have a LUN presented as a raw device to my VM guest and that LUN ID is over 128, although the ESX server can see it, when ESX exports it directly to the Windows guest, it can't be seen. I've a theory that it could be the generic device driver on Windows 2003 that VMware uses. I can't prove it (yet) and as yet no-one on the VMware forum has answered the question. Either they don't like me (lots of other questions posted after mine have been answered) or people don't know....

Monday, 13 November 2006

Continuous Protection

I've been looking at Continuous Data Protection products today, specifically those which don't need the deployment of an agent and can use Cisco's SANTap technology. EMC have released RecoverPoint (a rebadged product from Kashya) and Unisys have SafeGuard.

SANTap works by creating a virtual initiator, duplicating all of the writes to a standard LUN (target) within a Cisco fabric. This duplicate data stream is directed at the CDP appliance which stores and forwards it to another appliance located at a remote site where it is applied to a copy of the original data.

As the appliances track changed blocks and apply them in order, they allow recovery to any point in time. This potentially is a great feature; imagine being able to back out a single I/O operation or to replay I/O operations onto a vanilla piece of data.

Whilst this CDP technology is good, it introduces a number of obvious questions on performance, availability etc, but for me the question is more about where this technology is being pitched. EMC already has replication in the array in the form of SRDF which comes in many flavours and forms. Will RecoverPoint complement these or will, in time, RecoverPoint be integrated into SRDF as another feature?

Who can say, but at this stage I see another layer of complication and decision on where to place recovery features. Perhaps CDP will simply offer another tool in the amoury of the Storage Manager.

Wednesday, 18 October 2006

Cisco, Microwaves and Virtualisation

So I've not posted in October so far and the month is half over. To my defence I spent a week in South Africa doing an equipment (Cisco) installation. No it wasn't a holiday but I did see the fantastic table mountain, which is spectacular and well worth a visit.

Part of the installation involved deployment of SRDF between two symmetrix arrays using Cisco 9216i switches. These have two IP ports and can be used for iSCSI or in this case FCIP as we had to provide the SRDF links across IP. Whilst this all sounds fine, the IP link was in fact over a microwave connection rather than a fixed line installation. Telecoms are prohibitively expensive in SA and sometimes microwave is the only option due to the high cost of fixed line installations.

Installation and configuration of the links was simple, including enabling SRDF. However my big concern is performance (I should note that this is no way a synchronous implementation, but an async one). The IP link is not dedicated to storage and being used for other purposes so fcping responses from the Cisco switches were very variable. At this point we haven't fully enabled SRDF but the next stage is to test performance of the link. As I see it, we need to monitor SRDF stats with ECC Performance Monitor, SRDF/A response times and the lag time of unwritten I/Os not committed to the remote site. The solution needs significant monitoring to ensure all links are active; with SRDF it is possible to monitor the status of a remote SRDF link using the "symrdf ping" command; this needs to be being issued reguarly if not every 5 minutes or less. More updates as I have them.

I spoke to an analyst today (it's Storage Expo in the UK) on the subject of virtualisation and in particular HDS' Universal Volume Manager. He was asking my view on where virtualisation is headed. I still think the switch is the right place, long term, for virtualisating the storage infrastructure. In the short term though, discrete hardware array virtualisation is a good thing. UVM can provide cost savings (on significant virutalised volumes of data) and additional functionality, however HDS will have to up their game to retain the virtualisation crown as time progresses. Specifically they need to address the issue of a failure in the USP when so much storage could be dependent on one subsystem. A clustered solution could be the answer here. Time will tell.