Showing posts with label barry burke. Show all posts
Showing posts with label barry burke. Show all posts

Thursday, 13 November 2008

Obligatory Atmos Post

I feel drawn to post on the details of Atmos and give my opinion whether it is good, bad, innovative or not. However there's one small problem. Normally I comment on things that I've touched - installed/used/configured/broken etc, but Atmos doesn't fit this model so my comments are based on the marketing information EMC have provided to date. Unfortunately the devil is in the detail and without the ability to "kick the tyres", so to speak, my opinions can only be limited and somewhat biased by the information I have. Nevertheless, let's have a go.

Hardware

From a hardware perspective, there's nothing radical here. Drives are all SATA-II 7.2K 1TB capacity. This is the same as the much maligned IBM/XIV Nextra, which also only offers one drive size (I seem to remember EMC a while back picking this up as an issue with XIV). In terms of density, the highest configuration (WS1-360) offers 360 drives in a single 44U rack. Compare this with Copan which provides up to 896 drives maximum (although you're not restricted to this size).

To quote Storagezilla: "There are no LUNs. There is no RAID. " so exactly how is data stored on disk? What methods are deployed for ensuring data is not lost due to a physical issue? What is the storage overhead of that deployment?

Steve Todd tells us:

"Atmos contains five "built-in" policies that can be attached to content:

  • Replication
  • Compression
  • Spin-down
  • Object de-dup
  • Versioning


When any of these policies are attached to Atmos, COS techniques are used to automatically move the content around the globe to the locations that provide those services."

So, does that mean Atmos is relying on replication of data to another node as a replacement for hardware protection? I would feel mighty uncomfortable to think I needed to wait for data to replicate before I had some form of hardware-based redundancy - even XIV has that. Worse still, do I need to buy at least 2 arrays to guarantee data protection?

Front-end connectivity is all IP based, which presumably includes replication too, although there are no details of replication port counts or even IP port counts, other than the indication of 10Gb availability, if required.

One feature quoted on all the literature is Spin Down. Presumably this means spinning down drives to reduce power consumption; but spin down depends on data layout. There are two issues; if you've designed your system for performance, data from a single file may be spread across many spindles. How do you spin down drives when they all potentially contain active data? If you've laid out data on single drives, then you need to move all the inactive data to specific spindles to spin them down - that means putting the active data on a smaller number of spindles - impacting performance and redundancy in the case of a disk failure. The way in which Atmos does its data layout is something you should know - because if Barry is right, then his XIV issue could equally apply to Atmos too.

So to summarise, there's nothing radical in the hardware at all. It's all commodity-type hardware - just big quantities of storage. Obviously this is by design and perhaps it's a good thing as unstructured data doesn't need performance. Certainly as quoted by 'zilla, the aim was to provide large volumes of low cost storage and compared to the competition, Atmos does an average job of that.

Software

This is where things get more interesting and to be fair, the EMC message is that this is a software play. Here are some of the highlights;

Unified Namespace

To quote 'zilla again:

"There is a unified namespace. Atmos operates not on individual information silos but as a single repository regardless of how many Petabytes containing how many billions of objects are in use spread across whatever number of locations available to who knows how many users."

I've highlighted a few words here because I think this quote is interesting; the implication is that there is no impact on the volume of data or its geographical dispersion. If that's the case (a) how big is this metadata repository (b) how can I replicate it (c) how can I trust that it is concurrent and accurate in each location.

I agree that a unified name space is essential, however there are already plenty of implementations of this technology out there, so what's new with the Atmos version? I would want to really test the premise that EMC can provide a concurrent, consistent name space across the globe without significant performance or capacity impact.

Metadata & Policies

It is true that the major hassle with unstructured data is the ability to manage it using metadata based policies and this feature of Atmos is a good thing. What's not clear to me is where this metadata comes from. I can get plenty of metadata today from my unstructured data; file name, file type, size, creation date, last accessed, file extension and so on. There are plenty of products on the market today which can apply rules and policies based on this metadata, however to do anything useful, then more detailed metadata is needed. Presumably this is what the statement from Steve means: "COS also implies that rich metadata glues everything together". But where does this rich metadata come from? Centera effectively required programming their API and that's where REST/SOAP would come in with Atmos. Unfortunately unless there's a good method for creating the rich metadata, then Atmos is no better than the other unstructured data technology out there. To quote Steve again:

"Rich metadata in the form of policies is the special sauce behind Atmos and is the reason for the creation of a new class of storage system."

Yes, it sure is, but where is this going to come from?

Finally, let's talk again about some of the built-in policies Atmos has:

  • Replication
  • Compression
  • Spin-down
  • Object de-dup
  • Versioning
All of these exist in other products and are not innovative. However extending policies is more interesting; although I suspect this is not a unique feature either.

On reflection I may be being a little harse on Atmos, however EMC have stated that Atmos represents a new paradigm in the storage of data. If you make a claim like that, then you need to back it up. So, still to be answered;

  • What resiliency is there to cope with component (i.e HDD) failure?
  • What is the real throughput for replication between nodes?
  • Where is the metadata stored and how is it kept concurrent?
  • Where is the rich metadata going to come from?

Oh, and I'd be happy to kick the tyres if the offer was made.

Monday, 3 November 2008

Innovation

"Innovative - featuring new methods or original ideas - creative in thinking" - Oxford English Dictionary of English, 11th edition.

There have been some interesting comments over the weekend, specifically from EMC in regard to this post which I wrote on Benchmarketing started by Barry Burke and followed by Barry Whyte.

"Mark" from EMC points me to this link regarding EMC's pedigree on innovation. Now that's like a red rag to a bull to me and I couldn't help myself going through every entry and summarising them.

There are 114 entries, out of which, I've classified 44 as marketing - for example appointing Joe Tucci (twice) and Mike Ruettgers (twice) and being inducted into the IT Hall of Fame hardly count as innovation! Some 18 entries relate directly Symmetrix, another 18 to acquisition (nnot really innovation if you use the definition above) and another 7 to Clariion (also an acquisition).

From the list, I've picked out a handful I'd classify as innovating.

  • 1987 - EMC introduce solid state disks - yes, but hang on, haven't they just claimed to have "invented" Enterprise Flash Drives?
  • SRDF & Timefinder - yes I'd agree these are innovative. SRDF still beats the competition today.
  • First cached disk array - yes innovation.

Here's the full list taken from the link above. Decide for yourself whether you think these things are innovative or not. Acquisitions in RED, Marketing in GREEN. Oh and if anyone thinks I'm being biased, I'm happy to do the same analysis for IBM, HP, HDS etc. Just point me at their timelines.

  • Clariion CX4 - latest drives, thin provisioning?
  • Mozy - acquisition
  • Flash Drives - 1980's technology.
  • DMX4 - SATA II drives and 4Gb/s
  • Berkeley Systems etc - acquisition
  • EMC Documentum - acquisition
  • EMC study on storage growth - not innovation
  • EMC floats VMware - acquisition
  • EMC & RSA - acquisition
  • EMC R&D in China
  • EMC Clariion -Ultrascale
  • EMC Smarts - acquisition
  • Symmetrix DMX3
  • Smarts, Rainfinity, Captiva - acquisitions
  • EMC - CDP - acquisition
  • EMC Clariion - Ultrapoint
  • EMC DMX3 - 1PB
  • EMC Invista - where is it now?
  • EMC Documentum - acquisition
  • Clariion AX100 - innovative? incremental product
  • Clariion Disk Library (2004) - was anyone already doing this?
  • DMX-2 Improvements - incremental change
  • EMC VMware - acquisition
  • EMC R&D India - not innovative to open an office
  • EMC Centera - acquisition - FilePool
  • EMC Legato & Documentum - acquisitions
  • Clariion ATA and FC drives
  • EMC DMX (again)
  • EMC ILM - dead
  • EMC Imaging System? Never heard of it
  • IT Hall of Fame - hardly innovation
  • Clariion CX
  • Information Solutions Consulting Group - where are they now?
  • EMC Centera - acquisition
  • Replication Manager & StorageScope - still don't work today.
  • Dell/EMC Alliance - marketing not innovation
  • ECC/OE - still doesn't work right today.
  • Symmetrix Product of the Year - same product again
  • Joe Tucci becomes president - marketing
  • SAN & NAS into single network - what is this?
  • EMC Berkeley study -marketing
  • EMC E-lab
  • Symmetrix 8000 & Clariion FC4700 - same products again
  • EMC/Microsoft alliance - marketing
  • EMC stock of the decade - marketing
  • Joe Tucci - president and COO - marketing
  • EMC & Data General - acquisition
  • ControlCenter SRM
  • EMC Connectrix - from acquisition
  • Software sales rise - how much can be attributed to Symmetrix licences
  • Oracle Global Alliance Partner - marketing
  • EMC PowerPath
  • Symmetrix capacity record
  • EMC in 50 highest performing companies - marketing
  • EMC multiplatform FC systems
  • Timefinder software introduced
  • Company named to business week 50 - marketing
  • EMC - 3TB in an array!!
  • Celerra NAS Gateway
  • Oracle selects Symmetrix - marketing
  • SAP selects Symmetrix - marketing
  • EMC Customer Support Centre Ireland - marketing
  • Symmetrix 1 Quadrillion bytes served - McDonalds of the storage world?
  • EMC acquires McDATA - acquisition
  • EMC tops IBM mainframe storage (Symmetrix)
  • Symmetrix 5100 array
  • EMC 3000 array
  • EMC BusinessWeek top score - marketing
  • Egan named Master Entrepreneur - marketing
  • EMC 5500 - 1TB array
  • EMC joins Fortune 500 - marketing
  • SRDF - innovation - yes.
  • Customer Council - marketing
  • EMC expands Symmetrix
  • EMC acquires Epoch Systems - basis for ECC?
  • EMC acquires Magna Computer Corporation (AS/400)
  • EMC R&D Israel opens - marketing
  • Symmetrix 5500 announced
  • Harmonix for AS/400?
  • EMC ISO9001 certification - marketing
  • Mike Ruettgers named president and CEO - marketing
  • Symmetrix arrays for Unisys
  • Cache tape system for AS/400
  • EMC implements product design and simulation system - marketing
  • Product lineup for Unisys statement - marketing
  • DASD subsystem for AS/400
  • EMC MOSAIC:2000 architecture
  • EMC introduces Symmetrix
  • First storage system warranty protection - marketing
  • EMC Falcon with rapid cache
  • First solid state disk system for Prime (1989)
  • Reuttgers improvement program - marketing
  • First DASD alternative to IBM system
  • Allegro Orion disk subsystems - both solid state (1988)
  • EMC in top 1000 business - marketing
  • EMC joins NYSE - marketing
  • First cached disk controller - innovation - yes
  • Manufacturing expands to Europe - marketing
  • EMC increases presence in Europe and APAC - marketing
  • Archeion introduced for data archiving to optical (1987)
  • More people working on DASD than IBM - marketing
  • EMC introduces solid state disks (1987)
  • Storage capacity increases - marketing
  • EMC doubles in size - marketing
  • Product introductions advance computing power - marketing
  • HP memory upgrades
  • EMC goes public - marketing
  • EMC announces 16MB array for VAX
  • Memory, storage products boost minicomputer performance
  • EMC offers 24 hour support
  • Testing improves quality - marketing
  • Onsite spares program - marketing
  • EMC delivers first product - marketing
  • EMC founded - marketing

Tuesday, 19 August 2008

Off The Grid

I've been on holiday for the last week (sunning myself and the family in Cyprus). I had no Internet access - not even TV! Although I had no laptop (or Blackberry this time) I did take my iPod Touch, now configured with the mobile version of NewsGator. As I've mentioned previously, I have a 100+ RSS feeds (which I'll publish once I get around to it) on storage and others. My backlog was about 2500 entries, so I decided to challenge myself to get up to date and read as much as possible. Clearly I didn't read them all (there were plenty that could be skipped) but I read most and it provides for an interesting cross section...

EMC - blogs are run like a military machine; co-ordinating the news relating to new product releases and mercilessly hammering the competition. EMC have more storage bloggers than any other storage company and there are some good ones out there - one of my favourites is Information Playground by Steve Todd, where he discusses the design of Clariion.

Netapp - follows a close second to EMC with lots of bloggers and lots of competitor bashing. I particularly like Alex McDonald's postings.

IBM - doing a great job running the "resistance", fighting back against the continual onslaught of Barry (A Burke). Check out Barry (Whyte) and Tony Pearson. I'd like to see more from IBM though, especially their product developers working on DS arrays and XIV.

HDS - A jolly good bloke, but not really a player in the blogosphere. Only Hu contributes regularly, but doesn't engage in any serious debate.

Sun - quite literally on another planet with their storage strategy!

Dell - bought some toys, but doesn't know how to play with them. Unfortunately the older boy who could help them play with them has left...

Now there are more companies out there and I don't think I have any blog links from Brocade, 3Par, Compellent, Emulex, Qlogic, Pillar and others although I may be wrong (it is getting late). Any RSS link offerings gladly welcome - although I might not get around to reading them before my next holiday!

Wednesday, 5 September 2007

Invista

There's been a few references to Invista over the last couple of weeks, notably from Barry discussing the "stealth announcement".



I commented on Barry's blog that I felt Invista had been a failure, due to the number of sales. I'm not quite sure why this is so, as I think that virtualisation in the fabric is utimately the right place for the technology. Virtualisation can be implemented at each point in the I/O path - the host, fabric and array (I'll exclude application virtualisation as most storage managers don't manage the application stack). We already see this today; hosts use LVMs to virtualise the LUNs they are presented; Invista virtualises in the fabric; SVC from IBM sits in that middle ground between the fabric and the array and HDS and others enable virtualisation at the array level.



But why do I think fabric is best? Well, host-based virtualisation is dependent on the O/S and LVM version. Issues of support will exist as the HBAs and host software will have supported levels to match the connected arrays. It becomes complex to match multiple O/S, vendor, driver, firmware and fabric levels across many hosts and even more complex when multiple arrays are presented to the same host. For this reason and for issues of manageability, host-based virtualisation is not a scalable option. As an example, migration from an existing to a new array would require work to be completed on every server to add, lay out and migrate data.



Array-based replication provides a convenient stop-gap in the marketplace today. Using HDS's USP as an example, storage can be virtualised through the USP, appearing just as internal storage within the array would. This provides a number of benefits. Firstly driver levels for the external storage are now irrelevant (only requiring USP support, regardless of the connected host), the USP can be used to improve/smooth performance to the external storage, the USP can be used for migration tasks from older hardware and external storage can be used to store lower tiers of data, such as backups or PIT copies.



Array-based replication does have drawbacks; all externalised storage becomes dependent on the virtualising array. This makes replacement potentially complex. To date, HDS have not provided tools to seamlessly migrate away from one USP to another (as far as I am aware). In addition, there's the problem of "all your eggs in one basket"; any issue with the array (e.g. physical intervention like fire, loss of power, microcode bug etc) could result in loss of access to all of your data. Consider the upgrade scenario of moving to a higher level of code; if all data was virtualised through one array, you would want to be darn sure that both the upgrade process and the new code are going to work seamlessly...



The final option is to use fabric-based virtualisation and at the moment this means Invista and SVC. SVC is an interesting one as it isn't an array and it isn't a fabric switch, but it does effectively provide switching capabilities. Although I think SVC is a good product, there are inevitably going to be some drawbacks, most notably those similar issues to array-based virtualisation (Barry/Tony, feel free to correct me if SVC has a non-disruptive replacement path).



Invista uses a "split path" architecture to implement virtualisation. This means SCSI read/write requests are handled directly by the fabric switch, which performs the necessary changes to the fibre channel headers in order to redirect I/O to the underlying target physical device. This is achieved by the switch creating virtual initiators (for the storage to connect to) and virtual targets for the host to be mapped to. Because the virtual devices are implemented within the fabric, if should be possible to make the virtual devices accessible from any other fabric connected switch. This poses the possibility of placing the virtualised storage anywhere within a storage environment and then using the fabric to replicate data (presumably removing the need for SRDF/TrueCopy).

Other SCSI commands which inquire on the status of LUNs are handled by the Invista controller "out of band" by an IP connection from the switch to the controller. Obviously this is a slower access path but not as important in performance terms as the actual read/write activity.

I found a copy of the release notes for Invista 2.0 on Powerlink. Probably about the most significant improvement was that of clustered controllers. Other than that, the 1.0->2.0 upgrade was disappointing.

So why isn't Invista selling? Well, I've never seen any EMC salespeople mention the product never mind pushing it. Perhaps customers just don't get the benefits or expect the technology to be too complex, causing issues of support and making DR an absolute minefield.

If EMC are serious about the product you'd have expected them to be shoving it at us all the time. Maybe Barry could do for Invista what he's been doing in his recent posts for DMX-4?