July 8, 2009

Seattle Database Fire Unnecessarily Shuts Down Businesses and Online Services

Cascade failure. If you’re in IT, that’s a particularly frightening term. In the case of last week’s Seattle data center fire, the term is especially appropriate since it was literally a cascade of water that wrecked everything and sent a number of businesses and online services offline. Here’s a look at this disaster and a way it could have been prevented.

Background: Fisher Plaza, a Major Hosting Facility

Fisher Plaza is “a self-styled carrier hotel in Seattle, and home to multiple datacenter and colocation providers.” [source] A partial list of organizations hosted there includes: payment service provider Authorize.net (which itself has 238,000 merchant customers), Port of Seattle email system, Swedish Hospital’s internal IT systems, Pacific Science Center website, geocaching.com website, major TV and radio station KOMO, online Facebook game Bejeweled Blitz and dozens of other businesses [source].

The Problem: Fire Leads to Cascade Failure

Early on Friday morning, July 3, 2009, Fisher Plaza’s main generator/transfer switch failed. This caused an overload. This caused a fire. This triggered the fire suppression system and brought firefighters to the scene, both of which shot water into the generator room. The generators stopped, and we deduce that power from the grid was shut off too. The UPS and the cooling system also failed. Temperatures in the facility rose high enough to wreck some servers and destroy data [source].

Think about the downstream effects. 238,000 merchants potentially have their transactions interrupted or lost because Authorize.net’s servers are forced offline. One can only hope they had their own functioning backup plan. A hospital’s IT system became unavailable; I have no information on what impact this had on patient care. And apparently KOMO had to transmit from a mobile unit in their parking lot [source]. It is not hard to imagine the impact to these and other organizations.

The Solution: Fire- and Flood-Proof Hosting

No, ZeroNines does not wrap servers in asbestos. There is no way to know what bizarre little accident will happen next, so prevention is unlikely. Some will trigger chain reactions that become major IT disasters.

What we do is to prevent a catastrophe in one place from knocking out a business everyplace. In this case, if any of the clients or tenants at Fisher Plaza had been using our technology, their data, transactions, apps, and other assets would have all been processing simultaneously and in perfect replication in other data centers hundreds or thousands of miles away.

This is not a cutover scenario. Processing would not have “switched” from Seattle to elsewhere. It simply would have stopped in Seattle and continued in real time in San Jose, or Denver, or Singapore, or wherever else they placed their data centers. There would be no loss of business continuity. Their businesses would not have gone down, and the real disaster – lost connectivity, productivity, and revenue – would not have taken place.

Visit the ZeroNines website to find out more about how our disaster-proof architecture protects businesses of any description from downtime.

Alan Gin – Founder & CEO, ZeroNines

June 30, 2009

Uptime and the Cloud Crowd at CSIA

A few days ago, Jake Smith of Intel and I presented at the Colorado Software Industry Association (CSIA) monthly meeting in Denver (source). We talked about cloud computing and the elements that will determine its rate of adoption: the needs of businesses, their expectations of cloud performance, and the real-world limitations of the cloud that are currently stalling its adoption. The biggest issue is reliability, and I introduced ZeroNines’ technology as a potential solution. It was a great crowd, and their hunger for a reliable cloud was obvious.

Businesses need their applications and data to be available all the time. So far, clouds and cloud providers have not succeeded in proving that they can actually offer that. The industry needs to overcome the cloud’s downtime problems before serious business can be done on it. I believe the Big Three (Amazon, Azure, and Google) will refocus their efforts on providing highly available cloud infrastructures and market this capability accordingly.

The Cause is Academic

Of course every network is subject to threats and failures that can cause downtime, and there’s no getting away from that. It doesn’t take an earthquake to knock vital networked apps offline; some recent high-profile cloud provider outages have shown that all it takes is a failed OS upgrade. New and unexpected problems crop up every day. But the cause of an outage is really only academic for the business relying on the cloud. Service should simply continue because the business needs it to.

The scary thing is that the current disaster recovery paradigm (failover) is insufficient for protecting businesses when these things happen, and can’t be relied upon to prevent downtime or even a speedy recovery. In addition, there is an increase in catastrophic risk from poorly architected virtualized environments, and most notably in server consolidation, which is a core technology of the cloud.

The Solution is Continuity

At the CSIA meeting, we introduced the crowd to our Always Available™ technology, which maintains cloud continuity by synchronizing and protecting multiple private, public or hybrid clouds. It can mix cloud computing and physical hosting via datacenters hundreds or thousands of miles apart. The distance prevents any single regional disaster from damaging more than one data center. There is no server hierarchy, so all transactions run simultaneously and equally on all cloud and server nodes. Best of all, they update each other constantly in real time so if one goes down the others simply continue processing with no interruption to service.

To protect against an outage during an upgrade, I would postulate the following solution: Isolate one cloud or network node in an Always Available configuration and do your upgrade there, while the other nodes manage the clients’ transactions. Test the upgrade and slowly roll it out to the other nodes. If things start to go haywire, isolate the misbehaving node, solve your problems, and start the rollout again. There would be no need to risk the entire service on an untested upgrade.

Always Available works for cloud customers as well as service providers. It is provider- and platform-agnostic, so you can mix and match all you need to.

Visit the ZeroNines website to find out more about how our disaster-proof architecture protects businesses of any description from downtime.

Alan Gin – Founder & CEO, ZeroNines

January 12, 2009

Hurricane Charley Couldn’t Stop the Email

In this first Disaster Litany Posting, I look at a sequence of events that is near and dear to ZeroNines. Our own real-world experience with a hurricane, power outages, and an email system will show just how our downtime-preventing technology works, and serve as a pattern for the solutions we suggest for other disasters.

Background: MyFailSafe™ Email System

ZeroNines offers Always Available technology that can virtually eliminate downtime among networked applications, data, and other assets. To test our technology, we created the MyFailSafe Email Service and launched it on our Always Available network in July of 2004. This was specifically intended to test Always Available in the real world, by running MyFailSafe just like any other email service is run, with real customers and real traffic, and subject to the same threats that any other network or email system is vulnerable to.

The Problem: A Hurricane

All readers who remember Hurricane Charley please raise your hands… For those of you who don’t, Charley hit Florida on August 13, 2004. According to Wikipedia, it killed about thirty people and caused $15 billion in damage. Widespread flooding, wind damage, power outages, and other problems crippled much of the state for several days. I don’t have statistics on downtime among private business networks or service providers, but it’s a safe bet that it was serious.

Charley hit about a month after we launched MyFailSafe. It caused electrical grid fluctuations that drained the Orlando local exchange carrier battery backup systems, isolating the Orlando node of the ZeroNines Always Available infrastructure. Our own battery system prevailed and still had a 75% charge when commercial power was reliably restored, but the site could not communicate for 16 hours because of LEC downtime.

The Solution: Hurricane-Proof Architecture

During this 16 hours, when our Orlando node was effectively offline, the MyFailSafe email service did not experience any downtime at all. Any user whose power was still on and whose desk was not under water experienced true 100% uptime throughout, whether they were in Florida, Colorado, Canada, Asia, or anywhere else.

How? Our Always Available deployment has additional nodes and data centers in Colorado and California. All applications, transactions, data exchanges, and other network activities run equally and simultaneously on these multiple secure application servers, geographically separated by hundreds of miles. In IT parlance, all servers are hot, and all instances of all applications are active. There is no server hierarchy, and consequently no single point of failure. When the Orlando node fell silent, all MyFailSafe processing continued uninterrupted on the others. There was no need for failover or recovery because these other nodes were far from the storm, they never went down, and continuity was maintained.

Since activation on July 15, 2004, the MyFailSafe network has never experienced any downtime for any reason, including this and other hurricanes, two migrations from server collocation providers to clouds, a data center move, and an email worm attack that interrupted email service from AOL and other major providers. These potential disasters, which forced our servers offline, had no power to bring our applications down. All applications and information retained 100% availability throughout.

Contact ZeroNines to find out more about how our disaster-proof architecture protects businesses of any description from downtime.

Alan Gin – Founder & CEO, ZeroNines

January 3, 2009

A Litany of Disasters: Downtime Events and How to Avoid Them

“Aviation in itself is not inherently dangerous. But to an even greater degree than the sea, it is terribly unforgiving of any carelessness, incapacity, or neglect."
-- Anonymous

Years ago, I saw those words on a poster of a World War One aircraft stuck about 20 feet off the ground in the limbs of a tree. If we were to update this and adapt it to the business user’s desktop, it would lose its poetic charm but strike home with a whole new audience:

“Networked assets in themselves are not inherently dangerous. But to an even greater degree than stuff on your hard drive, they are terribly unforgiving of any carelessness, incapacity, or neglect."

The warning is clear: Disaster may be only inches away, particularly for the unprepared. It’s a lot harder to recover after some accident knocks out a hundred or a thousand users than it is to re-boot your own machine.

In this blog, we will be looking at some actual disasters that have struck organizations when their networks have taken a hit from storms, fires, attacks, and far more mundane threats like human error and equipment failure.

For a business, there may be little correlation between the physical effects of a disaster and its financial impact. Imagine a business dependent upon a distant data center in the U.S. Tornado Belt. One good storm could leave their personnel and property untouched, yet destroy their ability to do business by wiping out their data, applications, and transactions. Elsewhere, an earthquake could cause deplorable loss of life and property damage, yet leave a business relatively unharmed if its networked computing capabilities remain intact. And an otherwise strong corporation could suffer irreparable damage by something as quiet as a software failure or equipment malfunction, which to the outside world does not qualify as a “disaster” at all.

I’ll be describing some instances where ZeroNines’ solutions for networks, virtualized environments, and clouds could have prevented disastrous downtime, and helped avoid unwanted headlines and losses to productivity, reputation, and revenue. Our approach does not use any kind of failover or cutover, since those occur after the downtime event and are not true disaster prevention. After all, it’s far better to avoid the downtime in the first place than to try to recover from it afterward.

Next week: How MyFailSafe really did provide fail-safe email during Hurricane Charley.

Contact ZeroNines to find out more about how our disaster-proof architecture protects businesses of any description from downtime.

Alan Gin – ZeroNines, Founder & CEO

September 15, 2008

Cloud Failure: The Myth of Nines

Visit Reuven Cohen's blog

For about as long as there have been computer networks, administrators have attempted to keep these networks up and running. It seems to be a continuous battle between faulty hardware, poorly written software, unreliable connectivity and random acts of God. With the emergence of cloud computing we are now for the first time close to realizing a computing environment where we are able to focus less on keeping our applications up and more on making them run more efficiently and effectively.

In the era of cloud computing uptime guarantees and service level agreements (SLA) have started to become standard requirements for most cloud providers. Google, Amazon, and Microsoft have all started to implement some kind of SLA. They do this in an attempt to give their cloud users the confidence to utilize these systems in place of more common in house alternatives. The common goal for most of these cloud platform is to build for what I consider the myth of five nines. (Five nines meaning 99.999% availability, which translates to a total downtime of approximately five minutes and fifteen seconds per year.) The problem with five nines is it's a meaningless goal which can be manipulated to meet what ever you need it to mean.

In the case of a physical failure such as Flexiscales recent one, the hardware downtime might be small, but the time to restore from a backup might be considerably longer. A minor cloud failure could cause a cascading series of software failures causing further application outage of hours or even days for those who depended on the availability of the given cloud. Meaning your cloud may achive five nines, but your application hosted on it doesn't.

Lately it seems there are a number of people in the cloud computing community who are starting to discuss alternatives to the dreaded five nines concept and looking at ways that cloud based infrastructures could be configured / deployed in a mannor that is more proactive than reactive to disasters. There is a growing consensus that cloud based disaster recovery may very well be the "killer app" for cloud computing. To achieve this, we need to start creating reference architectures and models that assume for failure. One that doesn't need to worry when the next disaster will happen next, just that it will happen and when it does, it's going to be business as usual.

In a recent conversation with Alan Gin founder of a super secret stealth firm called Zeronines, Alan described an interesting philosophy. He said the problem with most disaster recovery plans is the recovery is reactive, it is what happens after a disaster has already harmed your business. He said on its face, this is an unsound strategy. He went on to say; That current disaster recovery architectures, which uses the synonym “failover,” is based on the cutover archetype: a system’s primary component fails, damaging operations; then failover to a secondary component is attempted to resume operations. The problem with current cutover approaches is that it views unplanned downtime as inevitable, acceptable, and so requires that business halt.

I really liked this quote from an executive from EMC, a leading computer storage equipment firm, “current failover infrastructures are failures waiting to happen.”

To be competitive in today's always connected, always available world. We need to reinvent the fundamental idea of disaster recovery. One of the major benefits to using cloud computing is that you can make these types of failover assumptions well before they happen using an emerging global toolset of cloud components. It's not a matter of if, but a matter of when, when you take into consideration that application components will fail then you can build an application that features "failure as service". One that is always available, one with Zero Nines.

Reuven CohenFounder & chief technologist for Toronto based Enomaly.