Showing posts with label disaster recovery. Show all posts
Showing posts with label disaster recovery. Show all posts

Thursday, August 14, 2008

VM infrastructure and disaster recovery

VMWare's ESX/ESXi Update 2 contained what they describe as a "build timeout" which caused patched machines to expire their licenses on August 12th. This meant that VMs could not be powered on or resumed on updated machines, and that VMotion couldn't be used to move VMs to those systems. VMWare has released a patch and a letter to their customers notifying them of the issue, and has flagged the isssue as an alert in their knowledgebase. The fixed patch requires a reboot of the VMWare host, potentially causing off-cycle maintenance to be required for those systems that were affected.

As more infrastructure moves to a VM environment, we create the potential for greater failures when the VM host systems have issues. In this case, a single patch could prevent DR from occurring if all of your VMWare systems were patched and a failure occurred. Workarounds were relatively easy if you knew what was wrong - system date and time could be changed in the short term, or, if necessary, a pre-patch backup could be restored to the system.

How can we best plan to handle issues like this? In many ways, the same processes that system administrators have used for years to test patches will continue to serve us, but we need to have plans in place for what to do when an issue effects all VM hosts at a given patch level. This reminds me of Hoff's talk about VM infrastructure at BlackHat - we're more vulnerable than we think we are with VMs, and this patch issue is a great, relatively low cost reminder.

So - how are you planning to handle VM infrastructure outages?

Friday, August 10, 2007

365Main - an example of great disaster recovery communications

Even if you weren't effected by the 365Main power outage, you should read the status update posted by their president.

There are a few things to note here:

  • The entire event is broken down with technical detail.
  • Details of problem solving, maintenance, and testing are all available and relatively transparent - thus providing customers with detail about the event.
  • The tone is professional and communicates issues and events clearly.
  • Customers are reminded of what recompense their contract provides for them.
  • There is a clearly explained troubleshooting process.
  • There is a clearly explained plan to prevent future issues.
  • The data is being made available to other data centers to help prevent similar issues elsewhere - thus giving back to the community.
Also worth noting is that the status has been regularly updated, and that each update includes current information and future steps.

If you are ever in a recovery situation, this is a great example of after action communication to follow.

On the technical side - remember that simply having backup power may not be enough - many of the customers who had power interrupted would have continued to function if they had dedicated UPS units - but without power to their network uplinks, they might not have been able to see the outside world, even if they had power to the machines themselves.

Thursday, May 3, 2007

The importance of secondary routes

Network World is reporting that a fire caused when a homeless man threw a lit cigarette onto a mattress under a bridge has taken down Internet2 access between Boston and New York.

Yes. A burning mattress took down an important link for a major high speed network. No, this probably wasn't specifically covered in their design and operations risk assessment.

While high speed networks are expensive, this does demonstrate the vulnerability that purpose built dedicated links suffer from. If you only have one link, either because of cost or because of specialization, you need to have plans in place for when it goes down. Copper and fiber aren't invulnerable, and even when you think you're safe you can still get hit. Just when you think it is safe to cross your physical paths, someone will go dig there with a backhoe and cut your fiber.

Two stories come to my mind in which relatively unlikely events threatened or took down Internet access.

In the first, a semi hauling a backhoe went underneath a bridge that was lower than the backhoe's retracted and stored arm. The high speed impact with the bridge severed the fiber running underneath it, cutting off Internet access to a large chunk of Michigan.

In the second, a crew working on a a sewer line hit a gas line. In the process of attempting to fix the gas line, they dug and hit a fiber conduit. Fortunately, the slack in the fiber allowed it to pull just enough to remain operational, but a series of unfortunate events might have resulted in loss of Internet access for a major institution - in addition to a pretty nightmarish repair scenario with fiber and gas lines both broken.

Lessons learned? Always ask about single points of failure, identify alternate routes if possible and financially reasonable, and make sure you have an outage handling and recovery plan.