The Cloud and Outages : Five Key Lessons for Customers(cloudsigma.com)
cloudsigma.com
The Cloud and Outages : Five Key Lessons for Customers
http://www.cloudsigma.com/en/blog/2011/04/23/21-cloud-outages-lessons-learned
2 comments
Yes, it's a pretty transparent advert saying "don't be afraid of the cloud, just their cloud"
"Lesson 2: Size is No Protection from Outages"
At this point it has become clear that EBS failed due to a cascading failure as EBS volumes were mirrored across datacenters. http://joyeur.com/2011/04/22/on-cascading-failures-and-amazo...
From a practical point of view, the solution is simple, allow EBS to fail in certain areas to cordon off the damage. Size was a key variable here that made the outage take so long to recover from.
On point 1, in my opinion, this was not a single point of failure. Thus the need for the adjective "Cascading" That is, it was no more a single point of failure as relying on "hard drives" would be. ie. the "sigle point" was the technology itself, not a physical single point, obviously.
So in my mind, the question is, can the technology be modified to allow for failures or to restrict those failures while still utilizing the huge pool of hardware available to it effectively? It's a difficult problem, but not unsolvable.
At this point it has become clear that EBS failed due to a cascading failure as EBS volumes were mirrored across datacenters. http://joyeur.com/2011/04/22/on-cascading-failures-and-amazo...
From a practical point of view, the solution is simple, allow EBS to fail in certain areas to cordon off the damage. Size was a key variable here that made the outage take so long to recover from.
On point 1, in my opinion, this was not a single point of failure. Thus the need for the adjective "Cascading" That is, it was no more a single point of failure as relying on "hard drives" would be. ie. the "sigle point" was the technology itself, not a physical single point, obviously.
So in my mind, the question is, can the technology be modified to allow for failures or to restrict those failures while still utilizing the huge pool of hardware available to it effectively? It's a difficult problem, but not unsolvable.
Content-free PR FUD. Flagged.