Amazon Web Services are down(status.aws.amazon.com)
status.aws.amazon.com
Amazon Web Services are down
http://status.aws.amazon.com/?a
332 comments
Amazon's EC2 SLA is extremely clear - a given region has an availability of 99.95%. If you're running a website and you haven't deployed across across more than one region then, by definition, your website will have 99.95% availailbility. If you want a higher level of availability use more than one region.
Amazon's EBS SLA is less clear, but they state that they expect an annual failure rate of 0.1-0.5%, compared to commodity hard-drive failure rates of 4%. Hence, if you wanted a higher level of data availability you'd use more than one EBS volume in different regions.
These outages are affecting North America, and not Europe and Asia Pacific. That's it. Why is this even news? Were you expecting 100% availability?
Amazon's EBS SLA is less clear, but they state that they expect an annual failure rate of 0.1-0.5%, compared to commodity hard-drive failure rates of 4%. Hence, if you wanted a higher level of data availability you'd use more than one EBS volume in different regions.
These outages are affecting North America, and not Europe and Asia Pacific. That's it. Why is this even news? Were you expecting 100% availability?
4/21/2011 is "Judgement Day" when Skynet becomes self aware and tries to kill us all. http://terminator.wikia.com/wiki/2011/04/21
I am just a little freaked out right now.
I am just a little freaked out right now.
A couple of hours into the failure, and no sign of coverage on Techcrunch (they're posting "business" stories though). It shows how detached Techcrunch has become from the startup world.
Edit: I tweeted their European editor about it and he's posted a story up now.
Edit: I tweeted their European editor about it and he's posted a story up now.
This feels the same way as hearing that the whole Internet just got shut down.
I guess this is one Reddit outage that can't be blamed on poor scaling
Why is ELB not mentioned at all on the Service Health Dashboard?
We're experiencing problems with two of our ELBs, one indicating instance health as out of service, reporting "a transient error occurred". Another, new LB (what we hoped would replace the first problematic LB), reports: "instance registration is still in progress".
A support issue with Amazon indicated that it was related to the ongoing issues and to monitor the Service Health Dashboard. But, as I mentioned before, ELB isn't mentioned at all.
We're experiencing problems with two of our ELBs, one indicating instance health as out of service, reporting "a transient error occurred". Another, new LB (what we hoped would replace the first problematic LB), reports: "instance registration is still in progress".
A support issue with Amazon indicated that it was related to the ongoing issues and to monitor the Service Health Dashboard. But, as I mentioned before, ELB isn't mentioned at all.
Quora says: "We'd point fingers, but we wouldn't be where we are today without EC2."
I just launched a site on Heroku yesterday and cranked up the dynos up in anticipation of some "launch" traffic. Now, I can't log in to switch them off. Thanks EC2, you owe me $$$s
I think this is a good example of how the "cloud" is not a silver bullet to making your site always up. AWS provides a way to keep it up, but it is up to each developer to ensure that they are using AWS in a way to make sure their site can handle problems in one availability zone.
I think we will see more of a focus from big users of AWS about focusing on how to create a redundant service using AWS. Or at least I hope we will!
I think we will see more of a focus from big users of AWS about focusing on how to create a redundant service using AWS. Or at least I hope we will!
Instead of enumerating who's down, I'd be more interested to hear about those that survived the AWS failure. We could learn something from them.
Quora is down, and evidently "They're not pointing fingers at EC2" --
http://news.ycombinator.com/item?id=2470119 -- I was going to post a screen shot, but evidently my Dropbox is down too.
Holy crap. An Amazon rep actually just posted that SkyNet had nothing to do with the outage:
https://forums.aws.amazon.com/message.jspa?messageID=238872#...
https://forums.aws.amazon.com/message.jspa?messageID=238872#...
I'm seeing 1 EBS server out of 9 having issues (5 in one availability zone, 4 in another). CPU wait time on the instance is stuck at 100% on all cores since the disk isn't responding. Sounds like others are having much more trouble.
Silver lining: Hopefully I can test my "aws is failing" fallback code. (my GAE based site keeps a state log on S3 for the day when GAE falls in a hole.)
AWS/S3 has become the new Windows - great SPOF to go for if you want to attack. This space needs more competition.
http://venuetastic.com/ - feel bad for these guys. They launched yesterday and down today because of AWS. Murphy's law in practice.
So when big sites deal use Amazon Web Services for major traffic, do they get a serious customer relationship? Or is it just generic email/web support and a status page?
It's a bit ironic that Amazon WS has become a SPoF for half the internet.
Yes, they are. :(
Assuming the problem is indeed with EBS, I would say this should be a warning sign to anyone considering going with a PaaS provider, which Amazon is quickly becoming, instead of an IaaS provider like Slicehost or Linode.
The increased complexity of their offering makes it more likely that things will break, leaving you locked in.
I did a 15 minute talk on the subject, which you can check out here: http://iforum.com.ua/video-2011-tech-podsechin
EDIT: here are the slides if you can't bother watching the video http://bit.ly/eqDNei
The increased complexity of their offering makes it more likely that things will break, leaving you locked in.
I did a 15 minute talk on the subject, which you can check out here: http://iforum.com.ua/video-2011-tech-podsechin
EDIT: here are the slides if you can't bother watching the video http://bit.ly/eqDNei
From EngineYard: "It looks like EBS IO in the us-east-1 region is not working ideally at this point. That means all /data and /db Volumes which use EBS have bad IO performance, which can cause your sites to go down."
They better start writing their explanation now. Multiple AZ's affected?
Had our blog go down. Didn't realize it was AWS wide..did a reboot. Now I am in reboot limbo. Put an urgent ticket into Amazon. They just said they are working urgently to fix the issues. Let's see how long this goes.
In case anyone is late to the party and missed the non-green lights on the AWS status dashboard, here is the page as of about 9:30 EDT...
http://screencast.com/t/p69xAoDJRSer
http://screencast.com/t/p69xAoDJRSer
Given that Heroku's parent company (Salesforce) owns a cloud platform, it seems kinda inevitable now that Herkou will perhaps sooner-than-later switch back-ends (or at least use both)
Everyone talks about SLAs but I believe it doesn't consider the fact that the EBS vols are still up (not on fire, and available) and are phantom writing or that the network is queued up the wazoo so writes don't even happen in a timely manner as you'd expect.
So do we get some credit on our AWS accounts? I haven't really read their SLA for EC2.
Being unable to get much done here, my co-workers have found other things to do in the office: http://www.youtube.com/watch?v=u1-oGxDHQbI :-P
"Netflix showed some increased latency, internal alarms went off but hasn't had a service outage." [1]
"Netflix is deployed in three zones, sized to lose one and keep going. Cheaper than cost of being down." [2]
[1] https://twitter.com/adrianco/status/61075904847282177
[2] https://twitter.com/adrianco/status/61076362680745984