Prime Down: Amazon’s sale day turns into fail day(techcrunch.com)
techcrunch.com
Prime Down: Amazon’s sale day turns into fail day
https://techcrunch.com/2018/07/16/prime-down-amazons-sale-day-turns-into-fail-day/
360 comments
Something no one has mentioned yet, could it be that the engineering force at Amazon is no longer what it used to be?
I can personally point to two friends who I consider top notch engineers and designers that have left Amazon because of its toxic culture. I'm sure I'm not the only with these anecdotal examples, we've all heard the stories. At the end of the day years of unbalanced work/life balance, overly aggressive management and frugal approach to everything makes for a weak argument for A players to stick around.
Could this be an example of crumbling engineering standard at Amazon?
I can personally point to two friends who I consider top notch engineers and designers that have left Amazon because of its toxic culture. I'm sure I'm not the only with these anecdotal examples, we've all heard the stories. At the end of the day years of unbalanced work/life balance, overly aggressive management and frugal approach to everything makes for a weak argument for A players to stick around.
Could this be an example of crumbling engineering standard at Amazon?
If only they had access to some kind of scalable cloud hosting service, they could've completely avoided this sort of outage. :)
Jokes aside, I admire the work of the team(s) responsible for Amazon's web site. I use it so often and encounter glitches so rarely that it really stands out when something does go wrong.
Jokes aside, I admire the work of the team(s) responsible for Amazon's web site. I use it so often and encounter glitches so rarely that it really stands out when something does go wrong.
Funny how this comes on the heels of aggressively expanding their workforce and trying to leverage themselves in a hundred different directions...
Maybe it's just me and my confirmation bias at work, but it seems that the core value proposition that Amazon provided -- high value, low margins on products -- has been eroding before our eyes.
Seems so much like the transition Microsoft made... too much focus on "synergies" and leveraging... not enough on keeping the bilges dry and the engine running.
It's funny... Fred Brooks wrote about this in 1975... and we're still making the same mistakes forty years later. There are real limitations to how quickly any organization can grow. Even awesome companies who are excellent at building organizations -- places like like Amazon and Microsoft -- can't organize this law of software development away.
Maybe it's just me and my confirmation bias at work, but it seems that the core value proposition that Amazon provided -- high value, low margins on products -- has been eroding before our eyes.
Seems so much like the transition Microsoft made... too much focus on "synergies" and leveraging... not enough on keeping the bilges dry and the engine running.
It's funny... Fred Brooks wrote about this in 1975... and we're still making the same mistakes forty years later. There are real limitations to how quickly any organization can grow. Even awesome companies who are excellent at building organizations -- places like like Amazon and Microsoft -- can't organize this law of software development away.
I wonder how Alibaba Cloud handles similar events [1], where there are bursts of 256k/s transactions and ~1bil packages being shipped out.
Do they just do brute-force massive scale out?
Amazon's US market is big, but my understanding is that number of online users in China (> 400mil) exceeds the population of the US (~325mil), which makes me wonder if the folks there think about data architecture a little differently than we do.
[1] https://qz.com/1127087/singles-day-crazy-stats-from-alibabas...
Do they just do brute-force massive scale out?
Amazon's US market is big, but my understanding is that number of online users in China (> 400mil) exceeds the population of the US (~325mil), which makes me wonder if the folks there think about data architecture a little differently than we do.
[1] https://qz.com/1127087/singles-day-crazy-stats-from-alibabas...
I actually managed to add an item on sale to my cart, but my cart is now empty after refreshing. Oh well, maybe next year.
I can't help but feel terrible for the team of people there that ultimately gets blamed for this, I hope they can get some sleep tonight.
I can't login to my root aws account right now. It's pretty annoying that the root account login for aws is tied to amazon.com retail accounts.
I work in a non-retail part of Amazon and I'm on vacation. Hasn't stopped friends and family from texting me about this though. As if I personally can go in and reboot a server or something. Hope we get it sorted soon!
If you're affected by this, please accept my unofficial thanks for your patience and understanding. (If you're a coworker in retail, good luck getting things up and running!) :-)
If you're affected by this, please accept my unofficial thanks for your patience and understanding. (If you're a coworker in retail, good luck getting things up and running!) :-)
Help strikers in Europe by NOT purchasing from Amazon today. Thank you.
It's also impossible to log into the AWS console. Surely something of such importance should be separate to the ecommerce site.
I know TFA is about technical failures, but the deals themselves are also incredibly lacklustre. I was expecting at least the Warehouse Deals part of Prime Day to come through, typically 15 to 20% off all used offerings. This year however, Amazon restricted it to only select listings which translated to a few hundred items total. Very sad.
What about the advertising money already spent?
I read somewhere (sorry, forgot where) that Amazon had been pushing sellers to spend like mad on ads within Amazon.com for Prime Day, apparently it gave you a big edge over whatever the algorithms suggest.
Those sellers will have missed their sales targets, and will consider the ad spend to have been wasted. Will they get it back?
I read somewhere (sorry, forgot where) that Amazon had been pushing sellers to spend like mad on ads within Amazon.com for Prime Day, apparently it gave you a big edge over whatever the algorithms suggest.
Those sellers will have missed their sales targets, and will consider the ad spend to have been wasted. Will they get it back?
... Is there anything even good for Prime Day this year? Or the year before? Two years ago I remember seeing at least some Dell Workstations that could be repurposed into cheapo home servers. Most of the stuff seems to be odds, ends, and stuff that the various Chinese product-clone companies couldn't get rid of.
Is there enough historical, public data available to estimate the amount of money Amazon is losing per second?
The web site is a bit of a mess. And it seems that, just since this morning, my entire Wish List has vanished. Foolish me for not having a backup...
I'm seeing automatic reloading the page every second or so. Maybe some bad javascript, though it isn't adding entries to the history. Looks like they have a script that is DDoS themselves.
What happens to people responsible for the crash today (infra, culprit services)? Does Amazon take some kind of "action" since Prime Day is a huge, once-a-year event for Amazon?
Well, everything I wanted to buy is simply not on sale. If Prime Day isn't supported by those things that I want or need then why would I participate?
changing the url to smile.amazon.com works
How about some love for their marketing guy who made it all happen?
Jeff probably said "make it rain, let's see if your hordes can take down Amazon.com," and this guy basically accepted, and succeeded at, the challenge.
Jeff probably said "make it rain, let's see if your hordes can take down Amazon.com," and this guy basically accepted, and succeeded at, the challenge.
Huh? Amazon seems OK now.
Edit: But hmmm, Quora just went down.
> 504. Gateway Timeout.
> Quora is temporarily unavailable.
> Please wait a few minutes and try again.
And they use AWS, right?
Edit: I just got as far as search, order, cart. But no account as Mirimir, so ...
Edit: Re Quora - http://downdetector.com/status/quora
I wonder what other AWS stuff is down. If that's it, anyway.
Edit: But hmmm, Quora just went down.
> 504. Gateway Timeout.
> Quora is temporarily unavailable.
> Please wait a few minutes and try again.
And they use AWS, right?
Edit: I just got as far as search, order, cart. But no account as Mirimir, so ...
Edit: Re Quora - http://downdetector.com/status/quora
I wonder what other AWS stuff is down. If that's it, anyway.
[deleted]
Industrial action is cool and good
https://twitter.com/kadybat/status/1018926864767676416
I’m curious what the reprocusions are, if any, once a post mortem is completed and teams or individuals that contributed to the outage are identified. Is “causing” this a fireable offense?
FWIW, it worked fine in Western Europe since 1 GMT, 8 hours ago.
This is the page everyone seems to be getting
https://i.imgur.com/vpIHDpA.jpg?1
https://i.imgur.com/vpIHDpA.jpg?1
It's up, aaaaannnnddddd now it's back down...
It probably doesn't help that the mobile app appears to load a random picture of a cute dog every time I press the "retry" button. So you can guess what I'm doing, trying to get it to load a new "Dogs of Amazon" pic.
Probably should have gone with goatse, reduce the load.
EDIT: do NOT search for "goatse" on your work connection. That alone, even if you've never heard the word, should tell you why I suggested it as an alternative.
Probably should have gone with goatse, reduce the load.
EDIT: do NOT search for "goatse" on your work connection. That alone, even if you've never heard the word, should tell you why I suggested it as an alternative.
[deleted]
My bet's on some unforeseen bottleneck that affects search and static pages. Almost everything within Amazon is crazy scaleable, but there are some bits where you scale them up and their behaviour changes radically. For instance, a service's cache misses might skyrocket as customers get distributed over a wider set of servers, causing service response times to increase just a little bit on average, tipping a dependent service over into more frequent timeouts, causing its downstream service to blow a timeout-percentage 'software fuse' and stop using that service... etc etc.
Given that each of those services (and many more possibly-related ones) will have an on-call engineer paged into a conference call when the manure hit the rotating ventilation apparatus, there are going to be a lot of unhappy people cancelling their weekend plans right now. I definitely don't miss that aspect of the job!