Correct me if I'm wrong, but you would still need to pay $3 USD/mo for Fastmail even if you use 1P. Whereas with Relay, it's 0.99 USD/mo, and no need to migrate my existing email to any other service.
The problem is, I don't have access to the email used to create this account. I created this account long time back, when hotmail was a thing. I haven't logged into that email account since 2004. I tried getting access to it, but wasn't able to do so. So basically, this account is in a weird state where it exists, but I don't have access to it.
I understand it. I've worked in AWS, and now in OCI, dealing with systems that affect hundreds-to-thousands of customers, which businesses are at stake.
Mitigation is your top-priority. Bringing the system back to a good shape.
If there needs to be follow-up actions, take the less-impactful steps to prevent another wave.
If there was a deployment, roll-back.
My concern here is, a deployment have been made months ago, and many other changes that could make things worse were introduced. This is the case. The difference between taking an extra 10-20 minutes to make sure everything is fine, versus taking a hot call and causing another outage makes a big difference.
I'm just asking questions based on the documentation provided; I do not have more insights.
I am happy Stripe is being open about the issue, that way many the industry learns and matures regarding software-caused outages. Cloudflare's outage documentation is really good as well.
Thank you for taking the time to respond to my questions.
I believe the high potential of causing a follow-up incident was left out of the post (or maybe I missed it?).
I hope that lessons are learned from this operational event, and invest towards building metrics and tooling that allows you to, first of all, prevent issues, and second, shorten the outage/mitigation times in the future.
I'm happy you guys are being open about the issue, and taking feedback from people outside your company. I definitely applaud this.
[2019-07-10 20:13 UTC] During our investigation into the root cause of the first event, we identified a code path likely causing the bug in a new minor version of the database’s election protocol.
[2019-07-10 20:42 UTC] We rolled back to a previous minor version of the election protocol and monitored the rollout.
There's a 20 minute gap between investigation and "rollback". Why did they rollback if the service was back to normal? How can they decide, and document the change within 20 minutes? Are they using CMs to document changes in production? Were there enough engineers involved in the decision? Clearly all variables were not considered.
To me, this demonstrates poor Operational Excellence values. Your first goal is to mitigate the problem. Then, you need to analyze, understand, and document the root cause. Rolling-back was a poor decision, imo.
I currently work for Oracle Cloud Infrastructure; ex AWS for Commerce Platform and Identity organizations. Let me tell you, OCI has a group of brilliant, industry mature developers and people. I've been working here for 2 years, and I've never been happier before. This org is nothing like Oracle Corp; started by ex-AWS/MSFT people, the environment feels just like those companies + a well funded start-up hype.
And the best thing, no assholes and backstabbing like in AWS. :)
I've thought about this problem, specially in LATAM countries were gov agencies are still ages behind in terms of security/compliance and sharing these type of documents. There is definitely a need for this.
I agree. But in my experience, even when VPs expose a vision of their culture, directors/GMs are the ones who end up exercising their vision/culture to their teams.
1. It depends on the company (Amazon vs AWS) and it varies from org to org, and team to team. Just like in every place there's going to be great and crapy people. In 3 years at Amazon I had the opportunity to work next to great people, but also with people at the complete opposite side.
2. In my case, it did not. I decided my happiness/wellness and health (physical and mental) were more important.
> "Daniel Imberman, a 26-year-old software engineer, drifts toward this pole, though he rejects the FIRE label (“sounds like classic tech-douche”). His target is $15 million."
How long do you think it'll take him to reach his goal? He talks about starting a start-up, selling or profiting, but all of that take a lot of work. And even then there's no guarantee his plan will work.
Saving 15M will probably take a developer anywhere from 10-20 years? Or am I very wrong?
I know this guy, he lives, or used to live in the same apartment complex as I did, and I would see him almost every day at the bus stop. He is a real warrior. Whenever the Front seats were busy he will encourage people to stay seated and he would find an empty seat.
It's very nice to see Amazon wrote the article about him!
I definitely agree with this, however, most people moved away from verbal communication because it's an async process, and it removes any awkwardness from the interaction. Video conferencing is used occasionally. I think text-based communication is here to stay.