Google Cloud Partial Outage(status.cloud.google.com)
status.cloud.google.com
Google Cloud Partial Outage
https://status.cloud.google.com/incident/zall/20003
13 comments
Is GCP actually less reliable than AWS or Azure, or is that only HN effect where everything bad about Google gets voted to the top?
I'm hearing a lot of talk from my contacts about AWS just simply denying and hiding any outages until they're undeniably proven - so perhaps the issue is just the fact that GCP is too honest?
I'm hearing a lot of talk from my contacts about AWS just simply denying and hiding any outages until they're undeniably proven - so perhaps the issue is just the fact that GCP is too honest?
While totally unscientific, my personal experience would suggest that GCP is less reliable than AWS. Our team has run into a few bizarre GCP networking issues that we've managed to work around, but I've still got some intermittent problems that I haven't been able to isolate or document for a ticket.
I have encountered one instance of "AWS just simply denying and hiding any outages until they're undeniably proven" so far this year, so I suspect it does happen. But overall, I do feel like I've had less unexpected behavior or service interruptions on AWS.
Although this is sort of like talking about which telco carrier is best, I suspect many have had a bad experience, and there are some who have never noticed a problem with whichever service they're using.
I have encountered one instance of "AWS just simply denying and hiding any outages until they're undeniably proven" so far this year, so I suspect it does happen. But overall, I do feel like I've had less unexpected behavior or service interruptions on AWS.
Although this is sort of like talking about which telco carrier is best, I suspect many have had a bad experience, and there are some who have never noticed a problem with whichever service they're using.
> I have encountered one instance of "AWS just simply denying and hiding any outages until they're undeniably proven" so far this year, so I suspect it does happen. But overall, I do feel like I've had less unexpected behavior or service interruptions on AWS.
Fewer spurious niche "affects only you" problems per customer, times far more customers, may still equal a greater total number of spurious niche problems, though.
For an awkward analogy, it's like being friends with a con artist. They might have conned you once, or maybe even never. But they might also have 1000 other friends that they've conned zero or one times. It may add up to a lot more conning than the average person gets up to!
Fewer spurious niche "affects only you" problems per customer, times far more customers, may still equal a greater total number of spurious niche problems, though.
For an awkward analogy, it's like being friends with a con artist. They might have conned you once, or maybe even never. But they might also have 1000 other friends that they've conned zero or one times. It may add up to a lot more conning than the average person gets up to!
Most evaluations from 3p sources like Gartner put Amazon a nose ahead in a virtual tie with Google regarding uptime and reliability with Azure pretty far behind. The latest numbers I saw were for 2018:
AWS - 99.9987
GCP - 99.9982
Azure - 99.9792
Source: https://www.geekwire.com/2019/microsoft-may-cloud-computing-...
There's a huge anti-Google groupthink here and negative news about Microsoft seems to be pretty muted.
AWS - 99.9987
GCP - 99.9982
Azure - 99.9792
Source: https://www.geekwire.com/2019/microsoft-may-cloud-computing-...
There's a huge anti-Google groupthink here and negative news about Microsoft seems to be pretty muted.
Curious about this as well.
Google hate here can be palpable
Wind back 5 years or more and this place was a Google love fest.
Somehow it seems that Google has managed to disappoint a lot of people.
Somehow it seems that Google has managed to disappoint a lot of people.
I am afraid you would need to rewind further than that.
I mean... they have. Remember when they pledged to "don't be evil"? Remember all of their cool initiatives that crashed and burned and were killed quickly?
> Somehow it seems that Google has managed to disappoint a lot of people.
Yeah, and those people aren't wrong - it's not like Google didn't do their own share of messups.
Still, the upplaying of Googles' faults and downplaying of other companies faults (even when they're the same!) has not been the best outlook of this community lately.
Yeah, and those people aren't wrong - it's not like Google didn't do their own share of messups.
Still, the upplaying of Googles' faults and downplaying of other companies faults (even when they're the same!) has not been the best outlook of this community lately.
[deleted]
One thing to consider with AWS is that a lot of the services are designed with multiple levels of isolation and independence beyond the region/availability zones. So a service that fails even 100 % for one customer may not fail even once for another. In such cases it is possible they do not post a general availability drop for the entire service since it may affect only a small number of customers. It may be posted to the specific customer’s personal health dashboard. It took years of operating such systems for Amazon to understand deeply how large scale systems can fail and incorporate those into the designs and even then it is an ongoing effort to make them ever more reliable. Now I don’t know how Google’s services are designed and if they don’t have same or higher layers of isolation then it is possible a failure affects most customers in a region thereby requiring a global status update.
Disclaimer AWS engineer.
Disclaimer AWS engineer.
Experience may differ but my first app with some modifications still working good, I think since 7 years.
GCP marketing needs to get on the ball and make AWS gaslighting a meme. It's very real, and it kills me to see them get away with it again, and again, and again.
Oh my! GCP is on fire?
https://status.cloud.google.com/incident/compute
It just looks horrific from the reliability standpoint. May be Google should consider adding columns for affected regions! Maybe they are too honest in admitting the failures than other providers, but that page paints a very bad picture.
It just looks horrific from the reliability standpoint. May be Google should consider adding columns for affected regions! Maybe they are too honest in admitting the failures than other providers, but that page paints a very bad picture.
I'm glad they do. Compared to AWS this transparancy is a breath of fresh air.
And I'm no Google fan either.
And I'm no Google fan either.
I like how I found out about this here and not through an email. Tons of customer reports from our site of being logged out randomly (we use Firebase) and we had no idea why.
The outage is going on for almost 12-hours. Quick math says that's 99.86% reliability.
On top of this, they want to start charging $70/mo for GKE because they "[...] are also introducing a Service Level Agreement (SLA) that's financially backed with a guaranteed availability of 99.95%" (pulled from pricing[0])
This is unacceptable.
[0] https://cloud.google.com/kubernetes-engine/pricing
On top of this, they want to start charging $70/mo for GKE because they "[...] are also introducing a Service Level Agreement (SLA) that's financially backed with a guaranteed availability of 99.95%" (pulled from pricing[0])
This is unacceptable.
[0] https://cloud.google.com/kubernetes-engine/pricing
The current/remaining issue is "Affected customers may experience delayed IAM modifications that surface across multiple Google Cloud Platform services." In simple terms, user login/auth is working fine, it's just modifications (new service accounts, access changes, etc) that are delayed. Obviously not great, but this doesn't really affect the critical path of any systems that are already up and running.
I'm guessing because we are going through a pandemic they might be taking more time than usual to apply fixes.
[deleted]
Already looking forward to that post-mortem. It might be morbid, but they are often my favorite things to read.
I guess the number of comments here reflect how many people are using Google Cloud - apparently not much?
There's nothing really to say about it.
I run about 8,000 CPU cores in GCP and what am I going to say?
I like the transparency, in my experience it's more reliable than AWS- although AWS does _not_ report problems until they're absolutely undeniable.
Other than that, shit happens, this is upsetting but eh.
I run about 8,000 CPU cores in GCP and what am I going to say?
I like the transparency, in my experience it's more reliable than AWS- although AWS does _not_ report problems until they're absolutely undeniable.
Other than that, shit happens, this is upsetting but eh.
Google Cloud was the least reliable major cloud provider in the last 2 years. May be they are too busy killing their other services, so they don't have time for other least important things like cloud.
Is this based on independent research, or officially-reported downtime? There are enough anecdotes in threads here about cloud reliability to suggest AWS is bad at reporting outages.
This is a disaster...
> We are currently investigating an issue affecting Dataflow, BigQuery, DialogFlow, Kubernetes Engine, Cloud Firestore, App Engine, Cloud Functions, Cloud Monitoring, Cloud MemoryStore, Cloud Spanner, Cloud Storage, Cloud Composer, Cloud Dataproc, Cloud KMS, Cloud Container Registry, Compute Engine, Cloud IAM, Cloud SQL, Firebase Storage, Cloud Healthcare API, Cloud AI, Firebase Machine Learning, Data Catalog and Cloud Console.
> We are currently investigating an issue affecting Dataflow, BigQuery, DialogFlow, Kubernetes Engine, Cloud Firestore, App Engine, Cloud Functions, Cloud Monitoring, Cloud MemoryStore, Cloud Spanner, Cloud Storage, Cloud Composer, Cloud Dataproc, Cloud KMS, Cloud Container Registry, Compute Engine, Cloud IAM, Cloud SQL, Firebase Storage, Cloud Healthcare API, Cloud AI, Firebase Machine Learning, Data Catalog and Cloud Console.
When you see that, it's almost certainly a low level networking outage (networking hardware failure or a very low level networking config change that isn't possible to canary with the vendor's current technology, possibly combined with something insufficient failover resources due to other issues).
The bad part here is that its impacting multiple regions. really hard for customers to build high-availability solutions when isolation boundaries can be breached this way.
(disclosure ex-aws)
What region is this? I have not seen a blip in central.
Apparently the outage is caused by a third party router bug: https://twitter.com/uhoelzle/status/1243398255083311105
However, as an engineer... I can empathize with their situation. GCP has less public cloud experience than AWS. They’re somewhat transparent about outages (so much better than being gaslit by AWS). Shit happens and at least the same problem doesn’t happen twice, which is what I really care about.
However my non technical manager won’t think this way and neither will other companies considering Google Cloud. GCP is in a tough spot from a marketing perspective if they get a reputation for being unreliable due to these incidents.