Agree. From 6+ years of experience it seems that we got fouled by the multi-az promise of being able to survive datacenter outage.
You can survive datacenter (AZ) outage IF you have separate stacks per AZ and don't mix traffic. If you have Kafka cluster spread out in 3 AZ don't get surprised if you just LOWERED your availability because any issue in one AZ makes your stack unstable. And issues in single AZ are quite common.
PromQL support (with extensions) and clustered / HA mode. Great storage efficiency. Plays well for monitoring multiple k8s clusters, works great with Grafana, pretty easily deployed on k8s.
If you're looking at scaling your Prometheus setup - check out also Victoria Metrics.
Operational simplicity and scalability/robustness are what drive me to it.
I used to to send metrics from multiple Kubernetes clusters with Prometheus - each cluster having Prom with remote_write directive to send metrics to central VictoriaMetrics service.
That way my "edge" prometheus installations are practically "stateless", easily set up using prometheus-operator. You don't even need to add persistent storage to them.
At the time of making this presentation AWS did not have anything in their offer that could match tuned MySQL on i2 instances. Aurora was just getting started.
From experience: after company grew to more than .. 200-300 people and user management/termination became a big burden we hired a person that would write tools to automate user management, and if something wasn't supporting SAML we did manage users via its API. If API was not available then we reverted to "Termination checklist" aka manual work.
Clarification: it wasn't that persons only responsibility, just one of many assignments to help automate Ops in the company.
To each of you guys having those extensive backup solutions (like NAS + cloud sync, second nas, etc)...
.. do you actually TEST those backups?
This questions comes from my experience as a system engineeer who found a critical bug in our MySQL backup solution that prevented them from restoring (inconsistent filesystem).
Also, a friend of mine learned the hard way that his Backblaze backup was unrestorable.
Been there, done that. AWS re:Boot in September 2014 showed us how good it was to invest in Ansible roles for all parts of our infrastructure. Still, a lot of hassle for Ops Team, especially that it was done during DevOps Days Warsaw ;-) AWS also said '10%' then, but for us it was 81 out of ~300 instances.
What is sad is that we learn about it from Hacker News and not from AWS, even when we have premium support and our own account manager. :/
Let's see how many of us did their homework after previous "xen update", and how much "10%" is now ;-)
I use MotionX GPS and download data from MotionX Terrain maps (which are OpenCycleMaps), good for getting around where I live. You might check coverage on http://www.opencyclemap.org/