Initially, it seemed like a DoS to us too, but it was not. This was confirmed by upstream provider metrics. No major traffic spikes. It was a combination of non-malicious things. More info later, some of us need sleep.
I feel like there is an element of "Body Doubling" here... a strategy used by those with ADD/ADHD. I recently looked in to this when curious about my own observation that I work longer and with better focus when working in close proximity of someone else.
This morning, I googled for issues with the firmware and the model of SSD, I got nothing. But now I am searching for "40000 hours SSD" and a million relevant results. Of course, why would I search for 40000 hours.
They were in two mirrors, each mirror in a different server. Each server in different racks in the same row. The servers were on different power circuits from different panels.
These were made by SanDisk (SanDisk Optimus Lightning II) and the number of hours is between 39,984 and 40,032... I can't be precise because they are dead and I am going off of when the hardware configurations were entered in to our database (could have been before they were powered on) or when we handed them over to HN, and when the disks failed.
Unbelievable. Thank you for sharing your experience!
It was part of a mirror of identical SSDs on an LSI MegaRAID RAID card. We see occasional "spectacular" drive failures that take the machine down with a single disk failure. Usually it's just a reboot to come back up, and a disk replacement, then some hours of time to rebuild the array and get back to situation nominal.
People guess the origin of our name often. Maybe this will give you even more of a chuckle. I was not aware of the name of this computer when I named the company. https://en.m.wikipedia.org/wiki/The_Ultimate_Computer
Thank you for sharing your positive experience! We can power cycle power outlets remotely and can connect a console (ip kvm)... and we are staffed 24x7.... in case you need another server. Thanks again!
Oh hi! Thank you for the kind words. I cant tell who you are by your name here, but if you've been with us since 2011, we have certainly spoken. Are you using our second San Diego data center for your failover location? If you and I aren't already talking directly, ask to speak with Mike in your ticket.
Unrelated issues, but I did hear from our other clients that O365 was having issues at the same time as our network outage affected HN and many others.
Founder and CEO of M5 Hosting here. We did have a network outage today that affected Hacker News. As with any outage, we will do an RCA and we will learn and improve as a result.
I'm a big fan of HN and YC in general, we host of other YC alum, and I have taken a few things through YC Startup School. During this incident, I spoke to YC personally when they called this morning.
M5 Hosting | San Diego, California | Sr. Systems Admin / DevOps | Work remote most of the time, but must be able to go to the data centers in San Diego regularly. Remote data centers in EU and AP. We provide public cloud, dedicated servers, colocation, and hybrid IaaS environments for clients ranging from a single VM to 1000+ servers.
Cogent and Cox are also having problems, but we are seeing a lot more successful traffic on Cogent than CenturyLink. It appears that CL is also not withdrawing stale routes. It seems CLs issues are causing issues on/with everything connected to it.
M5 Hosting here, where this site is hosted. We just shut down 2 sessions with Level3/CenturyLink because the sessions were flapping and we were not getting complete full route table from either session. There are definitely other issues going on on the Internet right now.
Founder/CEO of M5 Hosting here. M5 Computer Security pivoted to M5 Hosting in the early 2000s. We do host the servers that this site is on. I have been through YC Startup School recently. We host a few Fortune 500s and many startups.
Actually, we just bought bandwidth for a roll out at Equinix in Munich. $0.50 for Cogent (when added to several other 1G commit on 10G ports in our account. A single 1G commit on a single 10G port would cost more) and we were quoted $1.43 for Level 3 after rejecting a $1.70 quote. Both Cogent and Level 3 were 2yr terms, and in an "on net" location. We are going with another provider besides Level 3 there, but I used these as examples in the parent. You thought it was not "reality", and I refute that.