Like the human body the more you study the Internet the more amazing it is not that it sometimes breaks, but that it works at all. Especially for video/phone/etc.
Glad the content was helpful, I have links to some of them at avi.net (tutorials and old Boardwatch articles).
I swear my motive was pure (frustration with the content out there) but it was easy to see back then that helping people out with good content yields rewards ("Can I buy a T1?") or ("Come run my big global network"). So I still encourage everyone to write about what's confusing and frustrating...
As the others say in the comments, understand completely re: servers. And network. And some management overhead. I will write a blog post laying out our COGS at Kentik including all these factors, hope it'll wind up helpful.
We're (Kentik) a SaaS company and one key to having a great public margin is buying and hosting servers. In our case, we use Equinix in the US and EU to host, with Juniper gear for routing/switching, and the customary transit, peering at the edge.
One secondary factor is that we've only monotonically increased, and it's way cheaper to keep 10%-15% overprovisioned than to be on burst price with 50%+ constant load.
But the simplest math is - we have > 100 storage servers that are 2u 26x2tb flash, 256gb RAM, 36 cores. They cost $18k once, which we finance at pretty low interest over 36 months (and really last longer than that). Factor in $200-400/mo to host each depending (I think it's more like $200, but it doesn't matter for the cloud math).
That same server would be many $thousands/month on any cloud we've seen. Probably $4-6k/mo, depending on the type of EBS-ish attached. Or with the dedicated server 'alternate' they are moving to offer (and Oracle sorta launched with).
It'd be cheaper but still > 2x as expensive on Packet, IBM dedicated, OVH, Hetzner, Leaseweb (OVH and Hetzner the cheapest probably).
Three other factors for us:
1) Bandwidth would be outrageous on cloud but probably not as outrageously high as just the servers, given that our outbound is just our SaaS portal/API usage
2) We'd still need a cabinet with router/switch infra to peer with a couple dozen customers that run networks (other SaaS/digital natives and SPs that want to send infrastructure telemetry via direct network interconnect).
3) We've had 5-6 ops folks for 3 of the 6 years, 3-4 for the couple years before that. As we go forward, as we double we'll probably +1. It is my belief that we'd need more people in ops, or at least eng+ops mix, if we used public cloud. But in any case, the amount of time we spend adding and debugging our infra is really really really small, and the benefit of knowing how servers and switching stuff fails is huge to debugging (or, not having to debug).
All that said - we do run private SaaS clusters, and 100% of them are on bare metal, even though we could run on cloud. Once we do the TCO, no one yet has wanted to go cloud for an always-on data-intensive footprint like ours.
Good luck with your journey, whichever way you go!
Unless that's capex for building, or includes server capex that they depreciate, that seems way high given Uber's scale.
A pretty dense cabinet should only cost ~$1400/mo at wholesale (1MW+ rooms) rates, and $200M is 143,000 cabinets.
And it's public that Uber uses multiple clouds as well.
Disclaimer: I haven't reviewed public Uber filings, would be very interested if there's any data that indicates they're really spending $200M on opex for real estate (which would be equivalent to $400M/year on cloud, which is either opex or potentially mix if there is some reserved instance-type cloud spend).
Absolutely a believer in the power of single machine analytics on GB or TB vs PB datasets. But I've found that whether it's a new data system or just showing csv/tsv tool magic, if people aren't into data systems (for the former) or unix cli (for the latter), it's a whole culture/training thing vs. most technically efficient wins.
So I think the latter issues dominate (culture/training) for new implementations as well as existing.
Re: efficiency overall, and of distributed systems, it's interesting to see both that MapR has made a business selling more efficient Hadoop-ishness, but also depressing how infrequently I see them deployed, even for pretty massive clusters.
Unfortunately, it's hard to see a startup focusing on single-machine solutions for GB->TB sets (even scaled out with replicas) as the way many in the space get started is with an open core model and the thing they charge for is the clustering and/or monitoring needed to become a distributed system.
But... I am optimistic we'll see a generational effect over the next 10 years in openness/interest in composable ad-hoc analytics tools, especially with Windows incorporating unix cli components.
Many of the alternates (including linux cli stuff) that are much faster require a re-thinking of attitude, don't work where there are tens to hundreds of people submitting queries, or require different skills. It's tragic to think of all of the computrons and watts wasted with Hadoop-ish stuff (map-reducing without filters, Java itself for most implementations) - but still I wouldn't recommend to most CIOs they replace Hadoop in all or maybe even most cases, even for few-TB data sets and smaller.
Both because of familiarity with querying and the solidity of running a multi-tenant system.
But I do recommend that they switch to MapR [c++ core and a passable central FS for unix-based super fast queries] if they're concerned with efficiency.
[For context, in my day job we do multiple clusters of millions of network traffic summaries/sec and are often replacing Hadoop, or more recently, ELK, as people tried to use them for that use case. All well >>> will fit in ram. We have our own in-house column-store + streaming combo db done in go/c/c++ that started as clustering fastbit.]
C5/C6 herniation about 10 years ago. Did the standard range of NSAIDS and PT. Useless and wound up needing prescription pain meds to get+stay asleep. Declined steriod injection because I wanted to debug.
What fixed it (on first treatment though I did 8 sessions) was the DRX9000 traction machine. The VAX-D probably would have worked too. Had to go to a physiatrist who insisted on doing homeopathy/BS saline injections so he could bill insurance. Though I had a cervical issue, these machines at their core are designed for spinal and the cervical treatment is an extension, I believe. What they do is slowly stretch you for 20-30 mins as you watch TV.
No relapse in the last 10 years though I've been careful about posture, use 4 wheel roll-along luggage, and sleep with fewer pillows - generally try to keep my head more aligned.
I was about a month away from doing ablation to have the goo sucked out of the offending disc. That probably would have worked, but I am related to a number of doctors, all of whom recommended staying away from surgery except as very last resort, especially near the spine.
Interestingly, internationally they have silicon and other disc replacement techniques as well whereas in the US generally treatment still tends to be NSAIDs+PT, then steroid shot, then fusion despite the collateral stress that can ensue from that on adjacent vertebrae.
Post 'fix' MRI looks almost the same as 'pre', which is interesting and gets to the micro tolerances involved.
2 non-standard treatment notes:
Tried Chiropractic as 1st treatment on common friend recommendation but never again. The practitioner started yanking before xray even though it was pretty clear (I now know) from my specific pain sites what was going on. An interesting question to ask a potential Chiro is "what diseases can and can't be cured with Chiropractic?"
Interestingly, along the way, I went once to an accupuncturist recommended by my Tae Kwon Do master and with 4 pins all symptoms went away instantly - for a few hours. Successive treatments had much less effect so I only went twice.
Good luck with your recovery!
In the grand scheme of things it was minor vs. other medical stuff, but I was unhappy about pain meds (and fingers starting to get numb after many months) so it seemed a huge deal at the time.
Big lesson for me was - with the human body, even more than the internet (but true for both), the amazing thing is not that they break but that they ever work well in the first place.
Happy to discuss my course if anyone is caught up in this, but I have no formal medical background and my case probably differs from everyone else's. Email avi at freedman dot net
We do network analytics delivered as SaaS, selling to enterprise and service providers, but I'm happy to talk (for free, of course) about VCs and the funded life.
I have felt like a cultural anthropologist in the land of VCs (mostly positive, some not) over the last 4 years. I know a bit about the Chicago VC environment but more about SF, NYC, BOS.
Have you checked whether any of your customers are connected to VCs? Enough of ours are that in our seed, A, and B diligence they found folks to talk about us that I didn't point them to.
Re: selling vs raising, I would say it really depends on your fire to change the vertical market you're talking about (or more).
The other thing I've seen (in both directions) is that if you get hooked up to a few CEOs, they can make introductions for you. Can you introduce customers to any CEOs that don't compete that have great investors that you'd like to talk to? Get most CEOs I know a solid $100k+ ARR customer into and they'll listen to your story and make intros if I think appropriate.
It is a disadvantage to not be in an investor-dense area, and travel costs money. I moved to SF to start Kentik, but I think the same principles would work if I had stayed in Philadelphia.
At the late stage, investors usually want to see 3-10x vs 10-100x as a realistic range vs earlier stage (A round) investors.
Usually they also look to there being a MUCH lower change of going completely to 0 - more like 10-20% vs well over 50%.
And usually there it's because of the fundamentals of the business are starting to show (margin, cost of acquiring customers, customer churn and upsell, etc.)
Sometimes IP or assets add value/valuation as well, though.
In Docker's case, there is also probably a feeling that the asset (control of "Docker") is worth hundreds of millions.
In CloudFlare's case, it is millions of sites as users, and the ability (mostly untapped, I think) to monetize the data from that.
If it's just the 2 of you, it might be easier for the company to buy back his shares at the issue price. Need to check with your lawyer, of course, re: valuation issues, but if it's just the 2 of you it may not matter, and future investors will be understanding of what was going on.
I read "typical worst-case pause time" in go and don't know whether to laugh or cry.
We are seeing massively longer pauses in production and the suggestion is to go to 1.8, which is not ready.
Reminds me of ruby - religious people practicing on production to get the nirvana.
Yes, it's getting better, or at least less bad, but yeesh. We'll just waste resources architecting around long pauses, and rewrite critical stuff in C, for a few more years...
Agree that bare metal is effective and doable at your scale, and can if done right give better SLA and much much much better control - especially of Internet-facing network performance than public cloud, or combos thereof).
We are running a 50%+ gross margin mid-stage venture-backed startup in Equinix facilities (but started there vs. cloud), and have no people near our facilities, and have had 0 issues service-wise related to doing management remotely. Yes, people go out to set up cabs, etc, but we hired our ops folks as generalists who had some network experience, and our CEO and CTO do as well, though AFAIK I don't have network logins active right now.
2 high-level thoughts I'd share:
1) Try not to use Ceph unless you're committed to having 2 people with deep experience at the code level.
2) I'd use Juniper QFX or EX, or Aristas. You don't seem to be running at scale or functionality where SDN magic is needed and there is a large community of QFC, EX, and Arista users your folks can reach out to when problems happen.
The other comments are more tuning and FYI on what we do HW-wise:
Specifically re: HW, at Kentik we run tens of worker nodes + flow ingest servers, all SM 1us w a few SSD and 256-384gb RAM. 48 logical cores, 2 x E5-2650v4.
We run approaching 1PB of storage, and while we still have some 4u 36-disk 3.5" boxes, those are phasing out and all we buy now is 2u SuperMicros w/ 24x2TB Samsung Evo 850ss. Procs are 72 logical core, 2 x E5-2697v4.
The Evo SSDs have been great - but our workload is largely appends or create/writes - largely but not all sequential, with high read IOPS. Before Samsung I was a big fan of Intel but we have no data on the modern Intels - slower for sure, but a focus on reliability is great...
We use JBOD and ZFS on the storage nodes; the LSI 9300-8i. Have things tested so we can do TRIM.
They do make SuperServers for roughly those configs, but we go with SM resellers who assemble and burn-in for +10-15%. I had 50+ SuperServers that were great at my Usenet company, but we'd rather have our ops folks work on things other than burn-in.
Happy to explain why we went to SSD vs. spinning at 2x the cost, but basically it made enough of a different at 95th and 99th percentile in our query times, and we had access to venture debt on great terms (which you should too and happy to discuss, since we're both funded by August).
Last note re: gear - when we were doing spinning, we found a screaming deal on new 2TB enterprise SATA (Hitachi, I think) for $50 and took the power/space hit for the +IOPS and extra compute we got for firing up the additional machines. Not sure if those are still out there, or the IOPS of this kind of approach would be needed.
Agreed - LeaseWeb, OVH, Hetzner, and even SoftLayer if you call and negotiate can all be great options and have been very stable for many folks for dedicated servers. Generally I recommend that people not make long term commits, as it gives more leverage if there are network hot spots or other issues you need their help resolving.
I think they can be pretty efficient and if a handle is kept on what's deployed on top, can be not that much overhead over the cost of the infra (whether it's owned or cloud). But that keeping a handle on things is key - starting w/o a DBA is nice but if no one is tracking tables or how they're used things can get pretty expensive :)
Specifically re CF - have seen a few companies use CF to do multi-infrastructure, but a lot of the companies we work with have 5-10 roles and just run them via config mgmt to deploy or now docker +/- k8s, and don't use PaaS at all.
Absolutely agree that collateral damage vs many small-mid-sized hosting providers is 0 in Amazon, though you do still have to deal with the normal 'noisy neighbor' problem by re-creating instances in a different neighborhood.
They do manage those - though mainly for protecting themselves, not the point customer being attacked. In AWS the people I've talked to recently as well as historically say you'll get pretty uniformly rate-limited? vs. actually doing per /32 DDoS mitigation type limiting. Has your experience been different (for volumetric attacks)?
Like the human body the more you study the Internet the more amazing it is not that it sometimes breaks, but that it works at all. Especially for video/phone/etc.
Glad the content was helpful, I have links to some of them at avi.net (tutorials and old Boardwatch articles).
I swear my motive was pure (frustration with the content out there) but it was easy to see back then that helping people out with good content yields rewards ("Can I buy a T1?") or ("Come run my big global network"). So I still encourage everyone to write about what's confusing and frustrating...