A few quick thoughts---I stayed around for 3-4 mins so please read this from the perspective of a (techy) user didn't explore much:
The good:
The UI is simple and clean. Definitely feels like something that you can start with.
The bad:
Email went to spam folder.
Nobody within 100km of me (which is the max distance) and I live in a very busy city. Maybe for launch, allow a distance of infinity so people that just want to meet new people stick around and don't forget about the website.
Also, I am not really a fan of setting an intent and not being able to change it for 24 hours. It wasn't mentioned anywhere either.
I couldn't explore more as thats as far as I could explore without any people around me.
WHO's responsibility is not to a single person. Their policies are designed for the society as a whole. Minimizing the possibility of infection for critical jobs is much more valuable than minimizing the infection probability to a single person. Society is more than just me or you or our families.
They could have taken the path you are suggesting (possibly under inconclusive benefits of other types of masks) but that could have backfired and caused a lot more dmg than good.
You can't expect an entity like WHO to solely care about you as a person.
This. And one of the main reasons that the paper from Google are routinely of high caliber is due to the internal review processes and your peers.
Publishing a piece of work for the sake of publishing by ignoring the processes that are put in place is outright irresponsible and that's the end of it.
The statement that "Academic debate is, in fact, done through conferences and journals" is not strictly true. Specially given that a lot of reviews in more popular conference are very hit and miss. You can submit the same paper to the same conference multiple times and get wildly different opinions on the same paper.
The variation in reviewers' response is often due their lack of knowledge and unfamiliarity with the problem. Take a look at the recent reviews on some of the more popular conferences on OpenReview.net. Most of the reviews don't have any substance and are often vague/generic.
I'd take the reviews from peers that I trust and are aware of my work more seriously than reviewers of conferences.
How is this different than picking up the phone and having your convo go through ATT/Verizon networks? or using your ISP? Both parties can "legally" work with authorities to wiretap you?
Are you worried that the (training) algorithms that run on your voice somehow end up leaking your identity? Or are you worried that someone at Google knows your voice?
Also, not sure if Google has this fact in their ToS but if they do, what is the issue?
Do you have numbers or papers that support your argument? That these applications are bottlenecked by the network?
There is a 2015 paper [1] that argues that improving network performance isn't gonna help MapReduce/data analytics type of jobs much:
" .. none of the workloads we studied could improve
by a median of more than 2% as a result of optimizing
network performance. We did not use especially high
bandwidth machines in getting this result: the m2.4xlarge
instances we used have a 1Gbps network link."
Granted things might have changed by now, but I am curious to see how and by how much?
I am genuinely curious to see what types of "compute-intensive" applications fit the bill here. Outside of storage workloads (syncing data, etc.), why would you need a 100x improvement in data transfer rates between the machines? (We have TPUs and their specialized network architecture for ML-like workloads ...)
Physical distance between the machines in a DC prevents RAM style "shared-memory" architectures, at least ones that aim to have 30~60ns access times (10-20 meters). Unless there are new paradigms for computation in a distributed setting, I don't see the benefit for this ...
Also, what are the fundamental limitations/research problems of todays hardware that prevent us from building a 400G NIC? I cannot think of anything outside of PCI-e bus getting saturated. We already have 400G ports on switches ...
> I could ask them to go look up papers in Oakland, CCS and NDSS over the last couple of years and see if anything catches their fancy.
Isn't this exactly the major thing that is wrong with research today? Limiting work/creativity to a few well known conferences done by elites for elites? I read blog posts, posted daily here on HN, that are way more informative, honest, and replicable than many papers published in the three conferences you named.
> "elite" status. Some of this is deserved, because these people have done good work. But some of this is also just a publication cartel where everybody cites their friends' work and make it impossible for others to break into a field.
Elite status happens exactly because there are conferences like the ones you mentioned. If you work with an advisor that publishes in Oakland, your chances of getting a paper in Oakland gets increased multiplicatively. And hint, that's not because your ideas (or papers) are better than anybody else's.
> The larger point is that in a scenario where there are so many papers that nobody could possibly look at all of them will lead to a few groups accumulating all the citations and all the awards.
This is already happening. Look at all the "prestigious" conferences.
> where researchers don't have the PR muscle power to highlight unpublished stuff.
Who cares? If the work is worth anything, people will cite it. If not, it will remain as is. Why does it matter? Why do you care if 10 people cited your work or 100 people if you are happy with the work?
Unfortunately, nobody in this forsaken field (computer science) cares about the scientific aspect of the field anymore; everybody wants their name to be known and that's all there is to it. The measure of success is how many papers you publish in elite conferences ...
I do actually think that by breaking down all the barriers people care less about having their name in conference X or Y and more about the scientific aspect or citation cartels. First one is good, second one can be fixed (at least more easily than giving a few elites lots of power with no checks and balances).
You are assuming that everybody in this system is honest: researchers, funders, and companies. None of these entities need to be honest to do research or to come up with topics. This is a big fallacy with academic research: nobody needs to be honest or truthful, and as much as you like to believe they are, they have no reason to be honest. I have experienced this first hand during my PhD life.
> Researchers are normally happy to do important things and they have good ideas.
That's not how research works, however---at least not what I experienced. Usually a goal is the prerequisite of the funding. If that goal doesn't align up with the interest of the company/or the funder, there won't be any funding to begin with. That's a lie that researchers tell themselves, that they have the freedom to work on w/e they find amusing.
> Academia being driven by money has less to do with academia and more with lack of money it needs to work on independent research.
Sure. If academia was a utopia that wasn't so dependent on money (and fame), science would take over. That utopia isn't anywhere in sight, however. Try going through academic job market: a popular question you get is how you are going to bring money into the university. If you don't have a good answer, you won't get a position. As a professor, you will be spending 80%+ of your time writing grants or, as you put it, begging different companies for money. The other 20% is spent on classes and bureaucratic headaches. A good professor that I knew had a long rant on how she never gets to do research anymore (from a very well known university).
People twist words and data to make their point, get a publication, make a name for themselves, get more money, and repeat the cycle, all in the name of research and science.
My point isn't that the academic people as a whole are bad, but many of them are trying to survive and are willing to do w/e it takes to bring in money. It is very easy to abuse the bunch when people are trying to survive.
All I am suggesting is that you need to be critical when reading academic papers. Regardless of the content, topic, etc. read and decide for yourself---if it is a field that you are proficient in---or read opinions on counter points.
Unpopular opinion, but they aren't that far off with regards to academia. Academia has little scientific agenda at its core—most things boil down to money.
If you want to make a case for your idea on any topic, put money on that topic and professors and researchers work towards making a case for it (this is less true about math and (maybe?) physics but the farther you get from those topics, the more impact money has on results and ideas).
Academia (mostly? in the US) is just yet another institution that is solely driven by money.
1) Are you using or relying on DMA or SPDK to copy packet data? A single core, to my understanding, (assuming 10 concurrent cache lines in flight and 70~90ns of memory access time) doesn't have the bandwidth to copy that much data from the NIC to memory (assuming the CPU is in the middle). If so, RDMA and the copy methodology are not so different in how they operate.
I didn't look at the paper where you explained how you perform the copying.
Also, IMHO, RDMA itself is a pet project of a particular someone at somewhere that is looking for promotions :) . I don't really know if it's a good baseline. It could be more reasonable to look at the benchmarks of other RPC libraries and compare against the feature set they are providing.
2) As far as I remember talking with random people from Microsoft, Google, and Facebook, none of them use Timely or DCQCN in production. Microsoft may be using RDMA for storage like workloads and relying heavily on isolating that traffic but nothing outside that (?) . I could be wrong.
3) There definitely is a need to distribute the load unless you are assuming that the single core can "process" the data. That may work for Key/value store workloads but what percentage of the workloads in a DC have that characteristic? You say there are tons of communication-intensive applications, care to name a few? I can think of KV stores. Maybe machine learning workload but the computational model is very different there and you rely on tailored ASICs (?) What else? Big data workloads aren't bottlenecked by the network BW.
4) I won't dive into the CC/BDP discussion cause it's very hard to judge without actually deploying it. Sure, a lot of older works made a lot of claims about how X and Y are better for Z and W, but once people tested them they would fall flat for various reasons.
I completely agree with your observation. It's not that the Ph.D. prepares you for it, but that most people that pursue Ph.D. have that attribute.
I don't think all Ph.D. students or any ordinary engineer can take a vaguely defined research problem and make it succeed. Ph.D. doesn't necessarily prepare you for the research part; it does, however, prepare you to advertise a very /bad solution/ as a novel contribution. It wasn't always like this, but it has come to it. It takes credibility, curiosity, character, and of course, research skills to solve a vaguely defined problem---none of which are given to you by a Ph.D. degree.
I have worked at two FAANGs, and more often than not my interactions with Ph.D. degree holders have left a bad taste in my mouth. Speaking of which, there was one person that was "selling" an event timeline as a root cause analysis system that does "temporal" correlation (with no filtering or association at all) :). And another person that was advertising a DFS compilation of a neural net during the training phase (as opposed to the typical BFS that people do) as a superior and novel contribution that changes how we think about neural nets or something along those lines.
A good engineer would have laughed at both after carefully considering all aspects of the problem.
I suspect that Ph.D. "engineers" are more desirable because of their broader skillset (they have worked with more tools and have taken more classes) and also the fact that companies can hire them at almost the same cost as a BS/MS degree holders. Plus universities have already done some filtering on Ph.Ds.
Your repo is really nice for an academic paper. Thank you for that. It's rare to see a "networked system's" repositories that has readable code. I mainly checked large-tput example:
A few questions—
1) For your 75Gbps, what percentage of the payload of the RPC do you touch? I.e., what portion of the message is used on that core?
More directly, say you have a service that can sustain 100kQPS, if they switch to eRPC, what can they expect? Asked differently, what is the base overhead of today's RPC libraries? Especially ones that bypass kernel.
2) The congestion and flow control is debatable, and their efficacy is up for debate. Especially in a DC setting. Can you claim that eRPC would work for any types of the workload in a DC setting? How would it play out with other connections? At the end of the day, if you are forced to play nice, you may eventually add up branches in your code. Your fast path gets split depending on the connection type, etc. Is that something that you think is preventable?
3) How do you distribute the load across different cores at 75Gbps? How does the CPU ring, contention, etc. come into play? I.e., can you do useful work with that 75Gbps? or should I just read it as a "wow" number? Asking a different question, if I have a for loop that can do 10 billion loops per second and by just adding a function that drops down to 10k loops per second, why would I care about that 10 billion iterations?
4) You claim that it works well in a lossy network, yet your goodput drops to 18~2.5Gbps at 10^-4/10^-3 packet loss---I am still assuming the library is still flooding the network at 75Gbps. How does this play out in scale?
All in all, I do appreciate your work. My issue is that academic people like to make big claims, especially in an academic setting. People in the industry are aware of fast-paths. Kernel networking stack uses fast-paths rigorously. Sure it is heavy and it comes with a lot of bulk, but you can as easily cut it down.
Just to restate what I said---you cannot use this in its current state in production (or industry), ever. And by the time it becomes useful, it becomes a natural solution because the infrastructure supports it.
Every person that works on kernel knows the overhead that comes with abstractions. Everybody that has worked with DPDK knows that you can get 20Mpps+ on a single core.
All they have done is to "frame" the usage and the term RPC differently ... i.e., it's all story telling and no real meat :)
Look at the previous set of publications by the same author: e.g., Achieve a Billion Requests Per Second Throughput on a Single Key-Value Store Server Platform, etc. they are all based on a single assumption that if you don't implement X and Y in your stack you can get better performance. Of course, you can. If you use your calculator to only compute 2+2 you might as well hardcode 4 as the output of your calculator.
They are not setting a lower bound. They are hardcoding and bypassing the parts they don't find useful.
That is hardly the state of the art: they are basically sacrificing all abstractions that are rightly so required in the name of speed. This is no different than using vanilla DPDK with no congestion and flow control and being able to process 20 mil packets per core (better than the numbers in the paper). Getting 75Gbps per core using RDMA is hardly hard or new.
You can't use this abstraction for anything ... really unless you are on a completely lossless fabric that has enough capacity to avoid congestion.
Yes, Google/MS/FB are going towards such fabrics but the hard part is not building the RPC abstraction suggested in this paper---the hard part is getting to that fabric.
Just to put things into perspective, if you had a quantum computer you could do all sort of crazy stuff with it. You could "parallelize loops! and be super fast" is what this paper is suggesting (I specifically chose parallelizing loops cause there is nothing new in it).
> Instead, we see a thriving software industry that largely ignores research, and a research community that writes papers rather than software.
I have actively followed the NSDI and SIGCOMM community, and this is, for the most part, true: Research venues have become a hiring billboard for big companies (Microsoft specifically?), and that's all that is left of research.
Most papers published from these companies (and academia) are flawed at the core and primarily story driven (at least in the two conferences that I mentioned). Companies publish with data that is inaccessible to anybody but them---MSR, specifically, takes this to an extreme. The scientific contribution of most papers is close to nil. Writing and storytelling dominate the field. Experiments are cherrypicked, are rarely reproducible, and the software is seldom useable.
People rarely are willing to think outside the box and spend more than a year on a paper. Most people pick an idea from an outside field, apply it, and publish a paper. That's the end of it. Most people that I have talked with publish to get their name out, and rarely care about any scientific agenda.
Imagine, in a system's research community, a large portion of academic advisors cannot develop proper software---they can, however, pitch stories and write text for days.
The problem is not that people are evil... it's that the system/community has decided to take this path. And for one, I cannot fathom why. It is not even rewarding to publish a paper in these conferences anymore ... except to enjoy the trip and the dinners.
You hear stories on how people try to optimize their chances of getting in, e.g, I have heard and seen from good researchers that you should not register your paper early because you will get a two-digit paper number, indicating that this is a resubmission, and lowering your chances of getting in, etc. There are many such hidden gems there.