Author here. That's definitely my bad and not an intended user experience. The text was initially meant as a transcript accompanying presentation slides. I've compressed the images so it should be at least slightly better now.
The good news is that this already exists in the Redis search module [1], which allows you to do similarity search against indexed embeddings, among other features, and offers comparable performance against other ANN libraries [2], depending on your performance criteria.
I've been using it for a side project to do semantic search on books[3] and have been really happy with its performance. (Not affiliated with any of this, was mostly interested in exploring existing well-performing, fairly standard tools with low latency)
Author here. The primary difference is that Confluence is more like a reference manual. A P2 is more like a conversation and a record of decisions made, along with all the context for those decisions, since each post has a threaded comment section where people can reply to each other. It's not better or worse, but has a slightly different use-case in that it's essentially like email threads or RFCs that anyone in the company can read.
Editor here. Thank you for sharing your memories. And, for the note.
Some details about the origin: the vignettes were originally English-language comments in a private Facebook group.
It was a huge team effort to get them all collated, organized, and to check and re-check attribution and consent to share. We are so happy that we are able to share these collective memories and eyewitness testimonies, particularly outside of Facebook's walled garden.
Editor here. These were originally written in English, so we have plans to translate to Russian, which might take a while since we only have one fully-qualified translator. Help is welcome! Shoot me an email (included in my profile).
This is a rather interesting stance to take for a publication whose article quality has degraded and ad size has increased (thirty-eight trackers, including two from Facebook, detected by my adblocker when I tried to access the article) over the past five years to the point where I refuse to read them.
If you take a look at the homepage through the Wayback archive as it used to appear in 2005[1], 2011[2], and today, you'll see how content disappears and click-baity headlines rise over time.
The Atlantic is very much a part of the problem of "the race towards the bottom" the author describes, and instead of having a discussion about how to fix it and maybe trying different revenue models, it continues to un-ironically have share and tweet buttons at the top of this article.
I don't know that that's necessarily true. The most recent StackOverflow survey[1] shows a difference of 8%, which is not an overwhelming majority. Granted, that's not an unbiased sample size, but I think the OP above is correct...more data scientists use Python than Java.
So anyone wanting to use this library would have to think about tradeoffs: Are the efficiencies lost in data scientists learning to use Java for modeling worth the efficiencies gained in putting a model in production? For some, the answer may be yes, for some no.
Came across this yesterday. Can someone (maybe poster?) talk about when you would use this versus something like scikit-learn or any number of R libraries? Is the goal simply to have all machine learning in Java so it can be productionized easier?
Perhaps it's just my selective biases, but I feel like I am reading about more notable people in the tech industry speaking out in favor of privacy over the past couple months than even when the Snowden revelations happened in 2013.
Could this possibly be a tipping point for adtech as a revenue model? (Although I've been reading about it for almost a year now [1], [2], [3]) I'd like to hope so, and am also curious: how invasive can companies be before consumers start to push back?
Data scientist here. It is 100% possible to do things with kids, but you really have to be motivated to do it, AND it really helps if you have a support system of other people to help take care of your child. I wrote about the dangerous deception we have in American culture, and particularly tech culture, of people who "have it all," but in reality have a bunch of help in the background here[1] and here[2].
If you work full-time and you want to go above and beyond, you're essentially working three days: work, before school + after school, and then your third day is learning or development.
Whatever that means for you in terms of reshuffling energy and other commitments will vary on your personality, energy level, etc, etc.
When you have a small child, it is extremely hard to multitask. So I wait until she is asleep. All after work time and weekends are for her.
Here is the way my schedule works: I pick her up from daycare, do dinner, playing, and then she goes to bed. I then take half an hour break, and delve into whatever I have going on, for about three hours.
I'm currently taking a Java class, writing technical blogs, and working out some Python. So I'll usually do an hour of reading/Java homework, then start a blog post, then finish off with whatever else I was working on.
Over the past three weeks, I developed this talk on big data[3]. That was probably the hardest because I needed a lot of time to write the code, test the code and concentrate, and all of my energy was just sapped.
All of this is to say that you can do it. For me personally it takes a lot of reshuffling and work and giving up things, but that's how kids work.
There are definitely a lot of people who fall into those categories, but the articles/links give the impression that everyone is senior. Which is why it's great when articles like this come around.
This is a fantastic article for intermediate beginners. On HN, everyone is a senior data scientist working with Spark and Keras and Tensorflow and deep learning.
In the real world, there is a huge chasm of difference between people just learning Excel and developers, not many people even understand why you would switch away from the former when it's so convenient, which is why the difficulty v.s. complexity chart is so great, and may actually speak to people in an approachable way.
There are a lot of tutorials for how to do hard things and how to do easy things, but not a lot for how to think of the hard things in terms of the easy things, and this falls in that category. Another good book on this topic is Data Smart by John Foreman, where he goes over basic data science skills in Excel.
I am disappointed at the type of criticism of this post on Hacker News. This write-up is the kind of high-quality, original, analytical writing we so desperately need these days, in an online world that is completely saturated with clickbait that adds nothing to our understanding of the world.
Is the piece lacking in some semantics? Perhaps. But I was struck by the comments about market share and Android strategy, rather than directly discussing the article and the points it makes about the two maps at hand.
If your wife is already interested in writing about this, and already doing it, then there is absolutely value: to her. She gets to process what she's working through and pour it out on paper. That's already great. It will absolutely also look great on her portfolio. It might also be good tech exposure for her to start working with either Wordpress, or extra bonus points, Jekyll, and version control.
Speaking as someone who mentors other people in making their way through data to programming, yes, please, please have her write about this!
There are so many people looking to get into programming, but literally have no idea where to start because the whole universe is overwhelming to them and full of people who seem to be programming forever. HN is probably not the audience for this blog, but hundreds of thousands of data analysts and people who use Excel on a day-to-day basis are.
Finally, a nitpick, but I'd take issue with "These aren't enough to actually work in our industry" - if she's working with technical skills and wants to learn more, she's already in the industry and then some.
The more I read, the more I realize that fiction has a way of teaching the same lessons as nonfiction, only in a better, humane way that is not like someone lecturing at you from a podium, but more like a friend sitting down for a heart-to-heart conversation with you.
In this vein, I've started catching up on all the classics I haven't read yet, and there are a lot of them. At first, I was hesitant to dive into classics, because I thought they were hard to read and highbrow. This has mostly not been the case (with the exception of Wuthering Heights, which I found impossible to get through).
To wit, the best books I've read this year include:
A Tale of Two Cities - Impossible to get through the first several chapters, but after that, you're off to the races between France and England in the 18th century. Dickens used historical books as reference, but recreated the mood entirely from imagination. A better primer than anything in the news about Anglo-French relations, learning trust, and sacrifice.
No Country for Old Men - Want to understand the current fear of the Mexican border that has so bolstered Trump's popularity, as well as the ramifications of having a country full of veterans who don't have any medical support or care? Read this book.
Farenheit 451, A Clockwork Orange, Handmaid's Tale - Much better and scarier than any news clipping today about the possible future path for government interference in thought and action. Takes things to their logical conclusions.
And one non-fiction book that I did enjoy, Bringing Up Bebe, which is about the contrasts between child-rearing in France and America, but on a much larger scale, about the different things that seem culturally obvious to us, but are completely different to other cultures. This book made me really re-examine American food culture in a way Michael Pollan's diatribes have not.
> What's so difficult about just opening an incognito window and suffering the couple of seconds it takes for a computer halfway across the world to deliver you content while you sit at your desk?
The whole point of the post was that I'm doing something inordinately ridiculous to access a little bit of good content. Scraping and incognito browsing are both anti-patterns that are symptoms of a sick media industry. A media that is sick cannot provide us with good news and content that we can use to further our critical thinking, and we should be worried and thinking about how we can possibly solve this problem instead of trying to bypass it.
Thanks for reading. The Economist ads are a great point that I didn't address, and makes me even more worried for the high-quality news and content industry than I initially noted.