First of all, I don't think this is satire. I'll admit that the use of a gmail account by a researcher at a Chinese uni is facially suspicious, but it's not that odd given that cursory googling shows that both authors appear to be faculty members at Shanghai Jiao Tong University as claimed on the paper- though neither appears to have much, if any, background or expertise in machine learning.
I'm not much of a fan of a lot of the arguments made in Weapons of Math Destruction, but I do appreciate that in summarizing you draw the distinction between the biases of the engineer or (illogically, but oft-claimed nonetheless) the algorithm itself and the data which is used to train said model, and I think it's quite a valuable concept in regards to this particular paper.
For instance, the data set they're using here is fairly small, and while, they did use 10-fold cross-validation, that's still a bit on the less than ideal side generally speaking neural nets, especially CNN architectures, which are usually pretty deep. Furthermore, the dataset itself seems fairly questionable to me. I'm not sure how much I trust the Chinese criminal justice system to adequately adjudicate culpability in the first place, but even setting aside such admittedly conspiratorial notions, it seems rather odd indeed that nearly half of their positive samples are not in fact convicted criminals but merely suspects. I do not find their attempts at devil's advocate persuasive as it's not readily obvious exactly how they used or obtained any of their testing with the three different data sets.
As for the appropriateness of the broader topic, I'm more or less of the persuasion that all questions deserve to be examined, and that provided the work does not cause direct harm, it's hard for me to support a prohibition on examination of a given topic. That said, I do think that the more controversial the question, the higher quality of research required, and, good lord, does this mess fall well short of the mark. Perhaps if there existed a hypothetical criminal justice system free of systemic biases or, more realistically, a method by which to exactly define those prejudices and account for them in the composition of a data set, this could be a potentially useful question to investigate, but even then it seems to me quite unlikely that there's any particularly significant relationship between one's upper lip curvature and criminal disposition.
I don't read a lot of techcrunch to be honest, so maybe this is their standard fare and isn't worth the time I'll take to write this. On the other hand, I dislike that this man just painted my entire industry as xenophobes, without so much as a quote from the party in question. Is this acceptable now? I certainly reject the idea of a world in which psuedo-journalistic conjecture is taken at face value, without so much as an anonymous source.
can you explain more about how you want "coders" to do "missions", but you're charging 20k/yr for access to part of the data, and you have an obnoxiously low rate limit? are you bloomberg? I think we should all stick with quandl for now. But seriously though, it's disingenuous to brand yourself as "open data" but not offer bulk downloads and a realistic API limit.
deeplearning.net and the ufldl tutorial are excellent places to start. I've also perused this ebook recently and found it to be pretty solid in terms of giving mathematically solid but still intuitive.
last but not least, absolutely do no whatsoever "do a small project on deep learning or try out [a] few kaggle competitions." instead, pick up a paper that interests you and implement the methods they describe therein.
edit: here's the ebook link neuralnetworksanddeeplearning.com
I found that one particular phrase,"technically privileged", both hilarious and utterly horrifying. Hilarious because I suppose I never paused to consider that the folks in my major who were both curious and motivated enough to be involved in the field outside of class to be particularly "privileged". I suppose I should have given greater thought to how unfair it was that read API documents and I didn't!
Like I said though, it wasn't just good times and passive aggressive complaining about first world problems while I read this. I can't even begin to fathom the miserable, pathetic, and generally small, unexamined existence that would lead one to believe that h(s)e is somehow the victim of deep injustice at the hands of those people with their prejudiced technical abilities and natural curiosity! How dare they not level the playing field just because she never bothered to explore the use of code outside the classroom; clearly everyone missed that she's the subjugated one, with a comfortable liberal arts education and regular internet access.
But still, we should all take a moment to recognize the plight of the comfortable, generally satisfactory lives of those among us struggling silently with the burden of "technical un-privilege."
I'm not much of a fan of a lot of the arguments made in Weapons of Math Destruction, but I do appreciate that in summarizing you draw the distinction between the biases of the engineer or (illogically, but oft-claimed nonetheless) the algorithm itself and the data which is used to train said model, and I think it's quite a valuable concept in regards to this particular paper.
For instance, the data set they're using here is fairly small, and while, they did use 10-fold cross-validation, that's still a bit on the less than ideal side generally speaking neural nets, especially CNN architectures, which are usually pretty deep. Furthermore, the dataset itself seems fairly questionable to me. I'm not sure how much I trust the Chinese criminal justice system to adequately adjudicate culpability in the first place, but even setting aside such admittedly conspiratorial notions, it seems rather odd indeed that nearly half of their positive samples are not in fact convicted criminals but merely suspects. I do not find their attempts at devil's advocate persuasive as it's not readily obvious exactly how they used or obtained any of their testing with the three different data sets.
As for the appropriateness of the broader topic, I'm more or less of the persuasion that all questions deserve to be examined, and that provided the work does not cause direct harm, it's hard for me to support a prohibition on examination of a given topic. That said, I do think that the more controversial the question, the higher quality of research required, and, good lord, does this mess fall well short of the mark. Perhaps if there existed a hypothetical criminal justice system free of systemic biases or, more realistically, a method by which to exactly define those prejudices and account for them in the composition of a data set, this could be a potentially useful question to investigate, but even then it seems to me quite unlikely that there's any particularly significant relationship between one's upper lip curvature and criminal disposition.