I guess it is nice to know that he is also not perfect. But it’s still the case that his accomplishments outshine my own, so my imposter syndrome remains intact.
When I worked as a post doc, I wasn’t paid much but I got a direct return on extra work: another paper etc. when I got my first job as a data scientist, I was paid much more but there was seemingly little response to extra work. It was distressing. But later I learned that making a good impression on people would pay off long term, through recommendations for new jobs etc.
First thing it makes we want to do is qualify success rates among individuals. Eg investors. Some are quite successful, but more so relative to what you’d expect give equal randomness?
>> Some of the problems don’t matter as much if your goal for the model is just prediction, not interpretation of the model and its coefficients. But most of the time that I see the method used (including recent examples being distributed by so-called experts as part of their online teaching), the end model is indeed used for interpretation, and I have no doubt this is also the case with much published science. Further, even when the goal is only prediction, there are better methods like the Lasso, of dealing with a problem of a high number of variables.
I use this method often for prediction applications. First, it’s a sort of hyper parameter selection, so you should obviously use a holdout and test set to help you make a good choice.
Second, I often see the method dogmatically shut down like this, in favor of lasso. Yet every time I have compared the two they give similar selections — so how can one be “evil” and the other so glorified? I prefer the stepwise method though as you can visualize the benefit of adding in each additional feature. That can help to guide further feature development — a point that I’ve seen significantly lift the bottom line of enterprise scale companies.
Neat, I wasn’t aware of available data like this. I recently bought a used smart car — lots of fun, but I admit I worried it was a death trap. It’s not in this list but a google showed they are actually not much worse than average.
Love this. Very interesting that same amount of compression (samples) can give ever more accuracy if you do a bit more work in the decompression — by taking higher order fits to more of the sample points.
I submitted it because I saw a comment somewhere else about the rich being locked into unfair advantages and classes ossifying. Seems to not be true if these stats are to be believed.
Link below has some interesting plots. One shows child pedestrian deaths per 100k population going steadily down since 70s. Yet adult pedestrian deaths have recently ticked up.
And yet… the plug and play nature of many ML methods and their frequently positive impact when applied so has probably played a large part in the growth of the field.
Re temp, I’m glad we use F for daily life in the USA. The most common application I have for temp is to understand the weather and I like the 0-100 range for F as that’s the typical range for weather near me.
That all makes sense. On the other hand, next gen workhorses must often arise from people opting to do something new like this, rather than make do with the current gen workhorse.
A couple hours watching a data scientist / ml engineer at work might be a useful way to pick up the work process. I certainly learned a lot by copying my seat mate in my first job.
https://www.betonit.ai/p/cars-could-be-even-more-convenient