Not to mention that the entire class of Markov Chain Monte Carlo techniques only form a subset of general uses for Markov chains.
Markov chains form the basis of n-gram language models, which are still useful today.
Markov chains are also the basis of the Page-rank algorithm.
Hidden Markov Models (which are just an extension of Markov Chains to have unobserved states) are a powerful and commonly used time series model found all over the place in industry.
In the pre-deep learning model Markov chains (and HMMs) in particular had very wide spread usage in Speech processing.
They are probably one of the most practical statistical techniques out there (out side of obvious example like linear models).
Being reputable and being a vanity division are not mutually exclusive.
You could argue that at it's peak Bell labs was a vanity division. That research may have changed the world, but very little of it likely ended up benefiting AT&T in any major way financially. It's telling that once AT&T was broken up Bell labs, while existing in some form for years after, was never reestablished.
Facebook/Meta stock lost 30% of it's value a few months back and may not recover any time soon.
Many of these people are high level, so I'm guessing (based on a quick skim of levels.fyi) about 50% or more of their TC comes from stock.
This amounts to an effective 15% pay cut in a year with record inflation.
I think most of us would leave our job over a 15% pay cut, especially if we were well established in the field. On top of this it loosens those golden handcuffs quite a bit for anyone who was on the fence about being employed by facebook, but couldn't say no to the comp.
You've never had to any kind of factor analysis in your work or done any searching for latent variables that map to customer/stakeholder question? Given the number of people I've worked with that are interested in modeling "engagement", I find this hard to believe.
PCA is an incredibly valuable tool that I've used in most jobs I've had. It's just a terrible idea as a default part of a feature engineering pipeline (which is what the author is talking about in terms of "feature selection"), for reasons outline in this article.
I suggest you don't be quite so quick to dismiss important concepts in this area, and before criticizing this post, at least read through it (I noticed your comment about misunderstand what the author is discussing by "feature selection" is the top comment here).
It's very clear if you read the article that what the author is calling "feature selection" might be better termed "feature generation". He explicitly calls out what he means in the post:
> When used for feature selection, data scientists typically regard z^p:=(z_1,…,z_p) as a feature vector than contains fewer and richer representations than the original input x for predicting a target y.
I don't even think this is necessarily incorrect terminology, especially given the author's background of working primarily for Google and the like. It's the difference between considering feature section as "choosing from a list of the provided features" vs "choosing from the set of all possible features". The author's term makes perfect sense given the latter.
PCA is used for this all the time in the field. There have been an astounding number of presentations I've seen where people start with PCA/SVD as the first round of feature transformation. I always ask "why are you doing that?" and the answer is always mumbling with shoulder shrugging.
This is a solid post and I find it odd that you try to dismiss it as either ignorant or click bait, when a quick skim of it dismisses both of these options.
Am I alone in really disliking Towards Data Science?
While their articles always look nice, their content is all written quickly by data scientists wanting to polish their resume with the ultimate aim of rapidly generating content for TDS that will match every conceivable data science related search. This post clearly exists solely so that TDS can get the top spot for "Word2vec explained" (which they have). As evidence of this tactic you can see that there already is a TDS post "Word2vec made easy" [0], offering nothing substantially different than this one.
The problem is that content is almost never useful, it just looks nice at first skim through. The authors, at no real fault of their own, are just eager novices that rarely have new perspective to add to a topic. It's not uncommon to find huge conceptual errors (or at least gaps) in the content there.
I personally encourage everyone at every level to write about what they can, but the issue is that TDS has manipulated this population of eager data scientists in order to dominant search results on nearly every single topic they can cover related to DS, which has made searching for anything tedious.
Compare this post to the fantastic work of Jay Alammar [1]. Jay's post is truly excellent, covering a lot of interesting details about word2vec and providing excellent visuals as well.
I'm assuming TDS will fold as soon as DS stops being a "hot" topic (which I think we'll be in the relatively near future), and will personally be glad to see the web rid of their low signal blog spam.
I honestly think it all boils down to numpy being developed long before matrix libraries became a standard part of software development.
Ruby's early "killer app" (remember that term?) was Rails. Even to this day there is almost no major code out there built in Ruby that isn't ultimately related to building CRUD web apps. While Ruby may be losing popularity now, it moved the web-development ecosystem ahead in the same way that Python has moved the scientific computing world ahead.
20 years ago if you wanted to use open source tools to performant vector code there was Python and a hand full of oss clones of commercial products. Given the Python was also useful for other programming tasks in a way that say Matlab/Octave is not, it was the choice for more sophisticated programmers who wanted an OSS solution and need to do scientific computing. This creates a positive feed back that persists to this day.
Given that Python remains a decent language relative to it's contemporary peers and it has a massive and still growing library of numerical computing software it is extremely unlikely to be dethroned, even by promising new languages like Julia.
Even to this day there is nothing even close to numpy in Ruby. I do DS work in an org that is almost entirely Ruby, but we still use python without question because we know re-implementing all of our numeric code into Ruby would be a fools errand.
Had ruby had early support of matrix math, it wouldn't have surprised me if it would have replaced Python.
I'm surprised this argument is so hidden in the comments. All things being equal I would prefer permanent standard time, but it's pretty obvious that a very rational way to solve this problem is just choose whichever setting we use the majority of the time.
> They point to evidence, notably from the Nordic nations, that economies can continue to grow even as carbon emissions start to come down.
This a pretty laughable piece of "evidence" given that the most important export from the Nordic countries is petroleum.
But even if this wasn't the case, you can't talk about global systems on a local scale. There are countries that have continued to reduce fossil fuel growth and continued to increase economic growth, but at the end of the day these are just accounting tricks. The US has seen a bit of this, but we also exported a lot of our energy intensive manufacturing to other countries.
I'm sure someone will point out that some studies have tried to correct for this by measuring the carbon foot print of imports. But these studies are also flawed for several reasons. First the reduced cost of good is China is not only the lower labor cost but all of the other industrial services nearby that allow those goods to be created efficiently. It's very difficult to determine how far up the chain of production you need to go.
The other is that there's never an attempt to account for how much energy/fossil fuels go into the dollars that flood into an economy. If your finance sector is making a lot of money of foreign energy consumption, that counts.
The bottom line is that global consumption of energy has never decreased for any source[0]. We even burn more wood to produce energy than we did when that was one of of the major source of home energy!
We are already hitting the limits to growth, but the reality of that is too terrifying for most people to accept so we end up convincing ourselves that its not happening... while we're watching a war over fossil fuels develop in Europe.
> That if someone really wants to understand where we are with all those Libya, Syria, Iraq, ... Ukraine.
I appreciate Mearsheimer's perspective, but for a self professed "realist", I find it strange that he talks a lot about "ideology" and very little about "oil".
At the end of the day all of those conflicts are conflicts over the control of fossil fuels. I have a hard time taking any analysis seriously that doesn't focus on this primary issue.
Putin/The Russians aren't insane, they're looking after oil/NG interests. That's the entire reason they captured Crimea in the first place. The Ukrainian government then stopped the flow of water to Crimea to make it basically unusable.
NATO backed oil/NG companies also are annoyed because, despite investing a lot in gas fields in the East and the West, they haven't been able to produce much of anything because the region is under constant military conflict, funded by both the US and Russia.
I agree with the parent comment. If Russia can finally and absolutely control gas fields in the East and have assured access to not only Crimea but water to run the region they'll be getting much more out of this situation than they had prior. They will be interested in stability as much as NATO backed energy companies. Given the conflict in the region, and the recent (2014) overthrow of a government friendly to them (coincidentally right after Shell and Chevron signed contracts to develop gas fields), Russia is, rightfully, as wary of promises from the West as we are of them.
For purely, selfish, greedy reasons Russia would want peace in the region as much as NATO backed oil/gas companies have been hoping for it. There are still gas fields in the West that are waiting to be exploited and already have their resources under contract. Once these resources can be split up in a way everyone can deal with, there won't be conflict until the gas dries up because both parties make tremendously more money with stability after that.
I second this question. In the last few days I've heard pretty heavy reference to "people cheering on nuclear war" but haven't seen it first hand anywhere. A link to an opinion piece that explicitly states this, or even a few high-profile twitter accounts espousing this view would at least convince me it's happening.
Heck, even browsing reddit comments and hacker news comments I've seen no view encouraging nuclear war. Even 4chan doesn't seem to have any posts proclaiming this view.
> In an original and persuasive analysis, Mulder shows how isolating aggressors from global commerce and finance was seen as an alternative to war that worked precisely because of the pain it imposed on the target society. From the very beginning, it was civilians who suffered the most. Nevertheless, the League of Nations embraced sanctions and established an elaborate legal and bureaucratic apparatus to enforce them. Mulder argues that instead of keeping the peace, this form of economic warfare aggravated the tensions of the 1930s, encouraging austerity and autarky and restraining smaller states but backfiring against the larger authoritarian ones, such as Italy.
I think this is an interesting point to bring up, because, especially in the West, we hold a believe that "economic" violence is not as bad as "physical" violence. It's well worth at least questioning this assumption.
I think Bishop et al. WIP book Model-Based Machine Learning[0] is a nice step in the right direction. Honestly the most important thing missing from ML that stats has is the idea that your model is a model of something. That how you construct a problem mathematically says something about how you believe the world works. Then we can ask all sorts of detailed question about "how good is this model and what does it tell me?"
I'm not sure this will ever dominate. As much as I love Bayesian approaches I sort of feel there is a push to make them ever more byzantine, recreating all of the original critiques of where frequentist stats had gone wrong. So essentially we're just seeing a different orthodoxy dominant thinking with all of the same trapping of the previous orthodoxy.
No, this is the second volume of "Probabilistic Machine Learning", the first volume of which was just published this week. The 2 volume set can be seen as a complete rewrite/replacement for "Machine Learning: A Probabilistic Perspective"
For clarification, Murphy's first book is just Machine Learning: A probabilistic perspective this is his newest, 2 volume book, Probabilistic Machine Learning which is broken down into two parts an Introduction (published March 1, 2022) and Advanced Topics (expected to be published in 2023, but draft preview available now).
To answer your question. This book is even more complete and a bit improved over the first book. I don't believe there's anything in Machine Learning that isn't well covered, or correctly omitted from Probabilistic Machine Learning. This also has the benefit of a few more years of rethinking these topics. So between the existing Murphy books, Probabilistic Machine Learning: an Introduction is probably the one you should have.
Why this over Bishop (which I'm not sure is the case)? While on the surface they are very similar (very mathematical overviews of ML from a very probability focused perspective) they function as very different books. Murphy is much more of a reference to contemporary ML. If you want to understand how most leading researchers think about and understand ML, and want a reference covering the mathematical underpinnings this is a book you really need for a reference.
Bishop is a much more opinionated book in that Bishop isn't just listing out all possible ways of thinking about a problem, but really building out a specific view of how probability relates to machine learning. If I'm going to sit down and read a book, it's going to be Bishop because he has a much stronger voice as an author and thinker. However Bishop's book is now more than 10 years old an misses out on nearly all of the major progress we've seen in deep learning. That's a lot to be missing and it won't be rectified in Bishop's perpetual WIP book [0.]
A better comparison is not Murphy to Murphy or Murphy to Bishop, but Murphy to Hastie et al. The Elements of Statistical Learning for many years was the standard reference for advanced ML stuff, especially during the brief time when GBDT and Random Forests where the hot thing (which they still are to an extent in some communities). I really enjoy EoSL but it does have a very "Stanford Statistics" (which I feel is even more aggressively Frequentist than your average Frequentist) feel to the intuitions. Murphy is really the contemporary computer science/Bayesian understanding of ML that has dominated the top research teams for the last few years. It feels much more modern and should be the replacement reference text for most people.
Kevin Murphy has done an incredible service to the ML (and Stats) community by producing such an encyclopedic work of contemporary views on ML. These books are really a much need update of the now outdated feeling "The Elements of Statistical Learning" and the logical continuation of Bishop's nearly perfect "Pattern Recognition and Machine Learning".
One thing I do find a bit surprising is that in the nearly 2000 pages covered between these two books there is almost no mention of understanding parameter variance. I get that in machine learning we typically don't care, but this is such an essential part of basic statistics I'm surprised it's not covered at all.
The closest we get is in the Inference section which is mostly interested in prediction variance. It's also surprising that in neither the section on Laplace Approximation or Fisher information does anyone call out the Cramér-Rao lower-bound which seems like a vital piece of information regarding uncertainty estimates.
This is of course a minor critique since virtual no ML books touch on these topics, it's just unfortunate that in a volume this massive we still see ML ignoring what is arguably the most useful part of what statistics has to offer to machine learning.
Markov chains form the basis of n-gram language models, which are still useful today.
Markov chains are also the basis of the Page-rank algorithm.
Hidden Markov Models (which are just an extension of Markov Chains to have unobserved states) are a powerful and commonly used time series model found all over the place in industry.
In the pre-deep learning model Markov chains (and HMMs) in particular had very wide spread usage in Speech processing.
They are probably one of the most practical statistical techniques out there (out side of obvious example like linear models).