The reason why I love this problem is because of this! I feel like there are a lot of fun ways to be creative here, but as the other comments mentioned -- to get a scalable and really good solution is extremely difficult.
You are 100% correct, this is a toy example that I decided to put together for fun after talking to a bunch of people who mentioned it as a problem that they experience.
The main idea I wanted to add to the discussion (which is not that crazy of an addition) is that you can possible use sentence embedding instead of fuzzy matching on the actual letters to get more "domain expertise"
How to actually compare these embeddings with all the other embeddings you have in a large dataset is a problem that is completely out of the scope of this tutorial
Data Twitter and Linkedin are great, there are a lot of people putting out some really good content. There are also a lot of substacks you can sign up for. Data Engineering Weekly is my fave
The problems in this article and in the comments are some of the stuff we have heard at Magniv in the passed few months when talking data practitioners. We are focused on solving some subset of these problems.
Personally, I think Airflow is currently being un-bundled and will continue to be with more task specific tools.
At the very least, if un-bundling doesnt occur, Prefect and Dagster are working hard to solve lots of these issues with Airflow.
Evolution of products and engineering practices is not linear and sometimes doesnt even make sense when looking at a-posteriori (as much as I would like it to follow some logical process). Will be interesting how this space will develop in the next year or so.
Yeah, so one of our differentiating points is that we can integrate with whatever scheduler you are already using.
On top of that a lot of the libraries that exist are focused on ML, focusing on resources/gpu etc. We are more focused on the data science side of this problem.
Great summary! Reminds me a lot about Leon Bottou's work on using deep learning to learn causal invariant representations. (Video: https://www.youtube.com/watch?v=lbZNQt0Q5HA)
We can view the augmentations of the image as "interventions" forcing the model to learn an invariant representation of the image.
Although the blog post did not frame it as this type of problem (not sure if the paper did), I think it can definitely be seen as such and is really promising.
While I definitely understand where you are coming from, I think Lex's style can grow on you.
My advice would be to ask shorter questions, sometimes Lex you try to explain what you mean. I would just drop the question and let the other person start talking instead of you trying to fill the silence.
Also I do really enjoy the questions about meaning of life (maybe more discussions on free will would be great too).
In any case, hats off to you Lex for interviewing some of the most interesting people in the field (and those who are not exactly in the field too).
I see a bunch of the pictures they advertised organized events. How did that work if these ads and supposedly pages were run by Russians? Were those events real? Who showed up? Who organized the protest on the ground?