It's not a big data set that lends itself primarily to analysis, it's more like content. For example, a list of all US Presidents with a lot of metadata or text content fields about them collected/combined from different sources, cleaned, corrected, annotated, etc. (Pretend Wikipedia has only a subset of these fields and considers broadening them out of scope.)
As for Github, the data would still be under "my" account and I'm thinking about more of a platform that doesn't depend on one person. Maybe I would manage day to day version control in Github but I'd want to promote occasional releases to be more official and not reliant on my account.
In the GPT-2 era I created CouldReads, a big data set of generated book titles/synopses trained on thousands of e-books. It was a fun project in the naivete of 2020 but it's less amusing now.
A while back I wrote up a way to turn the big Wikipedia XML dump into a database. Not a generic table with articles but thousands of tables, one for each article "type". I'm not sure if this is still the best way to go about it.
That looks like a crowdsourced project for turning arbitrary sites into RSS which is very cool, but I don't see a way to get a large RSS data set out of it. And with about 5000 sources (I think) it's not as large as what I was hoping for, but it could be a good complementary source.
And also a new project to fetch all links seen in the Bluesky firehose and gather metadata to build a database of sites and pages at a more granular level than the domain. For example, is account X posting video links from one YT channel or many?
Just for fun I wanted to do a simple server-side version of this where the submissions would be truly hidden on my account, so it would take effect on mobile too. And avoid client side artifacts like messed up numbering.
If anyone's interested in an approach to processing the data set quickly, I got something working and wrote it up when I was curious about turning the content into structured data for database tables.
My company just doesn't have product or project managers and it's wonderful. Fewer meetings, no intermediaries, more agility. It could only work in practice, it could never work in theory.
I feel like I spend a troubling amount of time as a software developer dealing with and/or avoiding solutions that are much more complicated than the problems they're trying to solve.
When I take notes because I need to retain something, the writing part is more valuable than reviewing it later. And writing on paper is important, typing isn't effective for that part.
When I take notes just to record details that I can look up later on demand it's faster to type and there's no downside. But that matches work more than it does school.
At this point I already get my stuff from Amazon fast enough, and I don't even have a prime account. I'd rather see Amazon improve working conditions for its employees (if what that I've read about that is true) than chase this obsession with feeding my desire to get more material things as fast as possible.
Ironically their focus on customer satisfaction has started to make me feel dirty about buying from Amazon.
When someone gets entertainment value from playing the lottery, is it really about the chance of winning multiplied by the amount of happiness the money would bring? Or does it have more to do with the excitement that comes with each drawing? (Similar to how I can enjoy watching a football season even if my team doesn't win the Super Bowl.)
I think it's more likely the latter, though I don't personally see the entertainment value of playing the lottery. I wouldn't lose respect for someone for having a hobby I don't share or understand, though.
When the expected value is so definitively against you I guess it's implied that you're playing for variance. Maybe you also feel that you're playing for the entertainment value, which isn't subject to the same mathematical rationality.
I get that. When I play poker for 5 hours and come out even at the end of the night or slightly behind, I consider it worthwhile because I had fun for 5 hours.
When I saw the title my first instinct was, okay, what's the metaphor between a beehive and a development team or a startup? You could find some parallels, but I agree this article is better.
Maybe it's hard to find Facebook essential (as opposed to convenient) if you lived during a time when you had to actually know your friends' phone numbers and enter them manually. As I did.
Anyone you're close to you have a way to contact without Facebook. Anyone else shouldn't make you feel like you're missing anything important. Right?
I'd like to see a social news site that costs something like $25-50 to join. Enough to keep out kids, spammers, and trolls, but not enough to keep out serious adults.
I wonder if the PepsiCo Healthy Living program directed people to consume fewer PepsiCo products. It seems awkward for them to either address or ignore.
I don't know what's worse, the sense of entitlement to good grades or whatever thinking is behind the feeling your life is ruined if you can't work at Goldman Sachs.
As for Github, the data would still be under "my" account and I'm thinking about more of a platform that doesn't depend on one person. Maybe I would manage day to day version control in Github but I'd want to promote occasional releases to be more official and not reliant on my account.