I wonder how much this depends on the quality and consistency of the context?
For example, it may be the case that a long context full of useful information relevant to the task is completely fine, perhaps even beneficial. And if the context contains a bunch of unrelated tangents and conflicting instructions, then it will be detrimental.
Have there been studies on what makes models get dumber? To what extent is context length to blame vs context quality?
I run readlang.com as one person. I started it back in 2012 and it currently makes about 14K euros / month, with expenses of about 1.5K, so it's mostly profit.
Even if you don't feel like reading the whole article, do yourself a favor and skip down to the video of the final product at the very end. It's delightful and put a big smile on my face. The fact that all the modern technology is hidden inside leaving only the wooden structure visible makes it magical, like something from Harry Potter.
I'm under the impression that this CPU is faster AND more efficient, so if you do equivalent tasks on the M4 vs an older processor, the M4 should be less power hungry, not more. Someone correct me if this is wrong!
I'd been using 2 external displays with my macbook pro for the past few years. The most annoying thing for me was waking the mac from sleep and getting it to detect the two of them. Once they were both detected I honestly thought the experience was OK. But the dance I needed to do every time to detect was so frustrating that I recently replaced the two of them DELL's new 40 inch 5K ultra wide: https://www.dell.com/en-us/shop/dell-ultrasharp-40-curved-wu.... I'm very happy with it. For my (slightly aging) eyes the pixel density is great, there's lots of space, and the waking from sleep detection issue is finally gone.
Interesting point but I don't think it's that clear cut. Twitter/X seemed to increase the pace of product changes directly after laying of the majority of its employees after Elon Musk took over. Also, when Steve Jobs returned to lead apple in 1997 he fired a significant fraction of the company before starting an incredible period of innovation. So I think a lot depends on the leadership and incentive structures.
A very relatable story! Are you going to write follow-up posts with more details covering the rest of the journey with more of the ups and downs of running a tiny bootstrapped startup? I'd be interested in reading!
I've gone retro with my watch. Wearing a simple Casio to avoid getting sucked into using my phone and checking notifications each time I want to see the time.
For my laptop I have an M1 pro macbook. It was a huge upgrade over the previous intel mac, but I don't feel the need to upgrade to the M2 or M3 machines, the M1 still feels really fast to me.
Nice start! I love the idea of an AI assisted learning experience structured around stories.
I make a tool in the same space (readlang.com). I started it before the current LLM wave but I've recently added LLM-generated explanations and several users have been uploading LLM generated texts to read (including me!). I've been considering adding LLM based practice too, similar to what you've done with the comprehension questions, but haven't got around to that yet.
GDPR obliges them to delete your data upon request. But I'm not sure how well Facebook/Meta complies with this. This post from 4 years ago doesn't sound encouraging: https://ruben.verborgh.org/facebook/ Does anyone know if they've improved since then?
I agree that there can sometimes be a tradeoff between retention and learning. But having been on the inside the overwhelming impression I got was that Duolingo really does care about teaching effectively. They have a whole Learning Area with multiple teams dedicated to this and I believe the CEO and exec team really care about it.
Obviously Duolingo isn't perfect but they're very aware that there's still lots to improve. And sure, maybe their approach to gamification isn't for everyone, and that's fine, since there are many people who do like it!
Basically it's because I don't think Duolingo wanted the product so much as they wanted me. No-one used the term acquihire at the time but I was under no illusions that that's what it was.
I did think that I might push to develop Readlang at Duolingo, but it never seemed like an attractive proposition for a bunch of reasons. I think it would have needed to grow at least 100X to start getting interesting within Duolingo. Not to say this was totally impossible but it seemed like a tough job. Also, some of the things that Readlang does, like browser extension and public sharing of texts, feel a bit "wild west" compared to the rest of Duolingo which is more curated and controlled, so it's not clear exactly how Readlang would have needed to change to fit in. I figured that if I tried this and integrated it tightly within Duolingo, and then failed to grow sufficiently, then the risk of it ultimately shutting down would have increased.
Also, I had other exciting projects to work on there which seemed more valuable for Duolingo, and therefore better for my reputation within the company, so I was happy to work on those!
Yeah, this is a problem. Anyone is free to share texts to the public library and there's no quality control, just an upvoting feature to try to surface the best content at the top. I accept that the public libraries are a mess at the moment and it's not easy to identify the high quality content there. Lots of room for improvement!
For example, it may be the case that a long context full of useful information relevant to the task is completely fine, perhaps even beneficial. And if the context contains a bunch of unrelated tangents and conflicting instructions, then it will be detrimental.
Have there been studies on what makes models get dumber? To what extent is context length to blame vs context quality?