This reminds of sensitivity vs. specificity for medical tests. You essentially have a 2x2 matrix: Error vs. non-error on one axis, and detected vs. not detected on the other. If we assume that a reviewer LLM has 95% accuracy, meaning that it produces the correct result 95% of the time it (identifying errors as errors, and identifying non-errors as non-errors), you get the following: Of 100 cases, the output of 5 is wrong, and the output of 95 is correct. Of the 5 wrong outputs, the LLM correctly identifies 4.75 (95%) as wrong. But it will also create 4.75 false positives (5% of 95). Of the 95 correct ones, the LLM correctly identifies 90.25 (95%) as correct. Therefore it creates 0.25 false negatives. So the positive predictive value is low: When the reviewer LLM flags an error, it's actually only an error 50% of the time. On the other hand, the negative predictive value is high, so in ~99.7% of cases, there will be no error if the reviewer LLM didn't catch one.
All of that is assuming that the actual accuracy of the reviewer LLM is truly a flat 95% and that the base error rate is truly only 5%. I think that in reality, LLMs perform better with a certain type of tasks / errors and 5% error rate is just an average.
When we assume you have errors 5% of the time, and those errors are not caught 5% of the time, in my understanding, you will have errors 4.75% of the time. 100 cases -> 5 errors -> 4.75 uncaught errors.
Oh, the joy was never fully lost to me. I still listen to bands so obscure that Spotify and Youtube Music don't have them, so I have no other choice. I'm slowly reverting back to my tech use of the late 2000s and early 2010s, with the added bonus of being able to access my own music collection from anywhere. I've seriously been considering getting an MP3 player again.
Kombucha is the best "natural" soft drink for me, too. It's not entirely sugar-free, though (even though you can get it to low-sugar with longer fermentation).
Ah, good old Dragonspice.de, they have provided me with supplies for many of my experiments as well. I have many of the essential oils already, I might try this! Thank you for posting your recipe.
I agree with this observation. I often assumed it would be an USP if I highlight that I want to understand how things work instead of blindly following paradigms and frameworks, but it seems that (a sense of) uniformity just comes with too many perceived advantages, like you said.
I don't know the exact reason, but when they don't share the process or prompt, it seems like they're trying to gatekeep their results – which is very ironic from someone using a tool made possible by ingesting other people's work without their consent.
That's interesting to hear, because my impression was that software/web development is a field full of people who are self-taught or at least very enthusiastic about learning new tech. I am personally pretty undogmatic when it comes to languages or tools and I assumed most developers who care about solving problems are the same way.
Much better than last year, which is what I had been hoping for. There's still a ways to go, but it seems the worst is behind me, finally. It was a long and dark period with extremely crippling anxiety. Ironically, what helped most was to stop trying so hard to get better, but not letting myself go either. A very tricky balance to achieve that mainly hinges on your ability to talk to yourself in a kind, encouraging way. As someone who has an avid aversion to advice that seems to be superficial feel-good fluff I rejected the "be kind to yourself" concept for a while. Which, I realize, was my inner bully hijacking my logical brain making me believe I was doing something "right" by being cynical and unforgiving with myself.