I’ve been playing around with some similar ideas in a couple of pet projects. I truly believe that some type of “proof of work” system like this will be the only reliable way to have confidence in human writing. The various AI “detector” software that’s out there seems to be a dead end to me.
The comment I’m speaking of was from a different user, and it’s giving a strong smell that both users are bots unfortunately. Maybe I’m wrong, maybe it’s a coincidence those exact strings of characters were formed in isolation from each other in the same thread. I hope I’m wrong, but it’s a red flag
This is a great distinction that I don’t think is getting talked about enough. “Open source” is probably the wrong term to use. Open “weights”, sure.
If your only concern is how good is it at coding, then I don’t have an issue with using the Chinese models. Especially if you want to run it locally, they are kind of the only choice.
For any other use than coding, it’s going to have to be a hard pass from me.
Okay, first off, without honestly reading the whole article (I tried but I just don’t have the patience), a quick red flag is the final analysis is only through Fable? Isn’t that inherently going to introduce unwanted bias?
Whatever. Doesn’t really matter much. My next thought is, I get that requiring an SVG is adding an extra layer of complexity as far as the art goes, but why is nobody talking about that the actual art is absolute trash?
I get it. It’s basically a meme at this point and it’s a fun game to play with the models. But my thought is it should be illuminating to anyone who is an artist that LLMs are still a long way off from taking your job :)
I took the bait too and responded to this post, but it’s the third “green” screen name I’ve seen in the last couple days with the same text in the bio:
“When dang and tomhow ban all of us humans, only bots will remain”
If that’s what you took away from this, I’m afraid you missed the point. The question isn’t “will anyone actually play this?” No, they won’t.
It’s a demonstration of running web technologies (on the front end, and rust on the back) on a device that was never meant to accept them. And to run something outside of the assumed capabilities of the hardware for that era. I found this very interesting to be honest
I like this distinction “dirt notebook”. As someone who was never naturally good at taking notes, I’ve only later in life made it a consistent habit. That being said I’m realizing all of my notebooks of “dirt” notebooks. There’s no rhyme or reason to it. No consistent format. Just pure stream of consciousness trying to capture whatever was in my head at the moment, or whatever I’m trying to remember from a meeting or call or whatever.
I’ve been using Firefox on mobile (and desktop) for years. I still don’t understand all the hate thrown at them. Whatever the downsides/ shortcomings people see, they are irrelevant to me, because I hate Google more & refuse to give them my data. They are good browsers & work just a well as Chrome 99.9% of the time for me
I just mean no matter how hard anyone tries, I don’t see how useful these systems would be in practice. Sure they demonstrated a “working” system. Plenty companies sell products that “work” to one extent or another.
But how useful is it really to get a result of “This is 80% likely chance of being LLM generated”? Or 75%, or 95%? What if the text is a mix of human written text and LLM text? How would you even begin to test that?
I suppose a text that is half human half LLM would theoretically score in the 50% range, but do you see the problem? You can slap a confidence % score on a test run, but interpreting the results leads to a whole other can of worms.
Point is there are so many variables, and it’s not clear that the result from any of the systems is even valid or applicable to help you make a decision in a real life situation.
I could be wrong, but I just don’t see how trying to “detect” LLM generated texts is ever going to work. The only thing that makes any sense if you truly want to have confidence a human wrote it is some type of “proof of work“ system. I think there’s a lot of interesting ways to approach the proof of work problem with different pros and cons, but that is where our energy should be focused if we seriously want to solve this problem.
That actually makes a lot of sense, and it helps me wrap my head around why the technique exists at all. As someone who didn’t get into web programming until 2015 or so, I didn’t quite understand at first the usefulness of this. But for sites like this built in the 90’s it was a totally different world
I actually kind of liked the timer. It gave me a sense of urgency and exhilaration! Word games have never been my favorite, but this has a different feel. But I can see giving the option to turn it off, and fine tuning the timed mode. Seems like some real potential here
Yeah, I was lucky. My Samsung device is officially supported. As you said, since no one else has had the interest to do it for your Nokia, it’s basically up to you to make it work, since they rely on volunteers.
I wouldn’t begin to know how to take the base ROM and adapt it to a specific device, but I’d imagine any of the frontier LLMs could probably make quick work of it. It might be worth taking a weekend, and spending some tokens on coaching a model through the task?
I don’t disagree with what you said, but that same logic applies to every frontier model, or even lower tier models. You are subject to their creators bias whether you like it or not.
So what you are really saying is that you don’t accept SpaceXAI’s bias, and you’ll plant your flag elsewhere. It’s not that the other camps don’t have their own bias.