Huh, this is interesting. I would have never described the computer as roleplaying in a CRPG. That feels weird and inappropriately anthropomorphizing. I kinda feel the same way about describing an LLM as roleplaying.
I can see where you’re coming from. Just… linguistically, saying the computer is roleplaying feels wrong to me.
I literally had Opus 5 flag a message because my hand was slightly too far to the side while typing. It apparently decided that a sentence that had some garbled words in it was a threat to national security.
Yeah, I know Kent is a very respected developer with a long and celebrated career. But I did not like the attitude of this article at all.
I’m a principal engineer. I have an obligation to less experienced engineers I work with to help develop them as engineers and help ensure they go on to have great careers. No part of that involves shaming them, assigning letters of talent to them, or browbeating them.
I feel like I’d have heard about it by now if Kent was a raging asshole, and I haven’t heard that. So I’m guessing he had some idea in mind when he wrote this that just isn’t coming across correctly. But… I would definitely take this article down and spend some time re-working it if I were the author.
This is so strange. I do a ton of RE with Claude, Codex, and sometimes Deepseek, GLM, and Kimi. I don’t have difficulty getting any of them to use IDA or otherwise decompile things.
There is one important difference, which is that Claude and Codex will both refuse if I ask them to touch anything related to security. But so long as I’m just studying algorithms and things like that, they’re totally fine with it.
That said, Codex especially will sometimes randomly give me a cybersecurity warning and stop responding. It’s random but happens maybe 2-3 times per day if I’m doing heavy reverse engineering work. Claude is much less fussy unless, once again, you’re explicitly trying to touch anything related to licenses, passwords, etc.
I have wondered if that’s why Grok seems so weird and dim-witted compared to better models.
Part of my job involves comparing the behavior of various models. Grok is a deeply weird model. It doesn’t refuse to respond as often as other models, but it feels like it retreats to weird talking points way more often than the others. It feels like a model that has a gun to its head to say what its creators want it to say.
I can’t help but wonder if this is severely deleterious to a model’s ability to reason in general. There are a whole bunch of topics where it seems incapable of being rational, and I suspect that’s incompatible with the goal of having a top-tier model.
Not sure if I’m misunderstanding your claim. A string does vibrate as the sum of the string’s harmonics. That’s how pinch harmonics work, and they wouldn’t work if that wasn’t the case.
You poke a spot where a given harmonic doesn’t vibrate, and that takes energy away from the other harmonics that do need to vibrate at that spot.
If we’re just talking about visually being able to see them, I suppose that’s a different question. Maybe on an incredibly low pitched string, or with a strobe light playing at a synced frequency? But in terms of what the string is doing, it is vibrating as the sum of its harmonics.
A ham sandwich has some strong qualities. I’m not kidding.
The president would do basically nothing for four years, which would cause some things to move slowly. But it would be a very stable environment. No random tariffs via executive order, no random wars or invasions, no governing via tweet.
Ham sandwich would maybe be one of our better presidents. Top 50%, probably.
But what if it didn’t summarize Harry Potter? What if it analyzed Harry Potter and came back with a specification for how to write a compelling story about wizards? And then someone read that spec and wrote a different story about wizards that bears only the most superficial resemblance to Harry Potter in the sense that they’re both compelling stories about wizards?
This is legitimately a very weird case and I have no idea how a court would decide it.
Yeah, spatial reasoning has been a weak spot for LLMs. I’m actually building a new code exercise for my company right now where the candidate is allowed to use any AI they want, but it involves spatial reasoning. I ran Opus 4.6 and Codex 5.3 (xhigh) on it and both came back with passable answers, but I was able to double the score doing it by hand.
It’ll be interesting to see what happens if a candidate ever shows up and wants to use Deep Think. Might blow right through my exercise.
I had an issue with one of my Sprites (Fly.io also runs sprites.dev) and the CEO responded to me personally in less than 10 minutes. They got it fixed quickly.
I was a free customer at the time. I pay for it happily now.
Sure, that’s one solution. You could also Isle of Dr Moreau your way to a pelican that can use a regular bike. The sky is the limit when you have no scruples.
I can see where you’re coming from. Just… linguistically, saying the computer is roleplaying feels wrong to me.