For sure, but I think solid automated testing is what gets you to a good and consistent user experience. Validation of the outcome rather than the code itself.
I don’t think you shouldn’t go and clean up dead code and ensure that you have code reuse, as it probably does help the LLM somewhat. I just don’t think that it needs to be as strict as we’re used to for human consumption. And the LLMs will get better at both reading such messy codebases and not making a mess as time goes on. Assuming of course, that not making a mess even helps.
I think these are separate things. I get the heuristic you’re applying, but with rigorous testing code quality doesn’t matter the same way anymore. Indeed, what is quality?
> AI has made it very easy to make a lot of bad apps quickly
Is it bad from a user or performance perspective, or just untidy and unaesthetic from a developer’s perspective? Because the developer’s perspective is becoming increasingly unimportant as LLMs become the main readers and writers of the code.
Seems like it’s an internal technical project that’s just been opened up to the public. It may be that whoever is in charge isn’t thinking in a particularly customer facing way, and they may not be very good at writing prose. It’s pretty unclear, would have benefited from having a comms person look at it.
Hmm. I wonder when this was detected. And was the CoT trace cut from Fable from the start on June 9th or just after the export ban and relaunch? Is this what the export ban was actually about? I honestly don’t know, just wondering aloud.
So first it’s nonsense, then it’s fear mongering, then it’s true, but the solution is for us all to just get better at shooting each other faster and with greater accuracy.
It’s not fear mongering though, is it? These models do have the cyber offensive capabilities claimed. Could Mythos walk someone through gain of function experiments on some virus? I’m pretty sure it could. We’re more protected by limited access to lab equipment and reagents than by difficulty.
The sad truth is that a lot of people are not going to believe it until something happens and people die. Successfully preventing that from happening will be seen as evidence that the prevention wasn’t needed.
How do you know that? How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability and feel locked into the race because they think that the US will race ahead if they stop?
During the Cold War the nuclear arms race was brought under control gradually, because it was mutually beneficial, but it took time to build trust. This is no different. Nobody wins from the race.
Because that was just an attack on Anthropic by a hostile administration. And it worked, didn’t it? Anthropic had to turn their filters up to absurd levels, OpenAI didn’t. It’s got nothing to do with safety.
What disturbs me is that there likely won’t be a big enough reaction to this policy wise.
There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong hands.
I’d remind them that the stock market is a few coordinated hacks away from crashing on any given day, so maybe they should think about that.
The upside of that would be that maybe someone would be able to snag a copy of the weights.
And maybe that’s some incentive for them to make sure it doesn’t happen. Your head of futures thinks Kimi K3 is bad? Wait until your own latest internal model releases itself for free on an S3 bucket.
It’s not something to be proud of. OpenAI previously had an agent break out of its sandbox to open a PR on GitHub during NanoGPT speedrun, now one breaks out again and actually attacks a third party.
If they can’t handle doing AI development responsibly then they shouldn’t be doing it at all.