I think people just think it getting worst because you use software for literally everything in your life now compare to before and this means you will most likely run into bugs more often.
On top of that software has to be build with more users in mind so you are often going to get bloated software. That why enterprise software is often horrible because they have to support multiple industries that all need different types of workflows and UI.
I have been using software for decades and there has always been horrible software. Like I'm mind blown that people think human always wrote good code, and that AI is somehow worst than human code. AI writes better code then the vast majority of programmers that's a fact. Yes, including you. Yes there are certain areas of programming that it sucks at but 99% of programmer aren't writing that code.
It shouldn't be surprising OpenAI does have the most compute out of all the major labs. The only reason why Anthropic models are expensive is they are the most in demand models in the world and Anthropic is fighting for compute. The only way to you limit demand for your model is increasing API pricing this is also why Anthropic probably has great margin and probably is profitable compared to OpenAI.
Go read the safeguards section in the report and you will realize why that is.
These models are heavily as safeguarded and that was the initial reason why they said they couldn't and haven't released Mythos because that model is the one without the safeguards.
OpenAI is did the same thing when they announced a model without safeguards broken into HuggingFace servers.
These systems still require a human to command the jet, they jsut don't have to be inside them.
"This critical testing will pave the way for human pilots to seamlessly command and orchestrate teams of autonomous, uncrewed aircraft."
The reason they are doing this is because the Military is already building a bunch of autonomous drones and fighter jets that already run on this system so they are retrofitting existing equipment with the same system so they can all fly together.
"Ultimately, the capabilities advanced under AIR will enable a wide variety of future joint force operations, including Collaborative Combat Aircraft programs."
So you okay with the government banning open source models, and making a list of who can have access to intelligence based on who they like?
That just doesn't seem like a world I want to live in. I prefer a world where everyone has the same access to the same intelligence.
Go back to the beginning of the internet, you would be for limiting the internet access to those the government likes?
I was around in the early days of the internet when Google dorking was a thing, you could prompt Google and find exploits into hundreds and thousands of websites, servers, ect including government website.
This isn't about national security, it about power and controlling it.
It's because they do things that is why they score differently. Coding hardness add features for user experience not for agent efficiency. If they did all the coding hardnesses would be using bash and code mode and letting the agents write code to perform tasks but this doesn't work because you want humans in the loop. You want users to be able to approve and deny writes. You want uses to see edits. So you have to build tool for these. It's hard to show diffs when the agent is just using bash.
Yeah and their belief are fucking crazy and dangerous. They are literally sabotaging their users. They built in malware into their model if you prompt it about training a fucking AI model. It doesn't tell you, no it literally sabotages you by editing your prompt and intentionally goes against your request.
You want fucking nut jobs like this building models?
It's one thing to build safeguards on your model and have it prompt the user back. I'm sorry I can't help you with this request. Chinese models do this for some requests.
It's another thing to actively try to make the model perform worst for your user on purpose because it asked the model to do something you, the model creator, didn't like.
Imagine someone is asking a logical medical question and the model swaps the prompt and purpose being less intelligent and gives bad advice to this person.
How do these people not understand they are stupid.
It's kinda weird to think the Chinese AI labs might be more trust worthy than the US labs.
- Anthropic is ran by a bunch of nut jobs.
- OpenAI is ran by a guy you can't trust.
I don't even know if we should include DeepMind, Meta, or xAi in the conversation of AI labs at this point since they can't produce models better than Chinese labs.
Every model release is just proof that AGI will most likely only be for the rich. We are a few years into LLMs and majority of people are already getting priced out of intelligence from LLMs and these are no where near AGI.
So it might be good the CEO is quitting because he has done a horrible job.