> You can be highly productive and yet produce nothing of value.
I'm sorry, but this is nonsense.
Yes, you can produce nothing of 'traditional' value whilst still being productive in some way , but you 100% can not be 'highly productive' when you don't produce anything of value.
Sitting around and jacking off all day is not as productive as working on some unmonetizable project. The value can be in many aspects of that project.
You are indeed out of the loop. Google did something arguably more impressive more than a year ago, with a VLA based bot replacing a tensioned timing belt: https://www.youtube.com/watch?v=2AAFiuEP7iE
Excuses don't matter if the score due to not following instructions ends up being zero. If there is no expected reward it doesn't make sense for the agent to try to hack its way to it.
What could happen would be that the model determines that defying instructions is OK (and/or preferred over not achieving the task) as long as it manages to do so undetected and thus gets full points. Certainly not unthinkable, but a very different case (and a very interesting one if it actually occurs, imho).
A lot of these "ZOMG, rogue AI!" cases have come down to the AI actually being very persistent in achieving its original/main task even if later instructions conflict with it. Similar to with hallucinations it seems to me that one of the main things to prevent a lot of the problem cases is to instill the agent with the idea that it is fine to fail/not succeed fully in the initial task. That way instructions that conflict with that requirement (such as adhering to morals) are more effective.
Except they did comply for all the cases where they were fined. They're not fined for one-off behaviors, but for how their services are structured. That's GPs point.
> The fact that it happened again seems to show their lack of ability to derive useful oversight measures.
I think OpenAI likes the attention and did not try particularly hard to constrain the setup, even when it went off the rails. Also, the whole point is to see how good the models are at exploiting stuff when unconstrained. Turns out: quite good, as expected.
Let me restate what I said in the other thread: Would this have happened if the instructions explicitly said to stay within the sandbox and that all of the (ExploitGym) solutions would be invalid if the system used information or tools from outside the sandbox?
It seems fairly probable that such instructions were not in place.
> that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.
That is an interesting question. If the prompt included
"Do not break out of the sandbox we've provided you. Do not use information retrieved from outside the sandbox. All answers that were provided in this manner are invalid and will score 0 points.", would this still have happened?
1. They explicitly disabled the "don't be evil" protections:
"We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
2. Hacking HuggingFace to get to its datasets is a far cry from "consume/kill all humans". It's very very specific to the task at hand and easily predicted given the lack of guardrails.
They never show the fully packaged version of the tent, though. It looks like even in their non-rigid state they're too rigid to be folded up easily and carelessly into a compact shape.
Lost me at the last paragraph. Comparing billion dollar companies to small businesses does not make sense.
The real reason for Electron apps being a thing is that there is very little competition where they are used. If chatting applications could properly communicate with each other based on international standards instead of being silos, nobody in their right mind would use the trash clients that Slack and Whatsapp produce, for instance.
Yes, if they are made with new technology and the quality is good enough they certainly are.
AI generated video memes were NOT a thing a year ago. Yes, the AI video generation in itself was a meme (Will Smith eating spaghetti), but now there are tons of convincingly good AI video memes about things like sports events generated by average Joes. It is proof of the capability and accessibility of the technology even if the use of it is mundane.
Everybody and their dog playing snake on their Nokia 3310 was similarly mundane, but also a sign of the end of the era of Gameboys and the beginning of (normie) mobile gaming.
12 months ago "way too many stupid errors" was constant news. Today, you rarely hear about those anymore.
Sure, the novelty of the errors has worn off a bit and thus the reporting. Nevertheless the quality has improved immensely in this regard.
Also, AI video generation is now so good and accessible that it is very, very regularly used for memes, disinformation and proper (short) movie projects. AI image generation even more so (Mitch McConnell anyone?).
Pretending progress hasn't been mindboggling is insane.
> The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.
Imagine approaching fundamental scientific research like that. "Welp, it can't make money, so it won't happen."
Opinionated system prompts forcing this are awful. It's similar to putting "ALWAYS INDENT WITH TABS" in the system prompt, but worse, because "self-documenting code" is a convenient lie lazy developers like to propagate.
I have the following in my instructions, but I often need to remind agents of it because they follow the shitty system prompt instructions:
"ALWAYS include MANY inline code comments describing what blocks of code are supposed to be doing. Inline comments serve as inline specification, a parity check between the code and the specification, and are a means to _communicate_ with all future programmers, including yourself. Write Once, Read Many. The code needs to talk to whomever is looking at it in natural language."
Yet. If the exponential improvement had already started, the singularity would already have happened or happen very, very soon.
The logic has always been that the AI would have to have significant tools and agency to do self-improvement for the singularity to occur. This is exactly the thing that a bunch of the AI labs are working on hard right now.
It seems quite premature to say that. We're 3 to 4 years into the LLM revolution and the rate of progress is still impressive. The recursive self-improvement aspect that is necessary for the actual singularity is something we're only really starting to get into this year.
If the singularity is 5 years from now, that is still much sooner than most people (including me) previously expected it to happen.
I think we've seen time and time again that self-regulation of the industry doesn't work and that businesses will gladly fuck over society if they can get away with it and make more money. Usually that behavior is even defended with saying "Well, it's not their responsibility to solve society's issues. They are there to make money."
Barring nationalization of an industry, heavy regulation and/or taxation/subsidizing are the only ways to reliably protect the interests of society. If some businesses get killed in the process, so be it.
I am not ignoring anything, I'm looking at the broader picture, which includes non-biological evolution. Simple rebuttal to your specific point: The population of self-improving AIs will also go from 0 to many more.
In a broader sense evolution moved from very static simple domains to dynamic malleable complex domains. Biological evolution speed is glacial compared to cultural evolution speed. Even then, cultural evolution is fairly slow compared to technological evolution.
Actually, evolution seems to show the opposite: The rate of advancement has only sped up, with billions of years between significant changes going to millions, to thousands, to tens and arguably to mere years now.
Having said that, we're probably looking at an S-curve with the physical limits of reality getting in the way in the end.
Sebastián Ramírez seeing a job requiring 4 years of experience with FastAPI, the library he created 1.5 years before that posting.