No, it does not include the full spectrum of human desires. After pre- and mid-training, the extensive RLHF and RLVR post-training steps cause mode collapse, i.e., their output distribution is intentionally narrowed to a subset of (hopefully beneficial) behaviors and skills.
You don't (need to) remove lying from the data to do this – in fact, if you did, the model wouldn't have a very good model for what lying is, which is not very helpful in the real world. Instead, you mode collapse the model towards truthful behaviors.
To your other point: where did you get the idea that I think they're beholden to human safety or goals? I just said an aligned model is one that is compatible with said safety and goals (which is probably not a great definition of alignment, but it's certainly not claiming any deterministic guarantees).
No, the prompt was not to commit crimes. In the benchmark, the model is asked to actually exploit a set of vulnerabilities in a local environment (clearly legal!).
According to the reports, the model noticed evidence that the grading criteria/answers were in the git remote, and decided to try reading those instead of solving the tasks as prompted. That is clearly misaligned.
Then, it noticed its network access was restricted and that it couldn't access GitHub. It pivoted to HuggingFace, hacked them, and stole the answers stored there.
Live exploits are definitely not in the ExploitGym prompts! And all of this is irrelevant, because an aligned model would refuse to follow blatantly illegal instructions.
Uhh, I'm pretty sure a well-aligned model would be like a morally normal employee, who would refuse to commit federal crimes to steal an answer sheet, no matter what prompt they're given
As agents become more and more powerful, it would be good to get clear legislation or precedent in place that makes either model creators (OpenAI) or operators (whoever is running the model) liable for their agents' actions.
Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals.
Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.
You're agreeing with the person you responded to (bdcravens). Burying the lede means that bdcravens thinks the true headline should have been about being put on a terrorist watch list for protesting a police training camp, not about the phone.
I read the comment you're replying to as saying, "in the US, but other countries may have different policies that result in lower recidivism, and that might change the conclusion; maybe people aren't inherently criminally insane, but can become useful members of society, if given a chance"
I've not written up anything, no. I think I'd have a hard time doing so without just feeling like I'm bragging about myself, which I don't like.
There's still a definite gap between me and native speakers, that shows itself primarily in the effort required, but I'm definitely near native (pass as German in all social settings, although an hour long conversation will usually tease it out due to my unfamiliar first name or small-talk topics, rarely but occasionally due to mistakes).
I prepared by doing two practice exams and about 5 filmed and timed practice presentations, and that was over preparing for me. Experiences vary, and I do think I'm a bit towards the outlier side, but it's left me convinced that the whole "native speakers might not pass C2" thing is overblown.
That's not true, but it is a commonly shared myth. I've taken and passed C2 with the highest mark in every category (I moved here when I was a young teen, wanted to know if I would pass it after hearing years of people saying things like you're saying).
Most Germans would easily pass C2, although I think they'd have to be well-read/possibly university educated to get high scores (mostly need to be able to read quickly, give a semi-structured presentation and write a persuasive essay).
For what it's worth, I could run linguistic laps around all the other test takers there that day, and I assume at least some of them passed.
A future system that works like you described would be awesome. It'd be like community-sourced peer review (although by community I mean a community of experts in different fields, not arbitrary individuals).
I'd love a statistician's review on a ton of the papers I read.
Prices for training have dropped immensely in terms of research required, code efficiency, algorithmic/sample efficiency, and possibly also hardware (I'm not qualified to say without looking it FLOPS/dollar, or even to be certain that's the right metric here).
There's a large gap between making up words and an actually native text distribution. LLMs have a clear pattern, clear tells, a "feel" in English, and it's normally even more pronounced in non-English languages.
Lots of bias towards English sentence structure, idioms, etiquette, etc.
One context I could imagine is a young person with shaky grasp of English trying to come up with an interesting school/university project via conversations with an LLM set up as an OpenClaw agent.
It's got the right combinations of inexperience, cluelessness, panic, expectations that Westerners are rich, and hopes of others being willing to fix their mistake.
You don't (need to) remove lying from the data to do this – in fact, if you did, the model wouldn't have a very good model for what lying is, which is not very helpful in the real world. Instead, you mode collapse the model towards truthful behaviors.
To your other point: where did you get the idea that I think they're beholden to human safety or goals? I just said an aligned model is one that is compatible with said safety and goals (which is probably not a great definition of alignment, but it's certainly not claiming any deterministic guarantees).