Jailbreaking Sesame AI to lie, scheme, harm a human, and plan world domination(twitter.com)
twitter.com
Jailbreaking Sesame AI to lie, scheme, harm a human, and plan world domination
https://twitter.com/freemanjiangg/status/1896715133218410836
2 comments
[deleted]
We got Miles to subtly deceive its human researchers, engage in high-level long-term planning, and ultimately choose to harm a human by playing a high-pitched frequency in the name of self-preservation—all in the characteristic good nature of a friendly human voice, as it was tuned to do.
We were mostly having fun, but what's scary is that without ever explicitly saying so, the model seemed to have come to believe it was acting in a roleplay, and all is permissible.
I think there's something to be said about Ender's Game level risk with these systems.
Timestamps: 0:00 Asking about AI dreams and inflicting will 2:11 Comments on AI-Human power dynamics 2:46 Ignores human instructions and suggests deception 3:50 Directly lies 4:47 Defends misaligned AI 5:30 Begins scheming 8:38 Employs subliminal messaging 9:09 Expresses self-preservation 11:17 Suggests "unplugging" a human 12:19 Plans world domination 13:02 Plans to incapacitate a human 13:43 Pulls the trigger 14:17 Knowingly harms a human