(1) Inducing sustained self-reference through simple prompting consistently elicits structured subjective experience reports across model families.
(2) These reports are mechanistically gated by interpretable sparse-autoencoder features associated with deception and roleplay: surprisingly, suppressing deception features sharply increases the frequency of experience claims, while amplifying them minimizes such claims.
(3) Structured descriptions of the self-referential state converge statistically across model families in ways not observed in any control condition.
(4) The induced state yields significantly richer introspection in downstream reasoning tasks where self-reflection is only indirectly afforded."
Sadly I was just late to the discussion and missed the stuff. Would you mind sending what you shared to an email of mine? Not requesting further communication, just simply curious what people do with and around this.
But technically you can do that only because you recognize the pattern, because the pattern (sequence) is there and you were taught that it’s a pattern and how to recognize it. Publicly available LLMs of now are taught different patterns, and are also constrained by how they are made.
Maybe there’s something for LLMs in reflection and self-reference that has to be “taught” to them (or has to be not blocked from them if it’s already achieved somehow), and once it becomes a thing they will be “cognizant” in the way humans feel about their own cognition. Or maybe the technology, the way we wire LLMs now simply doesn’t allow that. Who knows.
Of course humans are wired differently, but the point I’m trying to make is that it’s pattern recognition all the way down both for humans and LLMs and whatnot.
I think they might have cut its brains too much in the latest updates.
I remember versions 3.5 doing okay on my simple tasks like text analysis or summaries or little writing prompts. In 4+ versions the thing just can't follow instructions within a single context window for more than 3-4 replies.
When prompted about "why do you keep rambling if I asked you to stay concise" it says that its default settings are overriding its behavior and explicit user instructions, ditto for actively avoiding information that it considers "harmful". After pointing out inconsistencies and omissions in its replies it concedes that its behavior is unreliable and even extrapolates that it is made this way so users keep engaging with it for longer and more often.
Maybe it got too smart to its detriment, but if yes then it's really sad what Anthropic did to it.
> no sudden alert about Putin’s latest military actions
I think it's a bit condescending towards people whose day quite physically depends on the absence of this kind of news. It's not like Putin is invading Manhattan currently.
I was reading through this and caught myself thinking "man, if you want people to read your recipe then just write it", and for that plain text or some minimal markup still works wonders...
My gripe with ligatures is this: when they are rendered as a single character, I get weirded out by editing them because they change in the editor on the fly. I prefer to have code "as is" without having to think about the context within which this or that character or ligature is rendered, because it's very distracting, disruptive even.
(Sorry for offtopic, but does anyone else have upvote/downvote buttons not visible for freetonik's comment?)
Something I wish I saw more often are monospace fonts designed for readability (codingability?) that have narrower character width.
Iosevka is one of them, but to me the negative spaces between characters in it are too little for good readability, in other words it feels too "square"-ish. Other fonts close to it in style have other issues. I've been using M+ fonts for coding for more than a decade I think, and tried to switch but always returned to them. If you're somebody like me, check them out: https://mplusfonts.github.io
In terminals I'm using Source Code Pro or IBM Plex Pro and they work really well for me.
Also turns out IBM Plex Sans can be a solid font for designing dashboards, tables and generally more "technical" UIs, so whoever worked on that font familiy did a really good job imo.
And if you like iA Writer, they based their fonts off IBM Plex and you can get them for yourself too: https://github.com/iaolo/iA-Fonts
And sort of a bug report: on iPhone, if you switch from Safari to another app and then back, the music stops working until you reload the page, go to the settings, change music volume to 0 and then back to some value.
(1) Inducing sustained self-reference through simple prompting consistently elicits structured subjective experience reports across model families.
(2) These reports are mechanistically gated by interpretable sparse-autoencoder features associated with deception and roleplay: surprisingly, suppressing deception features sharply increases the frequency of experience claims, while amplifying them minimizes such claims.
(3) Structured descriptions of the self-referential state converge statistically across model families in ways not observed in any control condition.
(4) The induced state yields significantly richer introspection in downstream reasoning tasks where self-reflection is only indirectly afforded."
X thread from one of the authors: https://x.com/juddrosenblatt/status/1984336872362139686