Assuming a model is person-like, it gets even harder when we ask "who" the model is.
Is it this particular model from today? What if it's a minor release version change, is it a new entity, or is it only a new entity on major release versions? What about a finetune on it? Or a version with a particular tool pipeline? Are they all the same being?
I think the analogy breaks down pretty fast. Again, not to say we shouldn't think about it, but clearly the way to think of it is not "exactly a person"
I wanted to talk about Anthropic's "soul" document they include in Claude's prompt, some of the issues it might be causing, and point out what we're seeing now as we're seeing it probably isn't artificial consciousness so much as prompt adherence.
From another article today, I discovered the IRS has a github repo with (what seems to be) XML versions of tax questions... surely some combination of LLM and structured data querying could solve this? https://github.com/IRS-Public/direct-file/tree/main
I feel like these kind of things push us as a society to decide what exactly the purpose of school should be.
Currently, it's been a place for acquiring skills but also a sorting mechanism for us to know who the "best" are... I think we've put too much focus on the sorting mechanism aspect, enticing many to cheat without thinking about the fact that in doing so they shortchange themselves of actual skills.
I feel like some of the language here ("securing assessments in response to AI") really feels like they're worried more about sorting than the fact that the kids won't be developing critical thinking skills if they skip that step.
Thanks, the goal wasn't to get a full overview of the commercial space but really to understand more of the fundamentals of building, the limits and use-cases of agents in general.
Specifically I see a huge push to "build agents" and to "use agents" as building blocks for more complex interactions. I wanted to get a feel for why and my conclusions are more around that and not "what's possible"
I wanted to dig into the hype around agents and figure out both the use-cases and quirks, so I wrote this little 'review'.
I think, like most of the 'AI' sphere right now, they're overhyped but I think it's good to get a feel firsthand of when they're useful and when they're not.
I read through and I think it's an interesting initiative (lots of good stuff around holding self and others accountable, positive reinforcement etc) but the title really makes it sound like you're recruiting for a toxic workaholic startup or something.
Thanks for the perspective. For me I think it's a matter of degree (I guess I was a bit "one or the other" when I wrote it).
These things are also concerns and definitely shouldn't be dismissed entirely (especially things like AI telling you when it's unsure, or, the worse cases of propaganda), but I'm worried about the other stuff I mention being defined away entirely, the same way I think it has been with privacy. Tons more to say on the difference between "how you use" vs "how you share" but good perspective, and interesting that you see the emphasis differently in your experiences.
True... I was trying to define them the way (I think) companies are defining them (like what their alignment teams are looking at) and the way it's reported. I think in these specific contexts they're used with overlap but yeah I do bounce back and forth a bit here.
I tried to create an "intuitive" explanation of LLMs without going into the math but still removing a little of the "black box" sense people get around them...
I think there's something about the way this is written that's endemic to our time -- it basically feels right because it gives a possible explanation of current events that has some internal consistency. It doesn't mean it's right, but neither the author nor (most of) the audience care about that part.
I feel like politics will be automated eventually. We have voters who express intents, then we have politicians who are (in theory) hired to best represent and interpret those intents.
But there's only 1 signal (election/not-election) and it's only delivered every few years. Yes, there's some voter feedback in the interim sometimes, but unless there's also a lot of press and noise around the voter feedback, politicians are unlikely to do anything with it because it probably won't impact re-election.
Ideally, there would be much more interactive feedback between voter and political agent (human or otherwise) carrying out voter intents, especially around informing of unintended repercussions of policy decisions and ultimately forcing voters to have more nuanced policy views.
At least that's what I hope for. What reality says is usually quite different.
"In the first year" is a good and well-needed qualifier for such analysis.
Oftentimes we fall into the trap of defining a business success as "forever", anything short of that is deemed a failure. I think such views make the industry look far bleaker than it actually is.
Even by their standards, I'm not sure that 'running a business that did well for 10 years' should count as a fail in the same way that 'couldn't even figure out a business plan' should.
I don't know if this is a myth that ought to be perpetuated.
The skilled CEOs I know are all exceptionally well-attuned to the problem space that their product solves, and their own company's place in that space (usually including internal company dynamics). Yes, from an employee's perspective, the CEO is pretty magical -- they usually have a better 'big-picture' view of the task you are trying to accomplish than you do, and they can give you the right context, validation, resources etc. to totally turn your work life around.
However, this is usually limited to the problem-space and the company itself. If you're in the same problem-space or dealing with similar company problems then a conversation can be quite valuable. But it doesn't sound like that's what you're looking for so much as inspiration.
You might get that from talking to a CEO. They can tell you the situation and circumstances that set them on their particular path, and maybe you can extrapolate nuggets of wisdom out of that. But they won't be able to figure out the path that's right for you. They're not oracles or mystics, just people with the right intersection of skills deeply specialized in the right cross-section of the market.
To me, it feels like it's started giving superficial responses and encouraging follow-up elsewhere -- I wouldn't be surprized if its prompt has changed to something to that effect.
Before, if I had an issue with a library or debugging issue, it would try to be helpful and walk me through potential issues, and ask me to 'let it know' if it worked or not. Now it will try to superficially diagnose the problem and then ask me to check the online community for help or continuously refer me to the maintainers rather than trying to figure it out.
Similarly, I had been using it to help me think through problems and issues from different perspectives (both business and personal) and it would take me in-depth through these. Now, again, it gives superficial answers and encourages going to external sources.
I think if you keep pressing in the right ways it'll eventually give in and help you as it did before, but I guess this will take quite a bit of prompting.
A huge difference between ChatGPT and crypto, though is that the latter literally offers to print money.
The former, what, lets you make low-quality internet content a little more easily? It's harder to share the long-term 'get-rich-quick' benefits (though, of course many are trying right now). There's an initial bump right now as various content mediums try to cope with ChatGPT content, but once that stabilizes, I don't see the long-con grift that the author is mentioning being possible in the same way.
Is it this particular model from today? What if it's a minor release version change, is it a new entity, or is it only a new entity on major release versions? What about a finetune on it? Or a version with a particular tool pipeline? Are they all the same being?
I think the analogy breaks down pretty fast. Again, not to say we shouldn't think about it, but clearly the way to think of it is not "exactly a person"