While this problem isn't exclusive to Claude, Claude does seem to be the most prone to it in my experience. I've had very few, if any, "WTF that's exactly what I told you not to do," experiences with other models. Codex in particular seems to be excellent at direction following and not breaking rules.
There's another layer to the non-determinism of LLM agents: what are the execution params the provider is using today?
I hate the feeling that a worn path that I've grown to trust will "do the right thing" over the last few months will suddenly start doing the wrong thing simply because an engineer at Anthropic or OpenAI found a way to save N million dollars by "optimizing" thinking token usage.
Having your "workstation" with monitors floating around you in space wherever you're sitting or standing with zero cable management. Whether you're at home in the comfy chair, at a treadmill getting your steps in, or at a hotel on a work trip.
Once the resolution and UX gets good enough a lot of people would love to have their entire office setup replaced by a portable wearable with next to zero cable management. Doubly so if that opens up space in your expensive SF apartment.
That's all good in theory, but we're still a long, long away from this being the future, let alone a future that everybody wants.
> This includes not clearing/compacting the context often. Opus now has a 1M context window, and quality is good to at least 200K. So each query is burning a lot of tokens until you clear/compact.
I see this repeated by others, including coworkers. It completely ignores caching. Caching itself is complicated, but the "longer context window = more expensive" is not 100% true and you are hampering yourself if you're not taking full advantage of large context windows.
You're talking about the "cultural collapse" of both Japan and Russia as if it was common knowledge. What exactly do you mean by this? Is this your personal opinion, or a reference to some quantifiable metric?
Japan is currently one of the hottest tourist destinations in the world. First because of the strength of the dollar vs the yen, but also because of their culture.
This is more than acceptable if it allows you to confidently send of an email in less than a minute that would otherwise take you 30 minutes of agony to write and still not be confident about.
Also, these aren't cold calls. The recipients aren't critical about how "botty" the email sounds.
ChatGPT and LLMs have had a significant impact on my wife's life. She's a second language speaker, and having ChatGPT available to draft and proofread professional sounding emails and text messages has drastically increased her self-confidence and ability to communicate with colleagues. I think that's amazing.
I feel like the fact that you are able to say this, and the sentiment echoed in other comments, is a pretty decent sign that the "movement" has peaked. It was just a few years ago that anybody voicing this kind of opinion was immediately shot down and buried on this very forum.
It will take a while for DEI to cool down in corporate settings, as that will always be lagging behind social sentiment in broader society.
I see it more as the simple truth. There's only so much influence you can wield working one day a week as an executive advisor. By his own admission he could have steered things better if he'd been more involved, but he didn't want to be more involved. He's got his own startup to work on.
Yup. I had the misfortune of being in Michigan, unemployed, at a particularly bad time. I applied to about 30-40 gas stations, movie theaters, fast food places... Everywhere I went I was told the same thing, "I'm required by law to give you this application form, but we're not going to hire you. Good luck." I didn't know anybody and couldn't get a job cleaning floors let alone flipping burgers.
> I find multi-letter variable names extremely old fasioned
Sometimes I read things here on Hacker News that throw me so hard I leave the site for a month or two. Congratulations, this time it's your fault. Goodbye.
I wasn't criticizing you. I was pointing out that an equally likely and more charitable interpretation is that you posted as a fan of AWS before you started posting as an employee.
Turns out I was wrong in this case, but you've explained the situation and everything is hunky dory.
Eh, I agree it's kind of a strange choice of words, but you don't really expect people to want to spend more time with somebody who doesn't provide any value? Just being pleasant company is useful and has value.
Agree except for the apologizing part. Especially if you're the kind of person that over-thinks things, like me.
There are two situations that I can think of where apologizing is appropriate and appreciated:
1. You did something bad, rather than just saying something. Like puking on a friend's couch. Go out of your way to make amends.
2. Immediately after you said something and realized how insensitive or offensive it was. Conversation moves fast, if it hasn't already moved on briefly retract and apologize, then let others talk for a while.
I tend to fixate on things I've said in the past that I regret. I have a rotating roster of my "most awkward moments" that my brain likes to randomly replay for me without prompting. In the past I used to go out of my way to find a way to apologize for these moments. Almost always the encounter was awkward enough to give me something new to fixate on. Most of the time they don't even remember the conversation in question.
Don't take yourself too seriously. There's a certain amount of hubris in assuming that something you said in passing deeply affected anybody else. Forgive yourself and let these small fixations go and others will too, probably much faster than you do.
Although it might annoy some users, a quick work-around might be to add a tag to the file. Tags on files/folders don't go away when you move or rename the file. Something like, "ghostnote:<ID>" where ID is some internal reference to the associated note.
In fact... using this method you wouldn't have to store the URLs of files/folders at all.
> I'd be interesting in hearing more Japanese/Swiss parents feedback on whether their country really is more liberal than the United States when it comes to letting their younger children walk to school in the presence of only another child, not older than 12 years old.
This isn't at all what you asked for, but I'm going to share it anyway, just for the sake of some heady juxtaposition. It's apparently a common sight in Iceland to see babies in strollers left unattended, outdoors, in sub-zero temperatures[1]. That definitely falls under the "you can't make this shit up," category for me. It's almost comically Viking.
The term is often used in common English with the meaning of a difficult initial learning process. Nevertheless, the Oxford English Dictionary, The American Heritage Dictionary of the English Language, and Merriam-Webster’s Collegiate Dictionary define a learning curve as the rate at which skill is acquired, so a steep increase would mean a quick increment of skill.
Arguably, the common English use is due to metaphorical interpretation of the curve as a hill to climb.
There's another layer to the non-determinism of LLM agents: what are the execution params the provider is using today?
I hate the feeling that a worn path that I've grown to trust will "do the right thing" over the last few months will suddenly start doing the wrong thing simply because an engineer at Anthropic or OpenAI found a way to save N million dollars by "optimizing" thinking token usage.