> It's theoretically possible for a bad actor to embed hidden adversarial behavior in a model. But if this happens, it serves the interests of responsible actors to find these exploits as soon as possible, and the best way to do this is to let anyone who wants to inspect them.
This is a bad argument. It isn't trivial to tell if the weights have poisoned:
I'm not arguing in favor of closed-AI; I'm simply saying poisoning may be subtle.
I was gonna say that at least the frontier labs may be motivated to not-poison their own models, but then Anthropic just attempted to poison Fable's LLM training capability so... sigh. I'll try not to derail.
No one is saying you can't have an opinion and you can't express it.
I'm just saying there's a predictable result when you express it with the level of detail and amount of effort that you _did_. And frankly, your comments are no better than pre-reasoning era LLM rage-bait.
What level of engagement are you looking for here; support for your lack of citations, or "yeah that's also my personal experience rah rah"?
I'm having a meta-level discussion, if you can't tell. I'm not "putting words in your mouth", I'm trying to discuss: the quality of your discussion. I'm discussing the quality of your arguments and your evidence. If you think that "racial profiling" is too hot, substitute in something else; that's not the point.
My man, if your comments can't be distinguished from a bot's, you're no better than one. Also if you can't tell that your comment's unsubstantiated bait, you really need to go touch grass.
Everyone has an annecdote. "it's common enough to be recognized as a trend" is equally justification for racial profiling, and at least racial crime statistics are easily citable. And you still haven't even put forth that modicum of effort.
> If you can't generate a report, it may be because you don’t have Memory turned on.
Has anyone found success with memory and Claude Code? I found Opus's 4.8 memories largely lacking value. Memories are too verbose/specific/yet-generic for the value captured.
Instead, I've been holding "retros" with the agent immediately after a session, and _those_ responses have been typically fantastic. It has ideas for coding changes, spots unmentioned small bugs, suggests invariants, principles to adopt, lint rules, tooling tweaks, skill-file updates, follow up work, all kinds of stuff.
Hah, 2 minutes before my remark. I feel like if anyone has 2+ ergomech keyboards, they should really try something with a pointing device. It's clear they don't mind having an expensive hobby :D
For people who like small keyboards, why do you NOT have an integrated trackball/pointing device? E.g. I love my Charybdis, but I rarely see people advocating for a keyboard with a pointing device. Does a trackball disqualify it as "small"?
I heckin love his /grill-me skill. Terse, to the point, and delivers outsized results.
Gonna take a moment to share my own generic "retro" prompt, which has found many areas of improvement IME.
> Let's conclude with a retro. Did you run into any issues during this session that you think could be improved? Any failed tool calls, confusing docs/prompts, or tricky wording that took you effort to figure out, etc? Any final thoughts that you want to raise? Anything minor you didn't mention? Help make this codebase easier for the next agent to work in.
It's somewhat doc-focused since I'm currently working on fairly dense design docs... but you can easily customize it for your own needs.
This prompt reveals how absolutely _ass_ the Claude Code harness is (so many stupid tool call failures), but not much I can do about that.
Feature request; `git diff --color-moved` uses colors to display moved chunks of code. Scanning https://diffs.com/docs it isn't obvious that yall support that; please add it :)
* Unclear TOS, citing Matt Pocock who sells a course on Claude (and therefore his interests are aligned with Anthropic):
> I have never before experienced, from any developer tool, such a frustrating lack of clarity over the basic terms of usage. I personally asked, 3 weeks ago, and have received nothing but delays. The recent @bcherny announcement did absolutely nothing to clarify things.
* And finally refusing to help a user debug issues with Dropbox, repeatedly saying "I'm best suited for software engineering" and "this is a question for Dropbox support" (paraphrasing Claude).
In the 14 hours since that flagged post, the OP has changed _nothing_ in that repo to address the feedback people gave him, and instead decided to just resubmit his proj to HN.