Show me a Gary Marcus essay, I’ll show you a few new LLM “gotchas” that will be fixed by the next version. Season to taste with self-assured confidence that all these tech goobers really don’t understand how totally overrated AI progress is.
So it has been for 10+ years, so it will be at least 5 more.
It's hard to reconcile how 59% of devs in their survey are "confident" AI is improving their code quality, with prior empirical research that shows a surge in added & copy/pasted lines w/ a corresponding drop in moved (refactored) lines https://www.gitclear.com/ai_assistant_code_quality_2025_rese...
My experience (using a mix of Copilot & Cursor through every day) is that AI has become very capable of solving problems of low-to-intermediate complexity. But it requires extreme discipline to vet the code afterward for the FUD and unnecessary artifacts that sneak in alongside the "essential" code. These extra artifacts/FUD are to my mind the core of what will make AI-generated code more difficult to maintain than human-authored code in the long-term.
Wish more people would report their placebo experiments. I have periodically run them on myself, and am consistently surprised how I have been unable to differentiate substances I thought were helpful (adderall, kratom) end up indistinguishable from placebo over 20+ trials. I guess that my main takeaway was that it is hard to pinpoint when subtle drugs work. My second takeaway was that the data I generated would prob be useful to share, but with no examples anyone cared I opted against. This story inspires me to potentially revisit some of my past placebo tests and show my data.
Original research author here. It's exciting to find so many thinking about long-term code quality! The 2023 increase in churned & duplicated (aka copy/pasted) code, alongside the reduction in moved code, was certainly beyond what we expected to find.
We hope it leads dev teams, and AI Assistant builders, to adopt measurement & incentives that promote reused code over newly added code. Especially for those poor teams whose managers think LoC should be a component of performance evaluations (around 1 in 3, according to GH research), the current generation of code assistants make it dangerously easy to hit tab, commit, and seed future tech debt. As Adam Tornhill eloquently put it on Twitter, "the main challenge with AI assisted programming is that it becomes so easy to generate a lot of code that shouldn't have been written in the first place."
That said, our research significance is currently limited in that it does not directly measure what code was AI-authored -- it only charts the correlation between code quality over the last 4 years and the proliferation of AI Assistants. We hope GitHub (or other AI Assistant companies) will consider partnering with us on follow-up research to directly measure code quality differences in code that is "completely AI suggested," "AI suggested with human change," and "written from scratch." We would also like the next iteration of our research to directly measure how bug frequency is changing with AI usage. If anyone has other ideas for what they'd like to see measured, we welcome suggestions! We endeavor to publish a new research paper every ~2 months.
Sure would be nice if Firefox desktop would join the browsers that support PWAs. We build an app that has been PWA-first, but it is unfortunate that this generally requires users to have a Chrome instance running. Would much rather point people to Firefox, and it seems like it would be to their advantage to give apps a reason to recommend FF, if they built a smoother PWA integration than Chrome.
I had never specifically considered the distinction between "intelligence" and "consciousness," but after reading this, I'd agree that it may be an important distinction warranting consideration.
My sense after reading the article and the wikipedia "Hard problem of consciousness" page is that consciousness is an evolved biological phenomenon that makes intelligent entities stateful. Being stateful is very useful from an evolutionary standpoint. Firstly it lets us pick through & prioritize long-term state. Secondly because it imbues the will to continue maintaining said state. The second is what makes consciousness a dicey prospect to aspire to w/ AI.
While the author asserts that consciousness and intelligence are separate concepts that can develop separately, I'm less sure. It seems plausible that intelligence/problem solving is always improved by having state. And more durable with state. That's presumably why evolution brought it along.
But we already have AI agents like Bard that build a history of responses akin to state. If "recognizing consciousness" is what happens when the speaker becomes aware of their state and able to meta-optimize it, then it seems that consciousness wouldn't ever travel too far behind intelligence.
Wow, a much more vitriolic first comment than I would expect from HN. The last six months has not lacked for Musk-hating enthusiasts howling “Twitter is dead.” But when I look at Google Trends, its story is that interest in Twitter is almost identical to where it was 5 years ago. Doesn’t seem so dead to me?
Have Musks changes have been net positive or negative? No shortage of internet opinion on that Q. Regardless of one’s personal feelings on it, it seems hard to dispute that Twitter moves faster than before. I consider that no small feat given how much legacy code and bureaucracy the company had at the time of his acquisition.
If he can free more of his time by making this hire, I am looking forward to seeing whether that translates to even faster iteration times. pg and all the others I followed are still there, so still plenty of potential to create quality entertainment/learning w less regret than Facebook
The "gaps that seem obvious"-notion describes exactly how I feel about pull request tooling these days. The status quo for PR review has very obvious-seeming improvements that have not been pursued (ex: de-emphasizing moved code vs. deleted code, making AI predict what comment will be left before dev starts typing, auto-reviewing trivial changes).
If I'm correct that PR review is currently much less efficient than it will be soon, it won't be because I'm smart or this is a new idea. It would just be because our company has spent last five years building code review tools, right place right time. Eventually there was enough infrastructure accumulated (and ambient events unfolding, ie OpenAI) that it became a small step to pass the edge of what PR review meant circa 2022.
As has often been Stripe's way as a company, they are setting the bar for what other companies should strive toward. Has there ever been a more generous severance package posted to HN?
As the owner of a (much, much) smaller company, I'm inspired by how the Collisons run their business, especially under adverse circumstances. Yes, they fucked up in estimating the future market, but they are in good company among CEOs and non-CEOs lately.
Interesting article, but could have used a definition of “F&B”. I presume it’s probably “food and beverage,” but it drives me crazy when an article uses an acronym (in this case, about 50x) without ever defining it.
Between Obsidian, Roam, Amplenote, and Reflect it has certainly been a golden age for note taking over the last few years. It's hard to remember that it was only 5 years ago that second generation note apps like Evernote, Notion and Bear were the only viable options unless you wanted a 1st gen app like OneNote or Workflowy.
What might be most interesting about the new set of fast moving note apps is that all seem to be built by teams of 3 or less people. Obsidian seems to have ascended to the top of the heap with a team of three and no apparent VC funding. Anyone that roots for small companies and passionate programmers should appreciate Obsidian proving that the best tools don't have to be built by the biggest teams. More the opposite.
I wonder if past generations have cared as much as we do about helping people feel better about their perceived failings? This essay fits snugly into a 2020s canon of NYT Best Sellers around "work less, enjoy more." My impression had been that previous generations principally aspired to get more done to advance their lot in life. If that's true, it will be interesting to assess later how long-term happiness compares between the "try harder do more" past generations and the "you're amazing why are you working so hard?" current one.
Cal Newport is the only contemporary personality I know of that advocates for getting more done, but his method still conforms to the zeitgeist, in that his "get more done" is more precisely "get more important stuff done by doing less overall."
Perhaps this is the natural evolution that occurs when a country has reached sufficient wealth where contributing gains to per capita GDP just ain't inspiring to a comfortable generation? It's interesting to try to contextualize the current popularity of anti-productivity literature vs what has come before it.
> Sure, I have customers to assist, servers to manage (not that they need much management), and business accounting to do; but professors equally have classes to teach, students to supervise, and committees to attend. When it comes to research, I can follow my interests without regard to the whims of granting agencies and tenure and promotion committees.
As a fellow developer/owner, this is the mantra I'm always repeating.
Is running a small software biz a great job? Meh. The job has higher highs and lower lows than when I was a dev for other companies.
At the end of the day, there's no better way to make sustained progress on groundbreaking, long-term dev projects than to run a company that dedicated to building such projects. If one is fortunate enough to find a stable cash source (like Tarsnap), it serves as a perpetual license to work on the most interesting dev projects one can fathom succeeding at.
It’s not possible to measure developer productivity, but “developer hertz” is still tied to commit count so not ideal. This is a problem I think is important enough to deserve a metric to measure the rate at which a repo evolves (emphasis on updating legacy code). When GitClear started building Line Impact 5 years ago (free as of 6 months ago) our hope was that someday devs would pick a metric that’s better than commit count, and one that prioritizes dev-control of long term data (as GC does). If GitClear sounds familiar prob because we also support open source projects like libinput, but eventually we hope to be known as the dev-centric metric provider
This dark pattern smacks of an IBM esque legacy co that will be vanquished by better competitors in due time. But as a user that keeps his process monitor open, Adobe’s greatest sin is their multiple mandatory Creative Cloud processes that soak up memory and CPU whether an Adobe product is open or not. Rotten at the core.
Love that dashboard! Makes me wonder what your Line Impact would be for the various projects, seems like you get a lot of business value from each line of code.
So it has been for 10+ years, so it will be at least 5 more.