Our more recent essay (and ongoing book project) "AI as Normal Technology" is about our vision of AI impacts over a longer timescale than "AI Snake Oil" looks at https://www.normaltech.ai/p/ai-as-normal-technology
I would categorize our views as techno-optimist, but people understand that term in many different ways, so you be the judge.
Thanks! HN was part of the origin story of the book in question.
In 2018 or 2019 I saw a comment here that said that most people don't appreciate the distinction between domains with low irreducible error that benefit from fancy models with complex decision boundaries (like computer vision) and domains with high irreducible error where such models don't add much value over something simple like logistic regression.
It's an obvious-in-retrospect observation, but it made me realize that this is the source of a lot of confusion and hype about AI (such as the idea that we can use it to predict crime accurately). I gave a talk elaborating on this point, which went viral, and then led to the book with my coauthor Sayash Kapoor. More surprisingly, despite being seemingly obvious it led to a productive research agenda.
While writing the book I spent a lot of time searching for that comment so that I could credit/thank the author, but never found it.
Thanks for the comment! I agree — it's important to remain fluid. We've taken steps to make sure that predictively speaking, the normal technology worldview is empirically testable. Some of those empirical claims are in this paper and others in coming in follow-ups. We are committed to revising our thinking if it turns out that our framework doesn't generate good predictions and effective prescriptions.
We do try to admit it when we get things wrong. One example is our past view (that we have since repudiated) that worrying about superintelligence distracts from more immediate harms.
We do not assume a status quo or equilibrium, which will hopefully be clear upon reading the paper. That's not what normal technology means.
Part II of the paper describes one vision of what a world with advanced AI might look like, and it is quite different from the current world.
We also say in the introduction:
"The world we describe in Part II is one in which AI is far more advanced than it is today. We are not claiming that AI progress—or human progress—will stop at that point. What comes after it? We do not know. Consider this analogy: At the dawn of the first Industrial Revolution, it would have been useful to try to think about what an industrial world would look like and how to prepare for it, but it would have been futile to try to predict electricity or computers. Our exercise here is similar. Since we reject “fast takeoff” scenarios, we do not see it as necessary or useful to envision a world further ahead than we have attempted to. If and when the scenario we describe in Part II materializes, we will be able to better anticipate and prepare for whatever comes next."
I appreciate the concern, but we have a whole section on policy where we are very concrete about our recommendations, and we explicitly disavow any broadly anti-regulatory argument or agenda.
The "drastic" policy interventions that that sentence refers to are ideas like banning open-source or open-weight AI — those explicitly motivated by perceived superintelligence risks.
This paper is being misinterpreted. The degradations reported are somewhat peculiar to the authors' task selection and evaluation method and can easily result from fine tuning rather than intentionally degrading GPT-4's performance for cost saving reasons.
They report 2 degradations: code generation & math problems. In both cases, they report a behavior change (likely fine tuning) rather than a capability decrease (possibly intentional degradation). The paper confuses these a bit: they mostly say behavior, including in the title, but the intro says capability in a couple of places.
Code generation: the change they report is that the newer GPT-4 adds non-code text to its output. They don't evaluate the correctness of the code. They merely check if the code is directly executable. So the newer model's attempt to be more helpful counted against it.
Math problems (primality checking): to solve this the model needs to do chain of thought. For some weird reason, the newer model doesn't seem to do so when asked to think step by step (but the current ChatGPT-4 does, as you can easily check). The paper doesn't say that the accuracy is worse conditional on doing CoT.
The other two tasks are visual reasoning and answering sensitive questions. On the former, they report a slight improvement. On the latter, they report that the filters are much more effective — unsurprising since we know that OpenAI has been heavily tweaking these.
In short, everything in the paper is consistent with fine tuning. It is possible that OpenAI is gaslighting everyone by denying that they degraded performance for cost saving purposes — but if so, this paper doesn't provide evidence of it. Still, it's a fascinating study of the unintended consequences of model updates.
OP here. Unfortunately this thread is mostly misinformation. There were a bunch of viral threads from the growth hacker / influencer crowd, including this one, within hours of the code release with a very superficial understanding of the code (and how recsys work in general). That's partly what motivated me to write this article.
OP here. The CNET thing is actually pretty egregious, and not the kind of errors a human would make. These are the original investigations, if you'll excuse the tone:
https://futurism.com/cnet-ai-errors
Summary: misinfo, labor impact, and safety are real dangers of LLMs. But in each case the letter invokes speculative, futuristic risks, ignoring the version of each problem that’s already harming people. It distracts from the real issues and makes it harder to address them.
The containment mindset may have worked for nuclear risk and cloning but is not a good fit for generative AI. Further locking down models only benefits the companies that the letter seeks to regulate.
Besides, a big shift in the last 6 months is that model size is not the primary driver of abilities: it’s augmentation (LangChain etc.) And GPT3-class models can now run on iPhones. The letter ignores these developments. So a moratorium is ineffective at best and counterproductive at worst.
We don't expect it to be free -- please read the article. That's not the issue at all. It's like if you subscribe to a product that you need to do your job, and one day the company tells you that the product is going away in three days and that you need to switch to a different product (that isn't at all the same for your use case).
Sure, but the article is talking about a completely different meaning of reproducibility, where a researcher uses an LLM as a tool to study some research question, and someone else comes along and wants to check whether the claims hold up.
This doesn't in any way require the training run or the build to be reproducible. It just requires the model, once released through the API, to remain available for a reasonable length of time (and not have the rug pulled with 3 days' notice).
We're under no such misapprehension and we're keenly aware that this is an uphill battle. The issue is that LLMs have become part of the infrastructure of the Internet. Companies that build infrastructure have a responsibility to society, and we're documenting how OpenAI is reneging on that responsibility. Hindering research is especially problematic if you take them at their word that they're building AGI. If infrastructure companies don't do the right thing, they eventually get regulated (and if you think that will never happen, I have one word: AT&T).
Finally, even if you don't care about research at all, the article mentions OpenAI's policy that none of their models going forward will be stable for more than 3 months, and it's going to be interesting to use them in production if things are going to keep breaking regularly.
"OpenAI responded to the criticism by saying they'll allow researchers access to Codex. But the application process is opaque: researchers need to fill out a form, and the company decides who gets approved. It is not clear who counts as a researcher, how long they need to wait, or how many people will be approved. Most importantly, Codex is only available through the researcher program “for a limited period of time” (exactly how long is unknown)."
OP here. Many people are reacting to the title of the paper. A few thoughts:
* The paper is 35 pages long and it's hard to convey its message in any single title. We make clear in the text that our point is not that predictive optimization should never be used.
* We do want the _default_ to change from predictive optimization being seen as the obvious way to solve certain social problems to being against it until the developer can address certain objections. This is also made clear in the paper.
* The title is a nod to a famous book in this area called "Against prediction". Most people in our primary target audience are familiar with that book, so the title conveys a lot of information to those readers. That's one reason we picked it.
* Despite its flaws, when might we want to use predictive optimization? Section 4 gets into this in detail.
I learned from one of the comments on my original post that many scholars have been saying this for a while, and that there's in fact a book that makes the same point!
OP here. The full title of this article is "Students are acing their homework by turning in machine-generated essays. Good."
The last word was edited out by the mods, presumably under the belief that it's clickbait. Unfortunately, the headline now sounds like I'm complaining about this development, whereas my post is about how it will force much-needed improvements to education and free students from the drudgery of pointless essays that ask them to regurgitate content (as opposed to essays that teach writing skills or critical thinking, which remain valuable).
OP here. I totally agree that ideally authors should report most of this information in the paper itself. One advantage of a standalone document (we suggest putting it in an appendix) is that it's easy for reviewers to check that all of this information has been reported. Of course, authors could answer some of the questions by pointing to the sections of the paper in which they have been answered.
It's possible you may have misunderstood the title of the post. It isn't about the science of ML, or GPT-3, or brains. Rather, it's about using ML as a tool to do actual science, like medicine or political science or chemistry or whatnot. The first sentence of the post explains this.
Princeton University Center for Information Technology Policy | Princeton, NJ | Onsite | Full Time
Princeton CITP is a leading research center at the intersection of technology and public policy. We've conducted groundbreaking work on privacy, government surveillance, net neutrality, algorithmic fairness, dark patterns, and other high-profile topics. https://citp.princeton.edu/
We're hiring a data scientist who will collaborate with our world-class faculty, fellows, and students on interdisciplinary research projects and policy impact. If you live in New York City or New Jersey, are passionate about the societal impact of technology, and have an impressive resume in data science (broadly conceived), we want to hear from you.
That's fair. We don't claim that this is a new problem; we are merely adding evidence and our perspective to a known problem. We do link to others who have reported similar problems when trying to disclose vulnerabilities. The sentence saying we "discovered two wider issues" was worded poorly; in the paper [1] we used the word "encountered", and I've now edited the post to use the same wording. Thanks!
Just as important, the post is a PSA that there are 9 websites whose users remain vulnerable, and people with accounts on these sites should check their 2FA and password recovery settings. The websites are: Amazon, AOL, Finnair, Gaijin, Mailchimp, PayPal, Venmo, Wordpress.com, and Yahoo.
Research: https://www.cs.princeton.edu/~arvindn/