Authors sue OpenAI for using their works without proper licensing(nytimes.com)
nytimes.com
Authors sue OpenAI for using their works without proper licensing
https://www.nytimes.com/2023/09/20/books/authors-openai-lawsuit-chatgpt-copyright.html
142 comments
https://archive.ph/Kn5xq
"But Defendants’ LLMs endanger fiction writers’ ability to make a living, in that the LLMs allow anyone to generate—automatically and freely (or very cheaply)—texts that they would otherwise pay writers to create"
This kind of luddism sees copyright as a way to enrich rights holders, as opposed to "promoting the progress of science and the useful arts".
This kind of luddism sees copyright as a way to enrich rights holders, as opposed to "promoting the progress of science and the useful arts".
It appears this lawsuit is complaining that ChatGPT can write fan fiction and they don't like that.
I was onboard initially thinking we were talking about OpenAI ingesting Game Of Thrones as training material, but it appears George et al are just mad because it can make stories with their characters.
This is far from the authorship/copyright problem of AI.
I was onboard initially thinking we were talking about OpenAI ingesting Game Of Thrones as training material, but it appears George et al are just mad because it can make stories with their characters.
This is far from the authorship/copyright problem of AI.
If you read the claims to relief (start on page 44 of complaint) it's mostly just standard copyright infringement during training. The claims about ChatGPT writing works that infringes their works seem to be to be to try and head-off a fair use defense. One of the tests for fair use is the effect on the potential market for original work.
Edit: This comment was for a different headline, doesn't apply anymore.
Terrible headline. They're not suing for theft, they're suing for copyright infringement.
The rest of the article is reasonable, and they link to the complaint, which is something every article about a lawsuit should do.
https://fingfx.thomsonreuters.com/gfx/legaldocs/xmvjlbqbnvr/...
Terrible headline. They're not suing for theft, they're suing for copyright infringement.
The rest of the article is reasonable, and they link to the complaint, which is something every article about a lawsuit should do.
https://fingfx.thomsonreuters.com/gfx/legaldocs/xmvjlbqbnvr/...
Theft is the word used by their lawyers. Seems fair to use in the title. The differences between theft and copyright infringement isn't important in this case anyway.
Either way, it will be interesting to see how this goes. Lots of weird arguments on both sides, so it will be interesting to see what rulings we get.
Either way, it will be interesting to see how this goes. Lots of weird arguments on both sides, so it will be interesting to see what rulings we get.
The complaint:
https://ia904703.us.archive.org/10/items/gov.uscourts.nysd.6...
Named plaintiffs include David Baldacci, John Grisham and Scott Turow.
https://ia904703.us.archive.org/10/items/gov.uscourts.nysd.6...
Named plaintiffs include David Baldacci, John Grisham and Scott Turow.
Here's a point that I struggle with. Let's imagine a point in the future where technology has progressed to the point that a machine can become "assisted memory" for a human. This could be useful for degraded memory conditions or even just to buff up human capabilities. In this scenario how do we deal with licensing and copyright? The "memory" is trained on books, artwork, etc and then human intelligence accesses that computer aided memory and constructs something new.
Seems like this lawsuit could set precedence that the future I describe would not be allowed.
Seems like this lawsuit could set precedence that the future I describe would not be allowed.
I think it depends on the use. The kind of AI you're suggesting wouldn't need to ingest the entirety of every book and movie you've seen, it would just need to store a summary and a location for the original file.
It's legal for me to make digital copies of copyrighted works I own so long as they're not redistributed. If I have a local AI maintaining my personal data for reasons other than generating competing works, that shouldn't involve copyright at all.
It's legal for me to make digital copies of copyrighted works I own so long as they're not redistributed. If I have a local AI maintaining my personal data for reasons other than generating competing works, that shouldn't involve copyright at all.
I think the question boils down to what is fair use in these situations. I actually think it would be useful to have a public effort to build a pan humanity model that claims fair use for public benefit and consumes all productions of humanity.
What specifically hurts OpenAI is the monetization and commercial benefit from other people’s work. They also fall clearly in the space of civil claims of copyright violation, if not criminal (however the use of models to produce copyright materials is likely criminal infringement). (IANAL, but a law professor friend made these claims to me, YMMV)
What specifically hurts OpenAI is the monetization and commercial benefit from other people’s work. They also fall clearly in the space of civil claims of copyright violation, if not criminal (however the use of models to produce copyright materials is likely criminal infringement). (IANAL, but a law professor friend made these claims to me, YMMV)
I think this is all mostly huff and puff and frankly going to be irrelevant. At some point anyone with a decent enough home computer will be able to train their own models and use them as they see fit. Any law that tries to stop that is going to be stifling and unwieldy in the extreme.
And it'll be cross border, so even more difficult to enforce.
And it'll be cross border, so even more difficult to enforce.
How is it different than a search index? It takes existing content as an input, processes it and then outputs data structures from it. Those data structures are then used to power full text search.
A LLM does the same thing but instead of search results it emits a stream of tokens.
A LLM does the same thing but instead of search results it emits a stream of tokens.
DMCA 2024 - A nice big report button for when generated content is too close to copyrighted content. It is then on the AI company to supplement the training materials around that content, to dilute the generation of content that could be seen as infringing. So instead of George RR Martin prequels with the same names and characters (because of a lack of training materials), it generates something more generic for the input prompt.
Win/win?
Win/win?
The actual complaint is about using their copyrighted works in the training of the LLM without a license. OpenAI is claiming it's fair use, the authors disagree. It's going to take a ruling from a judge to get clarity on the issue, and no matter what it'll be appealed until it hits the SC.
Do we know if OpenAi bought the book, or did they just "accidentally" pirate the book?
That's what discovery will be for, the complaint alleges that the likely source was libgen. Most of these authors haven't released DRM-free ebooks, and it seems unlikely that OpenAI has a large scale book scanning effort (and even if they did, that authors would likely claim that to be infringement itself.)
What if it never accessed the book, but read everything relevant like episode summaries, fan wikis, and forum discussions? It would still be as conversant. Is it still infringement?
oh right was it really proven that they were training on bittorrent book collections?
Or, just let people and computers be inspired
Ideas have never been the scope of copyright and it wasn’t in its democratic mandate. If creatives want that change, fine, advocate for a change of the law
Ideas have never been the scope of copyright and it wasn’t in its democratic mandate. If creatives want that change, fine, advocate for a change of the law
>Ideas have never been the scope of copyright
This isn't about ideas, it's about a specific individuals work given that the reproduced text lifts literal characters out of Martin's book. That has always been covered by IP law. Canonical example, you cannot write a novel about Harry Potter, you can write a book about a wizard going to a magical school.
This isn't about ideas, it's about a specific individuals work given that the reproduced text lifts literal characters out of Martin's book. That has always been covered by IP law. Canonical example, you cannot write a novel about Harry Potter, you can write a book about a wizard going to a magical school.
If a model generates large amounts of text that is very close to something you've written, because there isn't much else like it, how is that "inspired"? It needs more dilution.
We would have to change the law to allow the kind of ‘inspiration’ you are talking about, which is why there are multiple lawsuits here. That’s what OpenAI is asking for - redefinition of ‘fair use’. NNs aren’t copying ideas, they train on what copyright calls ‘fixation’ - they deal with text, audio, and pixels, not ideas. We keep hoping and looking for understanding in the NNs, but we have ample evidence that they don’t actually understand much, if anything, they are just really good at copying in a way that make understanding seem plausible to the layperson.
It’s a good idea to make this easier to report, but… shouldn’t it be on the AI company to train using legally acquired content in the first place? It’d be great if the training data was opt-in and curated. Wouldn’t that be better than a shoot first ask questions later policy? There’s definitely room to improve copyright and room to allow AI to exist, but do we really want to allow AI to ingest all copyrighted material and call it ‘fair use’? That would be giving them a ridiculous and unprecedented amount of freedom to take any and all content and turn around and auto-generate enough to obsolete the people who made the training material. It seems like the race is on to supplant Google as the portal for information, and it does feel like downloading everything in the world and then crying fair use after the fact is wishful thinking that more or less admits to copyright violation.
>shouldn’t it be on the AI company to train using legally acquired content in the first place
I don't think so. It's not illegal to look at or learn from copyrighted materials. If you start producing the materials it becomes a different question. I think the same applies to AI.
I don't think so. It's not illegal to look at or learn from copyrighted materials. If you start producing the materials it becomes a different question. I think the same applies to AI.
Your argument doesn’t work because OpenAI has admitted that ChatGPT is producing copyrighted material. They’re trying to carve an exception for AI, but have already acknowledged that training does copy the materials, literally, and that it does not “learn” from the the same way humans do. The intent with AI may be to remix them, but the whole reason there are multiple lawsuits here (as well as with Stable Diffusion and other NNs) is because they have repeatedly demonstrated they sometimes memorize the training data and can produce it more or less verbatim. They have violated current copyright law. In that light, we have two primary options: change the law, or enforce the current law. OpenAI is hoping to change the law, but whether they have copied some training data and produced it for the output is not even up for debate, this is already the different question you referred to.
There seems to be confusion. :)
Or disagreement anyway, about how comparable photocopiers & copyright are to generative models and protection from unauthorized automated style reproduction.
How I look at it:
1. In both cases, reproduced copies or reproduced styles, automation destroys economic incentives for creators to make any sustained effort.
Without economic protection, it isn’t even a question of less motivation. Creator’s like everyone else need to eat.
2. So we protect creative works from complete copies in order to have more creative works.
And it is primarily about automation and mass reproduction.
Nobody is worried about people hand copying Atlas Shrugged.
3. But we also protect copyrighted works from partial copying.
Only copying chapters 1-3? Not allowed
Only copying the plot but changing all the names, locations, fashion and colors? Not allowed.
4. So now it turns out a different substantial part of a work can be copied via automation. It’s style.
Well if you can protect a works plot from automated copies, why not a works style?
It is a substantial piece of a creative work.
Reasons for protecting style come down to protecting any major part of a copyrighted work.
The only thing different now is we have “style reproducers”.
So we have to decide, is this essentially the same situation as copyright addresses, or not?
5. It is.
The exact same trade offs between protection and incentivization exist for extracted & mass reproduced style as they do for extracted and reproduced plot.
Or disagreement anyway, about how comparable photocopiers & copyright are to generative models and protection from unauthorized automated style reproduction.
How I look at it:
1. In both cases, reproduced copies or reproduced styles, automation destroys economic incentives for creators to make any sustained effort.
Without economic protection, it isn’t even a question of less motivation. Creator’s like everyone else need to eat.
2. So we protect creative works from complete copies in order to have more creative works.
And it is primarily about automation and mass reproduction.
Nobody is worried about people hand copying Atlas Shrugged.
3. But we also protect copyrighted works from partial copying.
Only copying chapters 1-3? Not allowed
Only copying the plot but changing all the names, locations, fashion and colors? Not allowed.
4. So now it turns out a different substantial part of a work can be copied via automation. It’s style.
Well if you can protect a works plot from automated copies, why not a works style?
It is a substantial piece of a creative work.
Reasons for protecting style come down to protecting any major part of a copyrighted work.
The only thing different now is we have “style reproducers”.
So we have to decide, is this essentially the same situation as copyright addresses, or not?
5. It is.
The exact same trade offs between protection and incentivization exist for extracted & mass reproduced style as they do for extracted and reproduced plot.
How many books written have taken some style or concepts from other books?
Stranger things takes liberally from a lot of Stephen King as Spielberg elements not outright but in spirit and tone, why isn't Stephen King suing the Duffer brothers for reading his shit and coming up with ideas for books based on that?
Stranger things takes liberally from a lot of Stephen King as Spielberg elements not outright but in spirit and tone, why isn't Stephen King suing the Duffer brothers for reading his shit and coming up with ideas for books based on that?
One is that basic plots are copied all the time and there’s a meme that there are only seven basic plots. Of course there’s much more variety at the detail level.
Was Sword of Shannara pretty derivative of Tolkien? Yeah. But I assume it was pretty far from a copyright violation.
Was Sword of Shannara pretty derivative of Tolkien? Yeah. But I assume it was pretty far from a copyright violation.
So 7 basic plots. But an actual plot for an original story isn’t just a basic plot is it? It’s an original work.
Movies are sued all the time for copyright infringement due to substantially copying plot and character elements. [0]
Because these cases tend to each be unique, the line between infringement and non-infringement gets settled very much on a case by case basis.
As a result of this inherent unpredictability, most cases involve the accused settling with the aggrieved party to get the lawsuit dismissed.
This is common in many areas of civil law.
A few examples:
1. *"The Island" (2005)* - Accusation: Similarities to the 1979 film "Parts: The Clonus Horror." - Outcome: Settled out of court. [1]
2. *"Frozen" (2013)* - Accusation: Claimed similarities to a short film named "The Snowman." - Outcome: Disney settled the case. [2]
3. *"Coming to America" (1988)* - Accusation: Art Buchwald claimed the movie was based on his script. - Outcome: Paramount settled for an undisclosed amount. [3]
4. *"The Terminator" (1984)* - Accusation: Harlan Ellison claimed it was similar to an episode of "The Outer Limits." - Outcome: Settled out of court, and an acknowledgment was added to later copies. [4]
5. *"Disturbia" (2007)* - Accusation: Accused of being similar to Alfred Hitchcock's "Rear Window." - Outcome: Initially dismissed, but a settlement was reached. [5]
[0] https://movieweb.com/movies-accused-of-copyright-infringemen...
[1] https://en.m.wikipedia.org/wiki/Parts:_The_Clonus_Horror
[2] https://ew.com/article/2015/06/25/disney-frozen-lawsuit-the-...
[3] https://en.m.wikipedia.org/wiki/Buchwald_v._Paramount
[4] https://www.cbr.com/terminator-harlan-ellison-credit/
[5] https://www.flixist.com/new-disturbia-and-rear-window-lawsui...
Movies are sued all the time for copyright infringement due to substantially copying plot and character elements. [0]
Because these cases tend to each be unique, the line between infringement and non-infringement gets settled very much on a case by case basis.
As a result of this inherent unpredictability, most cases involve the accused settling with the aggrieved party to get the lawsuit dismissed.
This is common in many areas of civil law.
A few examples:
1. *"The Island" (2005)* - Accusation: Similarities to the 1979 film "Parts: The Clonus Horror." - Outcome: Settled out of court. [1]
2. *"Frozen" (2013)* - Accusation: Claimed similarities to a short film named "The Snowman." - Outcome: Disney settled the case. [2]
3. *"Coming to America" (1988)* - Accusation: Art Buchwald claimed the movie was based on his script. - Outcome: Paramount settled for an undisclosed amount. [3]
4. *"The Terminator" (1984)* - Accusation: Harlan Ellison claimed it was similar to an episode of "The Outer Limits." - Outcome: Settled out of court, and an acknowledgment was added to later copies. [4]
5. *"Disturbia" (2007)* - Accusation: Accused of being similar to Alfred Hitchcock's "Rear Window." - Outcome: Initially dismissed, but a settlement was reached. [5]
[0] https://movieweb.com/movies-accused-of-copyright-infringemen...
[1] https://en.m.wikipedia.org/wiki/Parts:_The_Clonus_Horror
[2] https://ew.com/article/2015/06/25/disney-frozen-lawsuit-the-...
[3] https://en.m.wikipedia.org/wiki/Buchwald_v._Paramount
[4] https://www.cbr.com/terminator-harlan-ellison-credit/
[5] https://www.flixist.com/new-disturbia-and-rear-window-lawsui...
This is good, and this is inevitable. Creators have control of their copyright, which should include permissions to be used in AI training.
Authors currently don't control who or what reads their works, of course.
Personally I currently feel that (at life +70 years) the copyright pendulum has gone too far towards the rights of publishers (not necessarily authors) as is.
That said, I'm open to good arguments to change my mind. Why do you feel that authors should be given this additional right to control what is used for AI training? What would be the public good or public trade-off here?
Personally I currently feel that (at life +70 years) the copyright pendulum has gone too far towards the rights of publishers (not necessarily authors) as is.
That said, I'm open to good arguments to change my mind. Why do you feel that authors should be given this additional right to control what is used for AI training? What would be the public good or public trade-off here?
And they don’t control fair use or promulgating the ideas in the book. That Wikipedia article summarizing the key contents in an editor’s own words? Perfectly legit.
What copyright buys is that no one else can distribute verbatim copies of large amounts of your work. But a lot of other uses are allowed.
What copyright buys is that no one else can distribute verbatim copies of large amounts of your work. But a lot of other uses are allowed.
Copyright indisputably covers more than the verbatim reprinting of a book
For example?
It covers the expression of ideas. Which in the case of a book is mostly the text as written. And, yes, doing some substitution of character names etc. may still violate copyright but you certainly can’t keep me from writing an article about the main points you make in your book.
It covers the expression of ideas. Which in the case of a book is mostly the text as written. And, yes, doing some substitution of character names etc. may still violate copyright but you certainly can’t keep me from writing an article about the main points you make in your book.
> What would be the public good or public trade-off here?
Consider the aesthetic landscape where creators do not have control over whether their work is used to train an AI versus one where they do. It's hard to predict with certainty, but my model is this:
No control: Anyone's work is fair game to be trained. If I want to make a prompt of "A graphic novel in the visual style of Moebius, written by Stephen King, set in Westeros" I can get something based on King and Martin's actual words and Moebius' actual drawings, without compensating them. Neat! However, potential new novelists see that quality novels can just be churned out for free or low cost and so, actually sitting down to write a new novel becomes a niche, geek thing to do. There's no money in it. These new novels just get thrown into the ml bin, fodder for the next version.
With control: novelists and other creators know they can make money from their work because they can make business decisions about how and when their work trains a model. We all get to see more new, professional quality creativity. Those who want to read Conan as written by Lord Dunsany can still see that, since those works are in the public domain.
Consider the aesthetic landscape where creators do not have control over whether their work is used to train an AI versus one where they do. It's hard to predict with certainty, but my model is this:
No control: Anyone's work is fair game to be trained. If I want to make a prompt of "A graphic novel in the visual style of Moebius, written by Stephen King, set in Westeros" I can get something based on King and Martin's actual words and Moebius' actual drawings, without compensating them. Neat! However, potential new novelists see that quality novels can just be churned out for free or low cost and so, actually sitting down to write a new novel becomes a niche, geek thing to do. There's no money in it. These new novels just get thrown into the ml bin, fodder for the next version.
With control: novelists and other creators know they can make money from their work because they can make business decisions about how and when their work trains a model. We all get to see more new, professional quality creativity. Those who want to read Conan as written by Lord Dunsany can still see that, since those works are in the public domain.
Training AI feels inherently commercial with intent to commercially distribute, which seems to be licensed specially in most domains.
In a sense it's like driving down the highway with a duffel bag of cannabis flower in a state where possessing and traveling with few ounces is no problem--something commercial is probably happening. Why is that prohibited? Perhaps another debate, but just trying to connect the implied intent aspect.
If an AI were being trained for strictly academic reasons then I'd agree with fair use and that type of arguments. But if the AI itself has a subscription fee, then whoever is subscribing is also probably using the work for real or anticipated commercial gain. Hence investing money.
True that hobbyists spend money with no intention to gain commercially, and we may do that at a higher than average rate as tech workers because we have usually have a decent amount of excess money from our work. But money is pretty scarce to most people and businesses with set non-investment budgets, so if they're spending it on AI there's little doubt it's with commercial intent.
So in conclusion, I do think there's both merit to the authors' case related to intent to commercialize and room for doing unlicensed non-commercial AI training.
In a sense it's like driving down the highway with a duffel bag of cannabis flower in a state where possessing and traveling with few ounces is no problem--something commercial is probably happening. Why is that prohibited? Perhaps another debate, but just trying to connect the implied intent aspect.
If an AI were being trained for strictly academic reasons then I'd agree with fair use and that type of arguments. But if the AI itself has a subscription fee, then whoever is subscribing is also probably using the work for real or anticipated commercial gain. Hence investing money.
True that hobbyists spend money with no intention to gain commercially, and we may do that at a higher than average rate as tech workers because we have usually have a decent amount of excess money from our work. But money is pretty scarce to most people and businesses with set non-investment budgets, so if they're spending it on AI there's little doubt it's with commercial intent.
So in conclusion, I do think there's both merit to the authors' case related to intent to commercialize and room for doing unlicensed non-commercial AI training.
What is inevitable is that copyright is dead. Do you think China will respect western copyrights? They already don't. People would just use LLMs hosted elsewhere.
This is a good thing. We're going to see an explosion of indie games, movies, and more that never could've been made before by a single, dedicated person.
This is a good thing. We're going to see an explosion of indie games, movies, and more that never could've been made before by a single, dedicated person.
“Should include” is exactly not the way the law works.
Is it illegal to SHA-256 a book without the author's permission?
If not, how are we going to legally codify the difference between that and an LLM?
If not, how are we going to legally codify the difference between that and an LLM?
You can't extract content from a SHA-256(book), but you can extract some content from LLM(book).
Can I, though? In my tests with ChatGPT I haven't been able to get it to regurgitate books or even small portions of them.
Yep, you can. I tested by typing the first 13 words of the book Frankenstein and GPT3.5 happily completed the rest. Feel free to try.
Here are screenshots of the completion and the diff with the real book.
https://ibb.co/album/sJBB11
Here are screenshots of the completion and the diff with the real book.
https://ibb.co/album/sJBB11
That's a public domain text. So, yes, it can do it. I stand corrected there. But can you coax a copyrighted text out of it?
When I ask it for anything from Game of Thrones it refuses beyond offering a summary.
But this is all begging the question.
How are you going to codify that? "You can't feed my text into an algorithm that has the ability to reproduce it"?
When I ask it for anything from Game of Thrones it refuses beyond offering a summary.
But this is all begging the question.
How are you going to codify that? "You can't feed my text into an algorithm that has the ability to reproduce it"?
I also tested this with Harry Potter books (which are still not public domain AFAIK). It starts auto-completing the correct words quickly, and then hangs and eventually stops producing output. You can call the API to generate more tokens, but it stops again after producing a few more correct words.
I think for a few high-profile authors (for books like Harry Potter and Game of Thrones), OpenAI probably installed some output filters in order to not get sued too hard. Of course, I can't definitively check that without access to the raw model. Which OpenAI conveniently doesn't provide.
I think for a few high-profile authors (for books like Harry Potter and Game of Thrones), OpenAI probably installed some output filters in order to not get sued too hard. Of course, I can't definitively check that without access to the raw model. Which OpenAI conveniently doesn't provide.
The complaint mentions LibGen, Z-Library and Bibliotik, and Sci-Hub in a footnote.
Thought experiment:
What if every person who dowloads materials from the above sources claimed that they were doing so only to "train AI".
Many such persons who download from those sources are probably doing so for noncommercial purposes, for example, academic research. Whereas, according to this compaint, OpenAI "intend[s] to earn billions from this technology."
Thought experiment:
What if every person who dowloads materials from the above sources claimed that they were doing so only to "train AI".
Many such persons who download from those sources are probably doing so for noncommercial purposes, for example, academic research. Whereas, according to this compaint, OpenAI "intend[s] to earn billions from this technology."
[dupe]
More discussion days back when this was news:
https://news.ycombinator.com/item?id=37585157
https://news.ycombinator.com/item?id=37599261
More discussion days back when this was news:
https://news.ycombinator.com/item?id=37585157
https://news.ycombinator.com/item?id=37599261
I don't get it.
So AI works can't be copyrighted but training AI using copyrighted materials are copyright infringement?
So AI works can't be copyrighted but training AI using copyrighted materials are copyright infringement?
Only humans can violate laws: the AI isn't using copyrighted materials, it's human owner/operator is.
The human owner/operator did not have proper licensing to perform this action, trying to argue "but with an AI" doesn't change the act.
The human owner/operator did not have proper licensing to perform this action, trying to argue "but with an AI" doesn't change the act.
Not disagreeing with your end conclusion, but surely the concept of a limited company exists exactly to have a distinction between legal entities, some of which are not humans but may still violate laws. Take it this way: if the EU fines a company for GDPR violations, it doesn't really fine an individual. Perhaps no individual broke the law explicitly, but as a collective the end result is a law violation.
Technically yes, but how that is handled is up to the country. In the US, a concept known as "corporate personhood" exists. A strange concept, because if a company murders someone, the company, nor any of its executives, go to prison.
The law is weird.
The law is weird.
GRR Martin, the author of Game of Thrones, had the audacity to join this lawsuit. The only thing I expect from AI in this context is NOT to reproduce the shitshow GOT ended up to be.
I'm wondering what will the authors do if we develop AIs that are able to find new artistic styles that are not in the dataset ? Would it still pose problem to use their content to learn how NOT to imitate them ?
Seems it is possible in collaborative filtering:
> Yes, in collaborative filtering, finding empty classes is possible. To recommend items for these gaps, utilize adjacent class information or employ techniques like matrix factorization, content-based filtering, or hybrid systems. These methods predict preferences based on observed patterns, similarities between items, and user preferences, filling in missing data.
I'm wondering what will the authors do if we develop AIs that are able to find new artistic styles that are not in the dataset ? Would it still pose problem to use their content to learn how NOT to imitate them ?
Seems it is possible in collaborative filtering:
> Yes, in collaborative filtering, finding empty classes is possible. To recommend items for these gaps, utilize adjacent class information or employ techniques like matrix factorization, content-based filtering, or hybrid systems. These methods predict preferences based on observed patterns, similarities between items, and user preferences, filling in missing data.
I see a business opportunity for someone who can produce a watermark that will visibly poison the learning set.
Interesting how this particular case is generating such conflicting political views
currently anything generated by llm is not copyrightable, right? so are they really threatened by public domain or what?
They are threatened by the fact that LLM can write better stories about their invented characters then themselves.
Maybe the default should be opt in instead of opt out. Why should copyright holders now have to do work to protect their already protected works?
slavboj(6)