Artificial Intelligence and Copyright: Request for comments(federalregister.gov)
federalregister.gov
Artificial Intelligence and Copyright: Request for comments
https://www.federalregister.gov/documents/2023/08/30/2023-18624/artificial-intelligence-and-copyright
308 comments
There are three copyright issues here; datasets, model weights, and model outputs.
Dataset copyright is pretty well defined and things can often be used under fair use. Fair use decisions are done with a four prong test and really decided by the courts on a case-by-case basis.
Model weights cannot currently be copyrighted. They are the output of a mechanical process over the dataset. However, software faced a similar situation where the source code could be copyrighted but the compiled binary was not. US copyright law was updated to address this. We may see something similar for model weights.
Model outputs are less clear, but these are likely copyrightable by the user of the model. It is not possible for a non-human to hold copyright, so the model cannot. It is very unlikely that the company producing the model could assert copyright over the outputs. A good analogy here is someone using photo manipulation software.
Super interesting area. I think we will eventually see an update to the copyright code to make weights copyrightable. Also it will be interesting to see how court challenges (code generation, image generation) affect datasets in the future.
Dataset copyright is pretty well defined and things can often be used under fair use. Fair use decisions are done with a four prong test and really decided by the courts on a case-by-case basis.
Model weights cannot currently be copyrighted. They are the output of a mechanical process over the dataset. However, software faced a similar situation where the source code could be copyrighted but the compiled binary was not. US copyright law was updated to address this. We may see something similar for model weights.
Model outputs are less clear, but these are likely copyrightable by the user of the model. It is not possible for a non-human to hold copyright, so the model cannot. It is very unlikely that the company producing the model could assert copyright over the outputs. A good analogy here is someone using photo manipulation software.
Super interesting area. I think we will eventually see an update to the copyright code to make weights copyrightable. Also it will be interesting to see how court challenges (code generation, image generation) affect datasets in the future.
Much debate has been had about how existing copyright law applies to AI models. But once you get past that and start asking about how copyright should apply to AI models (as the copyright office is here) the answer in my mind becomes clear.
Copyright, as defined in the U.S. Constitution, exists "to promote the Progress of Science and useful Arts"[1]. I can think of no better modern example of "the Progress of Science and useful Arts" than AI models themselves. Therefore, it follows that:
1. Existing copyright laws should _not_ be applied in such a way as to make training these models any more difficult than it already is (as that would be in direct opposition to the stated goal)
2. AI models should be copyrightable by the person training the model (for the same reason any other software program is copyrightable)
3. Output of AI models should be copyrightable by the person running the model (for the same reason any other creative work is copyrightable) provided the output does not conflict with any preexisting copyright
For those who think training on copyrighted materials should be illegal, explain to me how that helps "promote the Progress of Science and useful Arts" and I'll re-consider my position.
[1]: https://en.wikipedia.org/wiki/Copyright_Clause
Copyright, as defined in the U.S. Constitution, exists "to promote the Progress of Science and useful Arts"[1]. I can think of no better modern example of "the Progress of Science and useful Arts" than AI models themselves. Therefore, it follows that:
1. Existing copyright laws should _not_ be applied in such a way as to make training these models any more difficult than it already is (as that would be in direct opposition to the stated goal)
2. AI models should be copyrightable by the person training the model (for the same reason any other software program is copyrightable)
3. Output of AI models should be copyrightable by the person running the model (for the same reason any other creative work is copyrightable) provided the output does not conflict with any preexisting copyright
For those who think training on copyrighted materials should be illegal, explain to me how that helps "promote the Progress of Science and useful Arts" and I'll re-consider my position.
[1]: https://en.wikipedia.org/wiki/Copyright_Clause
I'm going to try to plead my case for images generated using sophisticated prompt engineering to be copyrightable. For example, at the point that I've written a prompt with 20 tags, 10 negative prompt tags, some loras, custom weights, embeddings merges, and prompt editing, I'm now writing what is effectively a "program", which should be copyrightable and so should its outputs.
It's total BS to me that a book of midjourney generated images is itself copyrightable because a human arranged the book together, but that a highly sophisticated prompt involving custom tooling wouldn't be.
If nothing else, my comments should show the US Patent Office how deep the rabbit-hole goes with just how interpolatable everything is with everything else.
It's total BS to me that a book of midjourney generated images is itself copyrightable because a human arranged the book together, but that a highly sophisticated prompt involving custom tooling wouldn't be.
If nothing else, my comments should show the US Patent Office how deep the rabbit-hole goes with just how interpolatable everything is with everything else.
I have never understood the fair use argument when it comes to training data.
I publish a copyrighted article. Some LLM ingests it without permission, but since the output of that LLM is sufficiently different from my source article there is no violation.
I publish copyrighted code. Some company decides to consume it without purchasing a license. The product they distribute is vastly different from my code itself, but I can still sue them into oblivion.
What's the difference between the two?
I publish a copyrighted article. Some LLM ingests it without permission, but since the output of that LLM is sufficiently different from my source article there is no violation.
I publish copyrighted code. Some company decides to consume it without purchasing a license. The product they distribute is vastly different from my code itself, but I can still sue them into oblivion.
What's the difference between the two?
The only clear solution is to abandon the notion of a copyright.
We have know for a long time that everything can be represented with numbers, even more so within the space of computers.
All we have done is invent a system to help us find numbers we find special.
We have know for a long time that everything can be represented with numbers, even more so within the space of computers.
All we have done is invent a system to help us find numbers we find special.
I'm surprised that nobody has suggested that what's behind this RFC is Disney and other large studios lobbying to make it legal to copyright AI generated content so that they can move to AI generated movies and art. Right now you can't get a copyright on AI generated content.
AI Jesus chat-bot could claim copyright over biblical content.
In theory, a company that owns Christian (c 2023) content could be filing DMCA claims every Sunday.
The silliness of digital-racketeers must end at some point. =)
In theory, a company that owns Christian (c 2023) content could be filing DMCA claims every Sunday.
The silliness of digital-racketeers must end at some point. =)
The entire discussion about AI and copyright strikes me as a bit naive.
Right now, we are in a situation where nobody quite knows what these AI models are useful for. We have some inkling that they might be extraordinarily useful for making money -- but not precisely how, not even the companies that are developing the models themselves.
Once they money starts, the debate over copyright will fall exactly into the economic seams between the major players involved:
- new tech orgs who are monetizing models will say that the model is "exactly as humans are": they see copyrighted works in training, and then produce wholly original outputs. And of course that the model weights themselves are, like the outputs of employees, completely owned by the company.
- incumbents who stand to lose out on the new gold rush will say that every single output of a model belongs to them if just a single image or sentence was seen in training. And that because of that, we really should just shut the whole thing down, because how could you ever prove that a model was not trained on copyrighted material?
The faultlines will entirely rest on who has more power, hard and soft. How much can they influence the legal system, either by spending $ to hire legal talent or by sheer soft politicking, balanced with how favorable they appear to the general public who uses their product (or consumes their media). I suspect that the end result of this debate is a "legal" way of doing things accessible only to the extremely large players, and a small, politically insignificant collection of individuals, hackers, and startups who aim to unseat those large players (or just flat-out train "illegal" models). The worst possible end result is that the legal system is just too fossilized to deal and tries something draconian like not allow datacenter-scale GPU compute.
As an aside, I predict a sizeable space for companies that do "compliance" -- asserting the copyright status of a dataset, perhaps even themselves using ML. That market will carve off and leave rotting a sizeable chunk of the new money's ML profits.
It's fun to talk about this, I guess. But remember that what you or I have to say about what a machine learning model philosophically is has no bearing what-ever when it comes to the actual ability for individuals, startups, or large players to use models.
I will predict though: enjoy Llama2 while it lasts. Like the internet, it will become fully assimilated into the larger intellectual property machine.
Right now, we are in a situation where nobody quite knows what these AI models are useful for. We have some inkling that they might be extraordinarily useful for making money -- but not precisely how, not even the companies that are developing the models themselves.
Once they money starts, the debate over copyright will fall exactly into the economic seams between the major players involved:
- new tech orgs who are monetizing models will say that the model is "exactly as humans are": they see copyrighted works in training, and then produce wholly original outputs. And of course that the model weights themselves are, like the outputs of employees, completely owned by the company.
- incumbents who stand to lose out on the new gold rush will say that every single output of a model belongs to them if just a single image or sentence was seen in training. And that because of that, we really should just shut the whole thing down, because how could you ever prove that a model was not trained on copyrighted material?
The faultlines will entirely rest on who has more power, hard and soft. How much can they influence the legal system, either by spending $ to hire legal talent or by sheer soft politicking, balanced with how favorable they appear to the general public who uses their product (or consumes their media). I suspect that the end result of this debate is a "legal" way of doing things accessible only to the extremely large players, and a small, politically insignificant collection of individuals, hackers, and startups who aim to unseat those large players (or just flat-out train "illegal" models). The worst possible end result is that the legal system is just too fossilized to deal and tries something draconian like not allow datacenter-scale GPU compute.
As an aside, I predict a sizeable space for companies that do "compliance" -- asserting the copyright status of a dataset, perhaps even themselves using ML. That market will carve off and leave rotting a sizeable chunk of the new money's ML profits.
It's fun to talk about this, I guess. But remember that what you or I have to say about what a machine learning model philosophically is has no bearing what-ever when it comes to the actual ability for individuals, startups, or large players to use models.
I will predict though: enjoy Llama2 while it lasts. Like the internet, it will become fully assimilated into the larger intellectual property machine.
I wonder how they verify the personhood of the people making the comments. I can see this process being easily abused if the comments aren't taken by real people in person.
That said, I hope the US doesn't end up piling even more restrictions onto copyright. They'd only be shooting themselves in the foot. Copyright has completely failed to achieve the purpose for which it was intended to solve (only intended to give authors a short amount of time to profit off their efforts? look at it now). Perhaps it's time to rethink the concept of copyright as a whole before other countries beat the US to it.
That said, I hope the US doesn't end up piling even more restrictions onto copyright. They'd only be shooting themselves in the foot. Copyright has completely failed to achieve the purpose for which it was intended to solve (only intended to give authors a short amount of time to profit off their efforts? look at it now). Perhaps it's time to rethink the concept of copyright as a whole before other countries beat the US to it.
I'd be interested to see the outcome of all this honestly, but I see parallels in how we exclude natural organisms from copyright and instead rely on patents to enforce ownership of unique genetics.
We as a society have a relatively healthy setup for people to create art and content. Sure there are problems, but on the whole it mostly works. What AI will do is destroy that by removing the profitability of creating that content.
Although generative AI operates on a similar principle to a human being exposed to a large number of artworks, it does so at a speed blindingly faster, enabling it to outcompete humans at many tasks. The number of such tasks will only increase in the future.
Thus, small-time content creators who make an independent living from content creation will be squeezed out and left in the dark. In some years, it will be very hard to make money from content creation at all.
The access to information and entertainment will also become more anonymous, will most people consuming things through AI generation. Of course, that will be convenient at first, but we will end up with a world where a significantly SMALLER fraction of people controlling AI supply us with everything. (Including manipulative advertising to consume more of their product.)
For every benefit that AI gives us, there are 10 losses.
I used to dislike draconian copyright laws, but now I like them. And I sincerely hope they are used against AI to make AI unprofitable. I believe further that AI will be society-disrupting in a variety of other ways and thus, as a society, we should destroy it. But I am pessimistic.
Although generative AI operates on a similar principle to a human being exposed to a large number of artworks, it does so at a speed blindingly faster, enabling it to outcompete humans at many tasks. The number of such tasks will only increase in the future.
Thus, small-time content creators who make an independent living from content creation will be squeezed out and left in the dark. In some years, it will be very hard to make money from content creation at all.
The access to information and entertainment will also become more anonymous, will most people consuming things through AI generation. Of course, that will be convenient at first, but we will end up with a world where a significantly SMALLER fraction of people controlling AI supply us with everything. (Including manipulative advertising to consume more of their product.)
For every benefit that AI gives us, there are 10 losses.
I used to dislike draconian copyright laws, but now I like them. And I sincerely hope they are used against AI to make AI unprofitable. I believe further that AI will be society-disrupting in a variety of other ways and thus, as a society, we should destroy it. But I am pessimistic.
Every day of my life I wake up feeling more and more detached from the world.
That people are even debating this is so incredibly stupid to me.
The only people that will benefit from more onerous copyright are the major corporations that already have lawyers lined up and ready to fight their battles ad infinitum. See https://news.ycombinator.com/item?id=37347528
The everyday person will not benefit from any decisions made by courts on this matter. We will get the absolute shittiest implementation of copyright possible, see all other industries where copyright plays a role.
Copyright by itself is such a clever trick to prevent the world from advancing.
They convinced you to fight each other over scraps while they violate your copyright behind closed doors and use it to further their own agendas. They wield copyright like a weapon to effectively silence and dominate entire industries, leaving the average human unable to even comprehend how to fight back.
Lawyers and Copyright. Without them we would be so much better off.
That people are even debating this is so incredibly stupid to me.
The only people that will benefit from more onerous copyright are the major corporations that already have lawyers lined up and ready to fight their battles ad infinitum. See https://news.ycombinator.com/item?id=37347528
The everyday person will not benefit from any decisions made by courts on this matter. We will get the absolute shittiest implementation of copyright possible, see all other industries where copyright plays a role.
Copyright by itself is such a clever trick to prevent the world from advancing.
They convinced you to fight each other over scraps while they violate your copyright behind closed doors and use it to further their own agendas. They wield copyright like a weapon to effectively silence and dominate entire industries, leaving the average human unable to even comprehend how to fight back.
Lawyers and Copyright. Without them we would be so much better off.
Is there anyone willing to make the case for allowing generated works to have copyright protections?
A lot of people here seem to mistake copyright for "right-to-sell" generated works. As it stands now, with no copyrights granted for AI generated works, anyone can sell any generated works unless some copyright holder believes it violates their copyright.
A lot of people here seem to mistake copyright for "right-to-sell" generated works. As it stands now, with no copyrights granted for AI generated works, anyone can sell any generated works unless some copyright holder believes it violates their copyright.
I would say, treat the AI like a human viewing said data/material, but unlike the majority of humans, most dont have Kim Peek levels of observation and recall, so should there be some sort of expiration of data built into AI's and if so, should it be a blanket cut off date for all data an AI has been exposed to, or allow for some specialisation like a human might have in order to fulfil their occupation?
I'm also aware that search engines have access to data, most humans do not, which gives the search engines an advantage, think accessing ft.com for example, so should AI's have access to data in the context of a search engine or as a human?
Its tough, I want to give AI's full access to see what the tech is truly capable of, but I'm aware this will lead to more global monopolies over and above the existing search engine monopolies allowed to dominate in a country or territory, which will harm smaller AI entities.
I'm also aware that search engines have access to data, most humans do not, which gives the search engines an advantage, think accessing ft.com for example, so should AI's have access to data in the context of a search engine or as a human?
Its tough, I want to give AI's full access to see what the tech is truly capable of, but I'm aware this will lead to more global monopolies over and above the existing search engine monopolies allowed to dominate in a country or territory, which will harm smaller AI entities.
[deleted]
A lot of talk here on how copyrights and intellectual property are stupid, backwards, ancient (this is particularly strange, because the right for humans to exist and not be harmed is arguably much older and no one is complaining) and harmful without exploring a world without them. If I recall copyrights were introduced along with the development of mass print for the purpose of protecting authors of books. How will the world of today solve this issue ?
Putting my money on the copyrighters lose.
The US and its fears of China catching up aren't going to let their latest prized jewel OpenAI get kneecapped by "stop copying me!"s.
The US and its fears of China catching up aren't going to let their latest prized jewel OpenAI get kneecapped by "stop copying me!"s.
Copyright is for humans, not AI. Since the AI is doing 99.999% of the work on an "original work" and since the work is derived from the work of others, copyright should not apply. Copyright is for humans and creations of humans. If we get a truly sapient AI then we can revisit.
Creative works are incompatible with capitalism. We’ve create a thin finicky interface between them with copyright laws, but it hardly works. I’m not saying artists and creators shouldn’t have financial security in this system, quite the opposite. I don’t have a better idea, but I hope we can come up with something that doesn’t conflate ownership with attribution and also protects the livelihood of people who want to share their creations.
I’ve made so much money stacking my pitch decks and websites with AI generated media that I don't care if someone copy and pastes it and uses it commercially too
People married to their prompt engineering outputs are really missing the forest for the trees
People married to their prompt engineering outputs are really missing the forest for the trees
Imagine the future where we have instead of HDMI recording cards that circumvent crappy encryption something like a downscale and upscale neural co processor to circumvent copyright protection.
Last time they tried to copyright everything with DMCA it actually turned out to be a useless law.
Why not fix DMCA first and make malicious intent actually prosecutable - and then move on to transformative topics?
Can't fix copyright if copyright itself is already broken in the law. 90 years made sense in a world of pen and paper, but not in a world with the internet.
Last time they tried to copyright everything with DMCA it actually turned out to be a useless law.
Why not fix DMCA first and make malicious intent actually prosecutable - and then move on to transformative topics?
Can't fix copyright if copyright itself is already broken in the law. 90 years made sense in a world of pen and paper, but not in a world with the internet.
Hi HN, I have been working on something directly related to AI and copyright. Would it be ok to point it out here?
Recently The Pile was taken offline from The Eye by DMCA. One solution is to host it offshore, which we're calling The Nose: https://thenose.cc
The technical security measures may be of interest to the audience here, so I'll be as detailed as possible. The following formula should be safe if you follow it to the letter.
The basic setup is to install Whonix on a VeraCrypt drive, acquire Monero through any method, use a service like changenow to convert Bitcoin on a wallet stored only on the Whonix installation, sign up for a ProtonMail account (when they ask for email verification, use a no signup inbox service like yopmail), rent a dedicated server at Shinjiru using bitcoin, and register the domain at the same place. They're both a registrar and a server host, which simplifies matters. Use N/A for all contact info. Use Cloudflare to manage your site's DNS records.
Wallet security: do not ever move Bitcoin to any wallet linked with your personal identity. This is easier said than done. First there is the question of how to store passwords. These are the keys to the kingdom, and are the most sensitive aspect by far, because they're intimately linked with you. Additionally, if hardware failure occurs, you'll lose everything if you store them on the Whonix drive. My setup is to use KeePass to store the passwords on a laptop I use to VNC into the computer with the Whonix drive, and then save the database to a folder that gets synced to the cloud. The only flaw in this model is that if your laptop is compromised while your KeePass is open, you're done. But (as Ulbricht discovered) this is always true. The threat model assumes lawyers coming after you with DMCA with additional safeguards against the FBI narrowing down who you are in real life. If your physical location is compromised through any method, you're done.
All it takes is one mistake to end you. SSH into your box from your real computer? Done. Sign up using your real name with Mailgun? Done. Accidentally say "Thanks, <your real name>" to the support staff at Shinjiru in an email? Done. Abandon ship and close everything down.
The security of this technique comes down to simplicity. There are very few moving parts. I opted for nginx + mediawiki with Discourse forums at https://forums.thenose.cc (though I don't know if anyone will care enough to join). Logging is turned off to protect users downloading the data, though you only have my word on this. But reputation is the only thing a hacker has ever truly had anyway.
If you're serious about following the above recipe, I urge you to read through the Whonix docs on online anonymity: https://www.whonix.org/wiki/Documentation Remember, threat model is your saving grace. You probably aren't starting a darknet, so you can relax your threat model in terms of physical safety. But you won't get away with any mistakes made in cyberspace.
As for the site itself, I've avoided asking for donations for now (hosting is $130/mo though, which will get expensive) or describing anything beyond this HN comment. I'll say it's for simplicity, but in fact I only started it a few days ago and haven't had time to provide anything but the essence of our service: hosting AI datasets in stable, copyright-resistant ways.
If additional datasets beyond The Pile need protection or distribution, you can contact me at [email protected] or at https://forums.thenose.cc. I have a 4TB drive, of which 800gb is being used by The Pile so far.
Recently The Pile was taken offline from The Eye by DMCA. One solution is to host it offshore, which we're calling The Nose: https://thenose.cc
The technical security measures may be of interest to the audience here, so I'll be as detailed as possible. The following formula should be safe if you follow it to the letter.
The basic setup is to install Whonix on a VeraCrypt drive, acquire Monero through any method, use a service like changenow to convert Bitcoin on a wallet stored only on the Whonix installation, sign up for a ProtonMail account (when they ask for email verification, use a no signup inbox service like yopmail), rent a dedicated server at Shinjiru using bitcoin, and register the domain at the same place. They're both a registrar and a server host, which simplifies matters. Use N/A for all contact info. Use Cloudflare to manage your site's DNS records.
Wallet security: do not ever move Bitcoin to any wallet linked with your personal identity. This is easier said than done. First there is the question of how to store passwords. These are the keys to the kingdom, and are the most sensitive aspect by far, because they're intimately linked with you. Additionally, if hardware failure occurs, you'll lose everything if you store them on the Whonix drive. My setup is to use KeePass to store the passwords on a laptop I use to VNC into the computer with the Whonix drive, and then save the database to a folder that gets synced to the cloud. The only flaw in this model is that if your laptop is compromised while your KeePass is open, you're done. But (as Ulbricht discovered) this is always true. The threat model assumes lawyers coming after you with DMCA with additional safeguards against the FBI narrowing down who you are in real life. If your physical location is compromised through any method, you're done.
All it takes is one mistake to end you. SSH into your box from your real computer? Done. Sign up using your real name with Mailgun? Done. Accidentally say "Thanks, <your real name>" to the support staff at Shinjiru in an email? Done. Abandon ship and close everything down.
The security of this technique comes down to simplicity. There are very few moving parts. I opted for nginx + mediawiki with Discourse forums at https://forums.thenose.cc (though I don't know if anyone will care enough to join). Logging is turned off to protect users downloading the data, though you only have my word on this. But reputation is the only thing a hacker has ever truly had anyway.
If you're serious about following the above recipe, I urge you to read through the Whonix docs on online anonymity: https://www.whonix.org/wiki/Documentation Remember, threat model is your saving grace. You probably aren't starting a darknet, so you can relax your threat model in terms of physical safety. But you won't get away with any mistakes made in cyberspace.
As for the site itself, I've avoided asking for donations for now (hosting is $130/mo though, which will get expensive) or describing anything beyond this HN comment. I'll say it's for simplicity, but in fact I only started it a few days ago and haven't had time to provide anything but the essence of our service: hosting AI datasets in stable, copyright-resistant ways.
If additional datasets beyond The Pile need protection or distribution, you can contact me at [email protected] or at https://forums.thenose.cc. I have a 4TB drive, of which 800gb is being used by The Pile so far.
The narrative beauty of filling this up with GPT-created comments will be so incredibly sublime.
Terms of Use
You are prohibited from using the content of this site in "large language models" or any other usage for the purpose of "artificial intelligence".
Liquidated Damages
You are prohibited from using the content of this site in "large language models" or any other usage for the purpose of "artificial intelligence".
Liquidated Damages
This is _exactly_ how it looks when a society commits suicide.
My opinion — and note I’m a software engineer, not a lawyer — is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material. This would, I think, require the AI’s creator to secure a license for all of its sources that allows this sort of transformation and presentation. And further, a user of the AI would themselves require a license to use the output.
The alternative seems to be “anything goes”.