Data helped build ChatGPT. Where's your payout?(theatlantic.com)
theatlantic.com
Data helped build ChatGPT. Where's your payout?
https://www.theatlantic.com/technology/archive/2023/03/open-ai-products-labor-profit/673527/
8 comments
I am very sympathetic to the idea that people doing labelling tasks in third world countries deserve more money. But I don't buy that I'm owed something for my tiny contribution to GPT-3's 300B tokens of training data, or Llama's 1.4T tokens.
> I don't buy that I'm owed something for my tiny contribution to GPT-3's 300B tokens of training data
I feel like this touches on one of the core issues in this discussion. Any individuals contribution is negligible. The sum of all those contributions definitely is not.
It's like stealing one cent from every bank account on the planet. I doubt anyone would go to the police about their missing cent, let alone any police officer actually investigating such a report. And yet ~70 million was stolen.
I feel like this touches on one of the core issues in this discussion. Any individuals contribution is negligible. The sum of all those contributions definitely is not.
It's like stealing one cent from every bank account on the planet. I doubt anyone would go to the police about their missing cent, let alone any police officer actually investigating such a report. And yet ~70 million was stolen.
Much like traditional mining, you get incredibly large companies doing everything they can to avoid paying society back for what they've taken.
>What do we do? The penny is the smallest bit of USD we have, and hyper-fractional parts that incrementally make up an unimaginably large whole are now the world we live in. It's difficult to imagine a world where you receive 175 billion royalty transactions of 1/1,000,000,000th of a $0.01 in a given year, but maybe that's a reasonable scenario to think about when it comes to the couple of bucks an average teen or adult should get from their presumed default contributions to large language models.
Remember trying to hit that minimum threshold payment for Google Adsense with your Blogspot, then finally getting your check after 15 years? If nothing else, we shouldn't blithely tolerate that again, because we didn't sign up for this. (We signed up for stuff even worse than this, technically, but in those cases at least we clicked "I Agree.")
Remember trying to hit that minimum threshold payment for Google Adsense with your Blogspot, then finally getting your check after 15 years? If nothing else, we shouldn't blithely tolerate that again, because we didn't sign up for this. (We signed up for stuff even worse than this, technically, but in those cases at least we clicked "I Agree.")
Very interesting the title on here drops the "YOUR" from the statement as it is sort of the whole point of the article.
That's probably the HN automatic title re-writer, not some nefarious plot.
Yes. Exactly.
Fair use cited to scrape the web & some pathetic clause of not to use the output to train a competitor is just horseshit
Data helped build the knowledge set of a lot of smart people too. Should the people who put the data out there also get paid?
I think there's a fundamental difference between using the data to create a product and using the data to improve people's minds. I get a return on my "investment" in the latter case because my neighbors being educated also improves my own standard of living.
That said, I'm not really on board with these "pay for using my data" arguments. I think they're just trying to work around that the use of people's data for things like gpt troubles them and they wish they could have stopped it.
Or, maybe I think that because the fact that my data was used to help train gpt bothers me greatly. No payout will fix that. So, instead, I've taken down all my public-facing websites and data repositories to prevent being forced to help these efforts in the future.
It's not much, but it's all I can do at this point.
That said, I'm not really on board with these "pay for using my data" arguments. I think they're just trying to work around that the use of people's data for things like gpt troubles them and they wish they could have stopped it.
Or, maybe I think that because the fact that my data was used to help train gpt bothers me greatly. No payout will fix that. So, instead, I've taken down all my public-facing websites and data repositories to prevent being forced to help these efforts in the future.
It's not much, but it's all I can do at this point.
It's not even clear to me that it's legal to use web crawls of copyrighted content to train an LLM.
It is, in the EU, UK, JP and SG.
If we wanted a return for our contributions we should have enacted sensible data ownership and privacy laws a decade ago. That ship has long sailed, at least for this generation of AI tooling.