openmic.social is an uncensored community. You may encounter strong language, controversial opinions, and mature or NSFW material. You must be 18+ to browse. Illegal content is prohibited and removed on sight — please report it. By continuing, you accept that you may see content you personally disagree with.
No one is saying they output of an llm is copyright (this is a different debate).
What people are talking about is they stole copyright work. They pirated from torrent sights mass amounts of data. That’s the violation at problem.
The output of the training, the model (LLM) itself, is transformative. I’m not talking about the output of the llm nor was the person I was responding to.
The pirating issue has nothing to do with the current lawsuit (the New York Times one mentioned in the article) nor did the government mention it in the current context.
That being said, I’m a pirate in my personal life, and I also think it’s kind of a necessary evil when it comes to building SOTA models. I wish we had a copy left solution to it, so you can use pirated data but need to open source it, so open source has a fighting chance and big AI can buy it if they want it. I’m okay with openAI and anthropic getting brought into court about it I guess, but the legislation ends up being market capture in their favor in the end.
I understand the issues though, very “rule for thee but not for me” currently.