this post was submitted on 10 Jan 2024
1243 points (96.6% liked)
Technology
61081 readers
2802 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related content.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, to ask if your bot can be added please contact us.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 2 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
Actually content laundering is the best term I've heard to describe the process. Just like money laundering, you no longer know the source and know it's technically legal to use and distribute.
I mean, if the copyrighted content wasn't so critical, they would train models without it. Their essentially derivative works, but no one wants to acknowledge it because it would either require changing our copyright laws or make this potentially lucrative and important work illegal.
Content laundering is not a good way to describe it because it's misleading as it oversimplifies and mischaracterizes what a language model actually does. It's a fundamental misunderstanding of how it works. Training language models is typically a transparent and well-documented process as described by the mountains of research over the past decades. The real value comes from the weights of the nodes in the neural network and not the source that it spits out in its entirety when it was trained. The source material is evaluated and wholly transformed into new data in the form of nodes and weights. The original content does not exist as it was within the network because there's no way to encode it that way. It's a statistical system that compounds information.
And while LLMs do have the capacity to create derivative works in other ways, it's not all that they do, or what they always do. It's only one of the many functions that it has. What you say would probably be true if it was only trained on a single source, but that's not even feasible. But when you train it on millions of sources, what remains are the overall patterns of language within those works. It's much more sophisticated and flexible than what you describe.
So no, if it was cut and dry there would be grounds for a legitimate lawsuit. The problem is that people are arguing points that do not apply but sound reasonable when they haven't seen a neural network work under the hood. If anything, new laws need to be created to address what LLMs do if you're so concerned about proper compensation.
I am familiar with how LLMs work and are trained. I've been using transformers for years.
The core question I'd ask is, if the copyrighted material isn't essential to the model, why don't they just train the models without that data? If it is core to the model, then can you really say they aren't derivative of that content?
I'm not saying that the models don't do something more, just that the more is built upon copyrighted material. In any other commercial situation, you'd have to license/get approval for the underlying content if you were packaging it up. When sampling music, for example, the output will differ greatly from the original song, but because you are building off someone else's work you must compensate them.
Its why content laundering is a great term. The models intermix so much data that it's hard to know if the content originated from copyrighted materials. Just like how money laundering is trying to make it difficult to determine if the money comes from illicit sources.