this post was submitted on 23 Sep 2023
209 points (91.0% liked)
Technology
59105 readers
4009 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related content.
- Be excellent to each another!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, to ask if your bot can be added please contact us.
- Check for duplicates before posting, duplicates may be removed
Approved Bots
founded 1 year ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
This analogy falls apart when you note that "overgrazing" these resources does absolutely nothing to harm them.
They're still there. They haven't been affected in any way by the fact that a machine somewhere has read them and learned a bunch of stuff from them. So what?
Only if you consider AI-supercharged misinformation to not be harmful.
Only if you consider the entropy of human interaction on the internet to not be harmful.
Only if you consider being unable to know who is real to not be harmful.
None of those things directly harm the resources being "grazed", and none of them are inevitable consequences of AI. If you think they are then you're actually arguing against AI in general and not the specific way in which they've been trained.
You think the internet being flooded with articles, comments etc. all being written by AI whose only goals are selling shit, disseminating misinformation, and manipulating elections and opinions - with no way to know what is human and what is AI - is going to be a great environment to continue to train your AI?
You might be interested to read about Model Autophagy Disorder.
That is not a problem caused by "overgrazing" those open resources. It's a separate problem with AI training that needs to be addressed anyway. You're just throwing out random AI-related challenges regardless of whether they're relevant to what's being discussed.
Simply put, quality control is always important.
If you pump toxic waste onto the field nobody gets to graze it.
Fuck you're being pedantic.
And you're completely missing the point.
Whether or not toxic waste is pumped into the field is completely independent of whether anyone is "grazing" on it. AIs are going to be trained and AIs are going to be generating content, regardless of whether those "commons" are being used as training material. If you wish to keep those "commons" high-quality you're going to have to come up with some way of doing that regardless of whether they're being used as training material. Banning the use of them as training material will have no impact on whether they get "toxic waste" pumped into it.
My objection is to those who are saying that to save the commons we need to prevent grazing, ie, that to save the quality of public discourse we need to prevent AIs from training on it. Those two things are unrelated. Stopping AIs from training on it will not do anything to preserve the quality of public discourse.
No mate, you're just being pedantic
And you're still keen to argue two days after the thread's moved off of everyone's feed.
Yeah man I login a couple of times a day to check notifications
While the analogy is not perfect, you can think that the harm is getting lost in the noise. If the "overgrazing" of content on the internet (content which has the purpose of being read/listened/etc. Often for a job) causes a huge amount of other content based on it (AI-generated), then the original is damaged by being lost in the noise.
AI-generated content is coming regardless, whether those open sources get "grazed" or not.
Yes, bit the qualitative difference of providing direct competition to the "grazed" material exists. There is a difference between AI generated audiobooks and AI generated audiobooks with the voice of X, for X. Once AI can perfectly reproduce X's voice, his/her value as a voice actor is 0, hence the "overgrazing". Is not the same thing compared to simply being able to provide audiobooks with any other voice.
api ._.
That was entirely self-inflicted damage.
?
Reddit is responsible for their own API changes. Not OpenAI or any other external agency who might have been using Reddit data for AI training. Only Reddit was capable of choosing to change their API, it's entirely under their control.