this post was submitted on 07 Oct 2024

565 points (98.8% liked)

Technology

63082 readers

3516 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

[email protected]

565

TikTok’s parent launched a web scraper that’s gobbling up the world’s online data 25-times faster than OpenAI (fortune.com)

submitted 4 months ago by [email protected] to c/[email protected]

123 comments fedilink hide all child comments

top 50 comments

sorted by: hot top controversial new old

[–] [email protected] 289 points 4 months ago (17 children)

It's illegal when a regular person steals something, but it's innovation and courage, when a huge corporation steals something. Interesting how that works

[–] [email protected] 105 points 4 months ago (11 children)

Honestly it’s fucking angering. So much regulation and geo-restrictions and licensing schemes… but it’s cool that there are data brokers, and shit like this. On top of it all Chrome screwing us with manifest v3 and killing ad blocking on chrome. It’s already in canary build.

WHAT THE FUCK IS WRONG WITH THIS SPECIES?!

[–] [email protected] 26 points 4 months ago

WHAT THE FUCK IS WRONG WITH THIS SPECIES?!

Yes.

[–] [email protected] 23 points 4 months ago

WHAT THE FUCK IS WRONG WITH THIS SPECIES?!

Capitalism.

load more comments (9 replies)

[–] [email protected] 53 points 4 months ago (1 children)

They're not stealing your data, they're pirating it.

[–] [email protected] 33 points 4 months ago (1 children)

They’re not pirating it. They’re collecting it.

[–] [email protected] 23 points 4 months ago (1 children)

They’re not collecting it. They’re archiving it.

[–] [email protected] 14 points 4 months ago (1 children)

Oh, like the way back machine?

[–] [email protected] 34 points 4 months ago (2 children)

No, that's stealing /s

load more comments (2 replies)

[–] [email protected] 36 points 4 months ago (3 children)

Aaron Schwartz killed himself over punishments for less

[–] [email protected] 21 points 4 months ago* (last edited 4 months ago)

Worse punishments. For far less.

[–] brbposting 11 points 4 months ago

RIP Aaron

load more comments (1 replies)

load more comments (14 replies)

[–] [email protected] 104 points 4 months ago (1 children)

We've had this thing hammering our servers. The scraper uses randomized user-agents browser/OS combinations and comes from a number of distinct IP ranges in different datacenters around the world, but all the IPs track back to Bytedance.

[–] [email protected] 38 points 4 months ago (1 children)

Wouldn't be surprised if they're just cashing out while TikTok is still public in the US. One last desperate grab at value-add for the parent company before the shut down.

Also a great way to burn the infrastructure for subsequent use. After this, you can guarantee every data security company is going to add the TikTok servers to their firewalls and blacklists. So the American company that tries to harvest the property is going to be tripping over these legacy bullwarks for years after.

[–] [email protected] 13 points 4 months ago

This has nothing to do with Tik Tok other than ByteDance being a shareholder in Tik Tok

[–] [email protected] 82 points 4 months ago (3 children)

Also it doesn't respect robots.txt (the file that tells bots whether or not a given page can be accessed) unlike most AI scrapping bots.

[–] kboy101222 52 points 4 months ago (7 children)

My personal website that primarily functions as a front end to my home server has been getting BEAT by these stupid web scrapers. Every couple of days the server is unusable because some web scraper demanded every single possible page and crashed the damn thing

[–] assaultpotato 16 points 4 months ago (1 children)

I do the same thing, and I've noticed my modem has been absolutely bricked probably 3-4 times this month. I wonder if this is why.

load more comments (1 replies)

load more comments (6 replies)

[–] [email protected] 21 points 4 months ago (1 children)

most

I doubt that...

load more comments (1 replies)

[–] Dindonmasker 57 points 4 months ago (8 children)

Not surprising that Bytedance would want to gobble up every bit of data they can as fast as possible.

load more comments (8 replies)

[–] [email protected] 35 points 4 months ago (5 children)

They're too late, there's going to be way too much AI generated garbage in their data and so many social media platforms like Reddit and Twitter have already taken measures to curb scrapers.

[–] chickenf622 18 points 4 months ago

Like those platforms aren't already full of AI garbage as well. Training new models will require a cut-off date before the genie was let out of the bottle.

load more comments (4 replies)

[–] [email protected] 32 points 4 months ago* (last edited 4 months ago) (2 children)

As for what ByteDance plans to do with a new LLM, a person familiar with the company’s ambitions said one goal has to do with the search function for TikTok.

Last week, TikTok released an update to its current search function focused on [keywords for ads], basically allowing advertisers to search in real time for words that are trending on TikTok. It allows marketers to build an ad with relevant keywords that would ostensibly help the ad show up on the screens of more users.

…

“Given the audience and the amount of use, TikTok with a search environment that is a completely biddable space with keywords and topics, that would be very interesting to a lot of people spending a ton of money with Google right now,” the person said.

A dark vision just flashed in my mind. And I am certain this is what will happen. AI-generated ads done in real time based on the latest “trending” thing. Presented to users basically as soon as the topic has the slightest amount of “trend”.

Just emitting untold amounts of CO2 to show you generated ads in near real time.

[–] [email protected] 14 points 4 months ago

No wonder Google ex-CEO was saying fuck climate goals.

load more comments (1 replies)

[–] [email protected] 31 points 4 months ago (3 children)

There it begins. Nothing good will ever come form this.

load more comments (3 replies)

[–] [email protected] 26 points 4 months ago (2 children)

I can not contribute to anything here, I just came to say I really really like the phrase "gobbling something up" :D

load more comments (1 replies)

[–] [email protected] 21 points 4 months ago (2 children)

from the article:

Robots.txt is a line of code that publishers can put into a website that, while not legally binding in any way, is supposed to signal to scraper bots that they cannot take that website’s data.

i do understand that robots.txt is a very minor part of the article, but i think that’s a pretty rough explanation of robots.txt

[–] Corkyskog 11 points 4 months ago (4 children)

Out of curiosity, how would you word it?

[–] [email protected] 19 points 4 months ago

i would probably word it as something like:

Robots.txt is a document that specifies which parts of a website bots are and are not allowed to visit. While it’s not a legally binding document, it has long been common practice for bots to obey the rules listed in robots.txt.

in that description, i’m trying to keep the accessible tone that they were going for in the article (so i wrote “document” instead of file format/IETF standard), while still trying to focus on the following points:

robots.txt is fundamentally a list of rules, not a single line of code
robots.txt can allow bots to access certain parts of a website, it doesn’t have to ban bots entirely
it’s not legally binding, but it is still customary for bots to follow it

i did also neglect to mention that robots.txt allows you to specify different rules for different bots, but that didn’t seem particularly relevant here.