this post was submitted on 29 Aug 2026
346 points (97.3% liked)

Technology

87685 readers
3024 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] grue@lemmy.world 63 points 1 day ago (12 children)

The scraping wouldn't be a problem if Reddit simply provided an RSS feed or other data-efficient API. The "ramfucking" is caused by the attempt to block bots; it is entirely self-inflicted.

Remember, it's all our content to begin with and Reddit does not have any right to try to lock it up for itself.


That doesn't mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.

[–] Toga77@lemmy.world 8 points 1 day ago (4 children)

But we all know AI companies are unethically scraping and selling shit back to us right? I just really need people to acknowledge that.

[–] grue@lemmy.world 0 points 1 day ago (3 children)

It is the "selling shit back to us" specifically, not the "scraping," that's the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue "no."

[–] rudyharrelson@lemmy.radio 2 points 11 hours ago (1 children)

The "scraping" part becomes unethical when the scraping is so aggressive that it takes down the website (or severely impacts its ability to serve actual clients).

Archive.org scrapes the web all the time, but it doesn't do it so aggressively that it becomes an issue for the websites they're scraping. The same cannot be said for AI scrapers.

[–] grue@lemmy.world 1 points 6 hours ago

Scraping more than necessary is so stupid that I just sort of dismissed it as a straight-up mistake that will eventually be corrected. I was arguing based on general principle, not specific current practice.

Obviously, yes, the AI companies should fix their (probably vibe-coded) scrapers so they stop misbehaving; that should've gone without saying.

load more comments (1 replies)
load more comments (1 replies)
load more comments (8 replies)