this post was submitted on 29 Aug 2026
346 points (97.3% liked)
Technology
87685 readers
3024 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
The scraping wouldn't be a problem if Reddit simply provided an RSS feed or other data-efficient API. The "ramfucking" is caused by the attempt to block bots; it is entirely self-inflicted.
Remember, it's all our content to begin with and Reddit does not have any right to try to lock it up for itself.
That doesn't mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.
But we all know AI companies are unethically scraping and selling shit back to us right? I just really need people to acknowledge that.
It is the "selling shit back to us" specifically, not the "scraping," that's the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue "no."
The "scraping" part becomes unethical when the scraping is so aggressive that it takes down the website (or severely impacts its ability to serve actual clients).
Archive.org scrapes the web all the time, but it doesn't do it so aggressively that it becomes an issue for the websites they're scraping. The same cannot be said for AI scrapers.
Scraping more than necessary is so stupid that I just sort of dismissed it as a straight-up mistake that will eventually be corrected. I was arguing based on general principle, not specific current practice.
Obviously, yes, the AI companies should fix their (probably vibe-coded) scrapers so they stop misbehaving; that should've gone without saying.