this post was submitted on 29 Aug 2026
272 points (97.6% liked)

Technology

87649 readers
3891 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] grue@lemmy.world 217 points 19 hours ago (14 children)

The entire fucking point of the Web was to make information as easily-accessible as possible, structured and semantically tagged, and consumable by humans and further machine transformation alike. "Scraping" is facilitated by design!

Using Javascript to deliberately break that is evil and every programmer who participates it is a piece of shit. No exceptions.

[–] artyom@piefed.social 55 points 17 hours ago* (last edited 17 hours ago) (10 children)

That's true but they probably didn't account for AI data scrapers ramfucking your server so they could steal all the value you assembled for general consumption and serve it themselves for profit.

[–] grue@lemmy.world 54 points 17 hours ago (9 children)

The scraping wouldn't be a problem if Reddit simply provided an RSS feed or other data-efficient API. The "ramfucking" is caused by the attempt to block bots; it is entirely self-inflicted.

Remember, it's all our content to begin with and Reddit does not have any right to try to lock it up for itself.


That doesn't mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.

[–] artyom@piefed.social 39 points 17 hours ago (1 children)

The scraping wouldn't be a problem if Reddit simply provided an RSS feed or other data-efficient API

That's simply not true. These bots are essentially DDOSing the entire internet, API or not.

[–] grue@lemmy.world 17 points 16 hours ago

Okay, if efficient APIs existed and they weren't incompetently failing to use them, it wouldn't be a problem. Happy now?

(I should've addressed that in my previous comment, as I was aware of how one of the Lemmy instances was taken down by scrapers the other day despite the fact that they could easily get all the content simply by consuming ActivityPub directly. But I was naively hoping it wouldn't be necessary because, as you can see from this text, it would've cluttered up my writing with double the words.)

load more comments (7 replies)
load more comments (7 replies)
load more comments (10 replies)