this post was submitted on 07 Jul 2026
1550 points (99.4% liked)
Technology
87575 readers
3768 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
Aren't these data centers built on our land, right next to our homes, using our water and electricity, funded by selling our data (that we didn't consent to), and the profits all go to giant corporations, not us? Fuck 'em. I hope they all get raided and stripped bare...
Do you believe we should be allowed to run open source / weight LLMs like deepseek locally, for our own gain, even though they too have been indirectly trained on our comments / articles / copyrighted books?
I'm sure everyone will have varying opinions on it, but if these models were fully open-sourced I'd have less of an issue. What does it for me is that these companies were unwilling to pay creators to use their work in the training data, instead choosing to pirate it to create expensive and locked down AI products which they expect us to pay for.
I believe it was logistical impossible. Many books used will probably be scanned and not even be available as ebook or drm protected or out of print. And e.g. Anna's archive has 64,416,225 books and 95,689,473 papers. Too large to even say what is pirated or nor, or buy every book in a lifetime. And if every book costs you ~$10 that's close to a billion dollars upfront. Basically creating LLMs wouldn't have been possible without piracy (or maybe the datasets aren't actually that extensive).
It's hypocritical, but ultimately the same argument for piracy that individuals use: IP laws creates unreasonable restrictions that prevent people from learning (or enjoy culture at a sensible price).
(I assume you're not saying you would need a negotiate a specific license to use a book or a public comment or article for machine learning).
Kinda reminds me of Year Zero by Robert Reid. The whole galaxy full of aliens loves Earth Music but only recently figure the concept of copyright. And how much quintillion moneys they now owe Earth lol.