Got it, copy right I'd dead then Right?
Technology
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots

Therefore copyrights mean nothing
The main challenge I see is that Aaron Schwartz and countless others have been prosecuted for access to information but AI companies have been rewarded. There is precedence that access like this is not legal. Had they gone through a library system or used a mechanic like that it may have worked but from what I understand, they used torrents and other mechanics to access the data. So you have companies that go after individual infringement but pursue their own mass infringement. Whether AI generated materials is infringement is above my pay grade but their consumption of the materials seems pretty straight forward as infringement.
But paying and asking for permission first would have been costly at the start and slow. They made a decision, maybe at the beginning, maybe as they realized ethical would mean lost time and placement in the race, and they said screw it, let's go. They also sidelined any AI safety research they were doing (some were making an effort on a difficult problem, but again, $$$ wins).
The NY Times and Pearson, two of the most valuable US publishers, each have market caps of about $10 billion.
Let's say Pearson went after OpenAI. They devote an unlimited legal budget to the fight. OpenAI is hoping to IPO as a trillion dollar company. If Pearson went after OpenAI, rather than fight them in court, a deal could be reached first. If that didn't work, if a deal couldn't be reached, OpenAI could bypass the problem completely:
- Spend $5 billion to buy a controlling share of Pearson.
- Fire the entire leadership team and install OpenAI minions in their place.
- Once they control Pearson, sign a long-term licensing deal with OpenAI with very generous licensing terms and huge early cancellation fees.
- Sell the shares back on the market at a (likely slightly reduced) value.
OpenAI would likely have to spend some money on net. The value of Pearson stock would likely be a bit lower after effectively giving away the rights to their works as training data. If signing a durable rights contract the new owners can't escape isn't practical, buying the company and simply holding it indefinitely would also be an option.
Schwartz was saving and distributing copies against the terms of the agreement by which he was able to access journals. What happened to him was heinous but it was pretty dissimilar to how models train on data. And the tormented material was Anthropic, which resulted in the largest copyright settlement in history. Because it was piracy. They briefly tried an argument that their intended use made it fair use, but...that's never how literally any of that worked.
All I'm hearing is piracy is legal now 🏴☠️
If you have copyright juggernauts gatekeeping the data through legislation, then you can't have open-source models.
I'm surprised how many want to shoot consumers in the face just to protect mafia like companies like Universal Studios and YouTube.
Remember how china said that a few months ago and the US was all like "communism is stealing our data"
Project Hail Mary but not for the good of all ...
lol
We’re in danger
Can you reasonably train a local model? Inference is one thing and can be run on a decent gaming graphics card but training requires a lot more compute.
You can absolutely train a nontrivial task model on a gaming gpu (image classifier, sound classifier, text model with a very structured input and output, etc). You can also post-train addon layers on top of existing open-weight models (LORA).
Can you train it better than anything already available on model zoos for free though? If not then beside learning I dont see the point.
Depends on the size of the training data, how big the resulting model is, and how fast you want it to finish.
Whatever, fuck AI slop.