this post was submitted on 02 Sep 2026
539 points (99.3% liked)

Technology

87814 readers
3457 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
(page 3) 50 comments
sorted by: hot top controversial new old
[–] Smaile@lemmy.ca 1 points 1 day ago

Got it, copy right I'd dead then Right?

[–] stumu415@lemmy.zip 1 points 1 day ago
[–] TropicalDingdong@lemmy.world 11 points 2 days ago
[–] Nagrom@lemmy.ca 1 points 1 day ago

Therefore copyrights mean nothing

[–] assembly@lemmy.world 10 points 2 days ago (3 children)

The main challenge I see is that Aaron Schwartz and countless others have been prosecuted for access to information but AI companies have been rewarded. There is precedence that access like this is not legal. Had they gone through a library system or used a mechanic like that it may have worked but from what I understand, they used torrents and other mechanics to access the data. So you have companies that go after individual infringement but pursue their own mass infringement. Whether AI generated materials is infringement is above my pay grade but their consumption of the materials seems pretty straight forward as infringement.

[–] Rhaedas@fedia.io 4 points 2 days ago

But paying and asking for permission first would have been costly at the start and slow. They made a decision, maybe at the beginning, maybe as they realized ethical would mean lost time and placement in the race, and they said screw it, let's go. They also sidelined any AI safety research they were doing (some were making an effort on a difficult problem, but again, $$$ wins).

The NY Times and Pearson, two of the most valuable US publishers, each have market caps of about $10 billion.

Let's say Pearson went after OpenAI. They devote an unlimited legal budget to the fight. OpenAI is hoping to IPO as a trillion dollar company. If Pearson went after OpenAI, rather than fight them in court, a deal could be reached first. If that didn't work, if a deal couldn't be reached, OpenAI could bypass the problem completely:

  1. Spend $5 billion to buy a controlling share of Pearson.
  2. Fire the entire leadership team and install OpenAI minions in their place.
  3. Once they control Pearson, sign a long-term licensing deal with OpenAI with very generous licensing terms and huge early cancellation fees.
  4. Sell the shares back on the market at a (likely slightly reduced) value.

OpenAI would likely have to spend some money on net. The value of Pearson stock would likely be a bit lower after effectively giving away the rights to their works as training data. If signing a durable rights contract the new owners can't escape isn't practical, buying the company and simply holding it indefinitely would also be an option.

[–] FatCrab@slrpnk.net 2 points 1 day ago

Schwartz was saving and distributing copies against the terms of the agreement by which he was able to access journals. What happened to him was heinous but it was pretty dissimilar to how models train on data. And the tormented material was Anthropic, which resulted in the largest copyright settlement in history. Because it was piracy. They briefly tried an argument that their intended use made it fair use, but...that's never how literally any of that worked.

[–] scottmeme@sh.itjust.works 8 points 2 days ago

All I'm hearing is piracy is legal now 🏴‍☠️

[–] Grimy@lemmy.world 7 points 2 days ago

If you have copyright juggernauts gatekeeping the data through legislation, then you can't have open-source models.

I'm surprised how many want to shoot consumers in the face just to protect mafia like companies like Universal Studios and YouTube.

[–] Alvaro@lemmy.blahaj.zone 7 points 2 days ago

Remember how china said that a few months ago and the US was all like "communism is stealing our data"

[–] einkorn@feddit.org 8 points 2 days ago (1 children)

Project Hail Mary but not for the good of all ...

[–] snooggums@piefed.world 12 points 2 days ago (1 children)
load more comments (1 replies)
[–] homesweethomeMrL@lemmy.world 6 points 2 days ago

lol

We’re in danger

[–] Not_mikey@lemmy.dbzer0.com 3 points 1 day ago (2 children)

Can you reasonably train a local model? Inference is one thing and can be run on a decent gaming graphics card but training requires a lot more compute.

[–] Dran_Arcana@lemmy.world 4 points 1 day ago (1 children)

You can absolutely train a nontrivial task model on a gaming gpu (image classifier, sound classifier, text model with a very structured input and output, etc). You can also post-train addon layers on top of existing open-weight models (LORA).

[–] utopiah@lemmy.world 1 points 1 day ago

Can you train it better than anything already available on model zoos for free though? If not then beside learning I dont see the point.

[–] frongt@lemmy.zip 4 points 1 day ago

Depends on the size of the training data, how big the resulting model is, and how fast you want it to finish.

[–] lechekaflan@lemmy.world 2 points 1 day ago

Whatever, fuck AI slop.

load more comments
view more: ‹ prev next ›