this post was submitted on 30 Aug 2026
80 points (97.6% liked)
Technology
87649 readers
3006 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
The idea that CSAM is used in training data is largely discredited. An AI doesn't need to have actual images of nude children to generate CSAM; it knows what humans look like naked, and it knows what children look like - it fills in the gaps from there.
Bullshit. They fed all that shit through in east african countries and the devil knows where else. Paying workers shit to id picture of csam and other abusive shit all day. You do not know what you are talking about, at best.
There was an article the other day saying that Grok used CSAM as training data.
https://arstechnica.com/tech-policy/2026/08/elon-musks-xai-used-child-porn-to-train-grok-models-lawsuit-says/
Looks like it's seriously reaching. I hope that poor women isn't being used by some unscrupulous ideologues. She is heading for trying times regardless.
It's important to note that to produce an AI-generated image of somebody, they do not need to exist in the training data beforehand. You can use an image as a prompt, and an AI can create new works based on that prompt, without ever referencing the person from training materials.
Most of these CSAM files are not readily available on the clearnet, and I don't think even Elon is stupid enough to let an AI scraper run free on Tor. One would have to go significantly out of their way to locate clearnet sources of CSAM to include into the training data.
Some would say it's a distinction without difference, but I think it's important to understand how these AI images are actually created, if you want to adequately legislate them.
The TL;DR is that there is a voluntary system that tracks court cases and will alert CSAM victims when it suspects images of them (either known images, or new ones) are the topic of criminal investigations. A Jane Doe was repeatedly raped in the early 2000’s to produce copious amounts of CSAM, which was broadly shared among pedophile groups. Her images were apparently very popular and prolific. Doe was alerted that Grok was producing new explicit images of her as a child. The images in question are novel (generated by AI, and not matching any known existing images) but are undoubtedly of Doe as a child. Meaning it was using Doe’s images as part of the training data for the images it generates.
Essentially, if Grok was trained on CSAM, Doe’s popular images were almost certainly included. And now Grok is producing new images of her. It would be like asking Grok to make an image of a popular egirl, and getting an image of what is undoubtedly Belle Delphine. And then X goes “no no no, we didn’t use any images of Belle Delphine in our training data. We promise!”
that is so wrong, don't spread that bullshit
they find it in models all the time!