this post was submitted on 28 Aug 2026
938 points (98.8% liked)
Technology
87649 readers
3548 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
I knew Elon was a pedophile (he begged to be on the island) but to actually train your AI on CSAM...
Well he should be arrested for possession of child pornography and Grok should be shut down until all of its CSAM is erased from its databanks
Unfortunately that's not really how the actual data is stored. Functionally there is no CSAM at all in its "databanks". Once Info goes through training what comes out the other side is just a mass of goop.
It's like if you took an entire cow ran it though a meat grinder. Then demanded that you remove only the chuck from the resulting ground beef.
It's not physically possible.
You can demand it's retrained entirely from the ground up with vetted data. To produce a higher quality clean dataset. And I would agree that is what should be done.
But you can't unground the beef
Do you really believe they just throw all the training data away after use?
on the one hand, it makes sense if your goal is to train AI to recognize child porn as a simple binary state (bool isCP). Social media sites used to have humans looking at that stuff moderating from afar and it really takes a horrendous toll on their employees.
On the other hand, Elon has repeatedly shown he refuses to censor child porn. They didn’t train it to stop making kiddie porn. They trained it to create more kiddie porn. And that’s why he’s rich. Elon won’t say no. He doesn’t care.
The detail about hash values is really important. The FBI maintains a database of known CSAM. Presumably, hers is in that database, hence the hash values. While not everyone has access to that DB, Xitter/etc does. There is no ambiguity of anything on that list; there's also no need for any human to review. It's already been confirmed.
While I'm not sure there's any case law about it, I would be amazed if using that to train generative AI (except POSSIBLY as content to block) was treated as anything other than possession, distribution, and maybe even production of CSAM.
Proving it might be difficult without full discovery, and AI is infamous for the massive corpus of training data. However, AI doesn't always generate truly unique works. Go to any image generator and prompt for a video game plumber, and you'll see an unmistakable image of Mario. It's possible that they can find a prompt that generates results close enough to her images.
Prompting AI for a video game plumber is most likely going to result in a Mario like entity just because it's the most common example.
That's a poor example of what your talking about.
You need someone niche and narrow that has a much smaller sample size. It would be more like asking for a generic description of one persons fursona with out naming it. And the model spitting out a almost 1:1 copy of a real preexisting art work on the fursona. Because the model only has that one picture to base things off of. Which is a real problem.
It's also the best way to tell if a model was trained on something specific. General prompts aren't going to get you anything beyond just the fact that yes. X thing is popular enough that everyone and their grandmother creates content on it.