this post was submitted on 28 Aug 2026
926 points (98.7% liked)
Technology
87627 readers
3732 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
In all seriousness, there are some very interesting legal questions that will be raised if this case makes it that far.
The problem is that there's no existing law that would effect this on its own. To my knowledge, no country in the world has a law on the books specifically dealing with AI models trained on CSAM. So the question, under existing laws, would turn on whether the data stored within the model itself would constitute CSAM.
The problem, in no small part, is that we have serious gaps in our public consensus knowledge about how LLMs actually work.
There's a case that, AFAIK, is still being argued in Germany pushing the theory that LLMs actually do, in effect, store a copy of all their training data, just in a compressed form. This certainly seems to hold some water given both the tests they relied on, and the situation with this Jane Doe where the model produced images so alike to real images of her that they tripped hash detections.
The German case argues that this is analogous to the difference between an MP3 and a WAV, or a JPEG and a PNG. That sharing a lossy copy of a work is no less infringing just because it's imperfect.
If the underlying claim - that LLMs function as a form of lossy compression - can be substantiated then there would be a real argument that the model itself would constitute CSAM. Since there would be no realistic method that I'm aware of for removing the offending material from the model - and presumably SpaceX would have to somehow prove that they've done so - that would make the entire model contraband. They'd have to retrain on a clean dataset.
Of course I said "if the case makes it that far" at the top because I don't think it will. SpaceX will do anything and everything to avoid handing over meaningful discovery in this case, including, I suspect, outright destruction of evidence. If there is anything that actually proves that they used CSAM in the training data then they are so far beyond fucked that there's simply no downside to further illegality in pursuit of concealing their crimes. They have the world's wealthiest asshole in a position to throw literal billions at making this go away. I genuinely wouldn't be surprised if people turn up dead off the back of this if that's what it takes.
CSAM laws are already overreaching in dimensions that are rife with abuse. It's rare to have a law where owning a picture of something is highly illegal. It's the only law I know of where somebody can just send you a picture on your phone, and suddenly, you're breaking the law. You could be arrested and thrown in prison for even admitting that somebody else sent you the picture.
Police use CSAM as an excuse all the time to search without a warrant. Remember the duress passcode case? Police claimed he had CSAM on his phone, when everybody knew they were targeting his activist work.
And what should we really care about? The CSA. Go after the CSA. Focus on the abuse. The Epstein files are right over there.
Yeah, it's absolutely valid to question the degree to which this is an outcome that we should even want. If someone hides a CSAM image in the Linux kernel should Linux become illegal?
I think there is a valid distinction to be drawn in this case, because removing individual components of an LLM isn't really something we know how to do. So there's a fair argument that a model which is built using illegal content should be illegal, and if that means they have to completely retrain from scratch, so be it.
But yes, I'd want to be very careful about the lines around a law or ruling like that and exactly what it's extent is. Child sex crimes and child safety are topics that tend to short-circuit all reasonable objections, and are frequently exploited as a means of getting bad laws onto the books. Bill C-22 up here in Canada is a great recent example.
Globally, there's powers at work to get all kinds of "age verification" laws on the books, which we all know is just a thin excuse for de-anonymization and identity gathering. These same powers were fucking with Steam and Itch.io's payment methods, because "what about the children", which then ties to similar events with PornHub a few years earlier. Congress passed three different "Internet child safety" laws in the 2000s and 2010s, and all three were struck down by the Supreme Court for being unconstitutional. Decades earlier, DARE abused "child safety" justifications for personal gains.
This excuse is used all the time, and the emotional weight it carries makes it shocking effective to the proles that buy it at face value.