isleepinahammock

joined 2 years ago
[–] isleepinahammock@lemmy.blahaj.zone 1 points 1 month ago (1 children)

True, but irrelevant. Why would they care at all about the terms of a license? Again, they'll happily engage in outright mass torrenting. Their legal theory is that using works to train LLMs is simply fair use. The ebook sellers may disagree, but it won't stop them.

You're confusing uniqueness with quantity. 4chan posts are not Ulysses, but they are still useful data.

My point was not that they would refuse to scan Book By the Yard quality works because they would consider them unsuitable. My point was that they don't need to scan mass market Book by the Yard stuff, as they already have digital copies of it.

And scanning books is not as cheap as you think it is. Even using labor in low income countries, the cost of scanning is still vastly greater than just downloading a file.

But imagine the cost of demolishing one of those behemoths. Hard to see how that would be more economical for a solar farm rather than just leasing farmland.

[–] isleepinahammock@lemmy.blahaj.zone 3 points 1 month ago (1 children)

Not the rare, out of print, and long forgotten. And that's what's so scary about this. They're probably not paying to scan books that would otherwise end up at a landfill.

[–] isleepinahammock@lemmy.blahaj.zone 1 points 1 month ago (1 children)

I think what they're saying is that the AI companies don't need to rinse and repeat. They've already created solid programs for how to make LLMS speak English and demonstrate basic reasoning. You don't need to retrain that part every time. Once you've trained "English.exe," you can just copy it endlessly.

Maybe with enough time, the hard-coded English of the LLMs could become increasingly anachronistic and sound old-fashioned and formal to most ears. But for that kind of for slow maintenance you could just pay people to write examples of modern language and train it on that.

[–] isleepinahammock@lemmy.blahaj.zone 4 points 1 month ago (1 children)

I envision the opposite:

In the future, I imagine authors will go to great lengths to prove their works to be authentically human. It may become standard practice for authors to live-stream their writing. Literally a years-long stream of them sitting at their desk. Comments turned off.

Perhaps libraries could even offer this as a service. Offer a space where authors can come write in public and have it simultaneously live-streamed. Maybe it becomes the norm for authors to mostly write publicly on a platform like Twitch, but to occasionally write in public spaces where they can be physically observed.

This will not be a legal requirement to publish a book. But no respectable author would do otherwise.

My light bulb is bigger than yours!

[–] isleepinahammock@lemmy.blahaj.zone 4 points 1 month ago (3 children)

Why would they scan that stuff though? That kind of mass market (human made) slop is easily available in digital form. We know the AI companies have engaged in mass digital piracy, including running massive torrenting operations. So if a digital copy exists, they probably already have it. And even if they have to buy it, purchasing an ebook is a lot cheaper than buying a physical one, shipping it, paying someone to scan it, etc.

I would think the old, the out-of-print, the rare, and never-before digitized are the only things worth buying and physically scanning in 2026. Everything else has already been scanned or was born digital-native.

[–] isleepinahammock@lemmy.blahaj.zone 6 points 1 month ago (3 children)

Are you sure they're digitizing Book By the Yard quality material?

Consider this. Purchasing, shipping, and physically scanning a book is the single most expensive way for an AI company to acquire the text of a work. It's been widely documented in court cases that these companies engage in mass IP theft. They're literally running massive torrenting farms, grabbing copies of every film, song, book, etc. that they can get their hands on.

What kinds of texts are most likely to be found pirated on the internet? It's the mass market stuff. The common stuff was digitized long ago. They can just download that. They can pirate it. They can buy the ebook. There's no need for OpenAI to purchase and destructively scan the works of Steven King. No shade on the man, but his works aren't exactly hard to find. I'm sure they can just find a torrent.

There's little value in scanning the mass market books that are produced in enormous quantities. They probably don't buy "Book by the Yard" books, because they already have digital copies of those.

But the rare, long out-of-print stuff? The stuff that you actually cannot find a legal or illegal copy of online? That's only stuff worth paying money to buy, ship, and scan.

Doing anything physical, especially at scale, is slow and expensive. I would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.

Or what Robert Picardo does with all his Star Trek: Voyager royalties.

Bruh, this is literally like scientists scavenging metal from shipwrecks that predate the era of nuclear bomb testing, cause they need steel that isn’t contaminated by nuclear fallout for certain measuring equipment

Except in this case, these scientists are working at a nuclear weapons lab. They need metal that hasn't been contaminated by nuclear bombs so that they can invent better nuclear bombs.

view more: ‹ prev next ›