What does it for me is that these companies were unwilling to pay creators to use their work in the training data
I believe it was logistical impossible. Many books used will probably be scanned and not even be available as ebook or drm protected or out of print. And e.g. Anna's archive has 64,416,225 books and 95,689,473 papers. Too large to even say what is pirated or nor, or buy every book in a lifetime. And if every book costs you ~$10 that's close to a billion dollars upfront. Basically creating LLMs wouldn't have been possible without piracy (or maybe the datasets aren't actually that extensive).
It's hypocritical, but ultimately the same argument for piracy that individuals use: IP laws creates unreasonable restrictions that prevent people from learning (or enjoy culture at a sensible price).
(I assume you're not saying you would need a negotiate a specific license to use a book or a public comment or article for machine learning).
Kinda reminds me of Year Zero by Robert Reid. The whole galaxy full of aliens loves Earth Music but only recently figure the concept of copyright. And how much quintillion moneys they now owe Earth lol.
Hardware compatibility is the major problem with alternative android. Basically I have to buy a 10 year old used phone as is. Would this only make this worse?
Anyway, maybe AI code generation could actually help with things like hardware compatibility. Things like an AI agent could iterate on reading about a smartphone model and testing various configurations and patch drivers to make it compatible for many more devices. Not sure if AI is suitable for that, but you should be able to define stringent test cases for this.