Re: DMCA?
"because of course the only way OpenAI could find a copy of any of their works in digital form is if OpenAI had themselves broken into the Kindle app store or something like that."
When I read this, I thought you might be sarcastic. I can't make that make sense with the rest of your comment, though, so I fear you're being serious. If you were, you're quite wrong. There are plenty of copies of books on the internet without permission, and they're not that hard to find by accident. Publishers go after particularly popular ones, but there are many that are on sites with small readership that publishers either don't know about or don't want to spend hours to fight when they can pop up again in ten minutes if they want to. Those sites are breaking the law, but just because the publishers haven't stopped them doesn't mean that a crawler can read the book from there and do whatever it wants with it. The LMMs have scraped a lot of the internet, and I'm sure they've included plenty of information that wasn't supposed to be on the pages it was on. Having stumbled on copyrighted works when I wasn't even trying to, I know it's not that hard.