Re: Why not try Finnegans Wake
Right, well I’m afraid that didn’t take long to figure out what’s going on, and I absolutely stand by my original statement, although with fractional nuance.
ChatGPT also knows the first line of Finnegans Wake, verbatim.
But if you ask it the second line, it claims that it can’t, because copyright issues. But it can *paraphrase*. And what it paraphrases with…..is wrong. Completely wrong, wrong subject and no words in common. The suggested second line does indeed reflect themes and plot of Finnegans Wake, but ChatGPT does not know what the second line of Finnegans Wake actually is, even in outline. Or the third line. Or the last line, it absolutely does not have a clue what the text is.
Famously, the last line is cyclical with the first line. ChatGPT knows this, can explain the link, can do a thematic analysis of the final monologue. But it does not actually know any of the words in it. ChatGPT knows famous facts about Wake, but it simply does not know the text, apart from the famous first line.
In short, ChatGPT knows the Wikipedia article on Finnegans Wake. It may also have memorised a statistical average of blogs and essays about Wake, that I haven’t yet figured out, and am continuing to play with it. Now, you may have an opinion that storing (a compressed form of) Wikipedia without attributions is itself copyright violation. I haven’t thought that through yet. But one thing I am absolutely certain of, is that ChatGPT has not memorised any significant portion at all of the actual text of Finnegans Wake. Even though it is public domain and one of the most famous books in the English language (which nobody has read). Far less has ChatGPT memorised the text any of the much larger corpus of English literature.
And I’ve even got more direct confirmation of that: if you ask it the most common word in Wake, it says “the”. It can even give you the top three most common words. But if you ask it for the tenth most common word *it does not know*. It is not counting words in the text. And it says “I can’t find a reliable source for that information”. It literally does a web search in front of your eyes, posts a publically available link to a word-frequency analysis essay, and then tries to parse the document in that link. This is really clear.
You called out a different LLM. I can’t and won’t do a systematic study to see which of the LLMs might have actually memorised more. Maybe you can find another one. But to be honest, this is like perpetual motion machines. There are really good fundamental information-theoretic reasons to believe that a 7B parameter model is not storing the compressed text of a 70T corpus as a small part of its storage. I’ve debunked one perpetual motion machine, I’m not going to debunk all of them separately.
This is a thing that just is not true. It cannot be true. And it is not.