@andreaslindholm Best I can tell in gpt3 was trained on the whole internet + books1, books2 which appear to include smashwords which is a bunch of amateur writing.
So we asked a bot to read a million pages of unedited, slush and create docs just like that and lo!, it did what we asked.
Openai AFAIK is keeping secret what it trained 4 on. I expect to see opensource models trained "just on public domain" or "just on good literature" and that is likely to be less cheesy.