@Grizzlysgrowls from docs "You will likely want to run GPT4All models on GPU if you would like to utilize context windows larger than 750 tokens" - so max tokens, which I think also gets called context window. I've seen reports that on all models there is a max_tokens where the model goes from generating text faster than you can read to waiting 10 minutes.
Post
January 18, 2024