It's the generation speed. Internally LLMs use tokens which represent either words or parts of words and map them to integer values. The model then does it's prediction on which integer is most likely to come after the input. How the words are split up is an implementation detail that can vary from model to model
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments