No, it's not arbitrary, the learning is done by completely different software at a completely different time on completely different computers.
I'm pointing out that the LLM is the product of the machine learning, where you feed all of Wikipedia and reddit and stack overflow in and calculate the LLM from it.
It's like if you wrote a book about the wildlife of antarctica. First you would learn about the wildlife and then you would write the book. The book isn't learning anything and it isn't intelligent. The book represents knowledge but it doesn't itself know or understand anything.
Similarly, the LLM isn't learning anything and it isn't intelligent, it's just regurgitating randomly selected words from its training data that look like they usually occur after the other words in the conversation so far.
It genuinely doesn't understand a word you said, using its guess about what sort of conversation it's supposed to be having and it's always just guessing what word it's supposed to say next.