the learning is done by completely different software
Wait, so you do think there is a software that "does learning"?
Isn't that the whole ballgame? You're saying LLM companies have made software that learns math and coding and translation at a near-expert level, you just happen to be under the impression that this software isn't present in deployed LLMs?
Unfortunately, the training algorithm that turns stochastic noise into trained LLMs is dumb as bricks. Engineers have done some tricks, but it's basically gradient descent over the vector space of text tokens/phonemes. What allows the stochastic noise to become something that can make novel statements about wildlife (like giving a halfway coherent answer to a question that has never been asked in its dataset) is the patterns that form within the weights and biases. It's vaguely like slime mold exploring a maze; the "teacher" can just be a stupid lump of sugar, it's the "student" that does the emergently complex task of finding the best route by executing simple algorithms in every part of itself.
Effectively, this means the LLM is the software that wrote the LLM. It is the software that did the learning.
It genuinely doesn’t understand a word you said
What is the difference between understanding a thing (to a certain degree) and being able to generate an accurate narrative about the thing (to the same degree)? Can you give any evidence that I've understand any word you've said? How can you tell I haven't just been very good at guessing what word I'm supposed to say next based on the many arguments I've read or participated in?
Do you mean more by 'understanding' than expressed in the Chinese Room argument?