How it currently exists, yes in most cases it is trained on stolen cognitive labor. Do you think this is inherent to the technology itself, however? Consider a model trained on entirely public domain data, or non-copyleft liscence not requiring attribution. E.g., talkie
Totally agree that we need strict regulation.
If only we lived in a society where people could be freely able to produce cognitive labor while also being guaranteed a dignified life with universal basic services and income, regardless of what they produce. Then, like with piracy, LLM training, in my opinion, could be trained on anything without harming original authors.