Millions of articles from The New York Times were used to train chatbots that now compete with it, the lawsuit said.

you are viewing a single comment's thread
view the rest of the comments
[–] 3 points 2 years ago (1 child)

These models can still be trained on data that they're allowed to use, but I think that what we're seeing is that the better LLM services are probably trained with shocking amounts of private data, whereas the less performant probably don't use stolen data.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 2 years ago* (last edited 2 years ago)

    Textbooks are a big one that I suspect we'll probably see a set of suits over. Particularly because they seem to be some of the most valuable training data.

  • source
  • parent