cross-posted from: https://nom.mom/post/121481

OpenAI could be fined up to $150,000 for each piece of infringing content.https://arstechnica.com/tech-policy/2023/08/report-potential-nyt-lawsuit-could-force-openai-to-wipe-chatgpt-and-start-over/#comments

you are viewing a single comment's thread
view the rest of the comments
[–] 0 points 2 years ago (1 child)

It's not repeating its training data verbatim because it can't do that. It doesn't have the training data stored away inside itself. If it did the big news wouldn't be AI, it would be the insanely magical compression algorithm that's been discovered that allows many terabytes of data to be compressed down into just a few gigabytes.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 2 years ago* (last edited 2 years ago) (1 child)

    Do you remember quotes in english ascii /s

    Tokens are even denser than ascii. simmlar to word "chunking" My guess is it's like lossy video compression but for text, [Attacked] with [lazers] by [deatheaters] apon [margret];[has flowery language]; word [margret] [comes first] (Theoretical example has 7 "tokens")

    It may have actually impressioned a really good copy of that book as it's lilely read it lots of times.

  • source
  • parent
  • hideshow 2 child comments