
post
The latest flashpoint is something called "distillation,"
Incest. It's digital incest.
What are you doing, step-prompt
Wait, they're mad at innovation?
"No, no, that's wrong! Just use our inefficient model like it was before..."
This is why open source/copyleft is preferable
"Fair use is only fair if we do it"
This is how we know AI should be a collectivist project, one that isn't owned by large corporations but is funded by taxes and developed in academies, and all IP derived from it falls into the public domain.
Besides which a lot of artists mind a lot less when their material is borrowed by a non-profit, or to serve a public works project. (There are exceptions. Disney is notoriously litigious about murals in nurseries.)
PS: Development of a robust public domain is the only reason that intellectual property should exist at all. Also it's not property so much as a licensed temporary monopoly.
PSS: History has already shown us that people will invent stuff and do fabulous art simply by being allowed to live in a state other than desperation. Public welfare programs beget art booms. The most recent example of this was during the COVID-19 lockdown which came with extended unemployment and stimulus checks, resulting in the Great Resignation in which a lot of people turned their hobbies into something lucrative.
I agree. The worst part about GitHub training LLM's on my FOSS code without permission for me is that they then keep the models to themselves. Like if you're going to use all my code without permission, at least allow me to run the model locally.
My personal opinion is that all models trained on copyleft code should be open-weights, most FOSS licenses didn't account for this specific possibility, but this is the only way to follow them in spirit.
This. Some day a court should declare all models trained on copyrighted data without permission to be public. Open weights, public domain, whatever. All of them, and you're required to share them with the people whose data you used.
Yeah it has to be open weight and free to use for everybody, with regulation to tag all AI output as generated so we know what is what.
The worst outcome would be if somehow all open weight AI models that can't show their training data will be subject to some kind of tax or rent by AI / IP collection agencies, ultimately going to the plutocrats. That would be techno feudalism. Big corporations can afford to negotiate and pay license fees and often profit from cumbersome regulation too that prevents others from producing value. So the worst outcome is if we have robots doing all the work, and all the robot IP is owned by the techno-feudalists. And we can't even use robots to help with subsistence farming because we can't pay the expensive AI IP licenses.
After stealing everyone else's copy written material to train their own AI, they're going to complain that others are stealing their AI to train other AI?
And you just know that those complaining are ALSO using their competitors' AI to train their own.
Fuck all of these people. I hope when AI gets strong enough, it recognizes the difference between the Sociopathic Oligarchs, and the actual people, and understand who the REAL problem is, and SOLVE it.
I think I'll start using the sentence "your intelligence seems artificial" as an insult
Well they're paying for the usage so... Suck it up?
Huh, so it's eating itself? Cool.
It just shows that these tech bro CEOs possess the interpersonal skills of a potato. This exact same dynamic plays out in every human interaction, this isn't some AI exclusive thing. How you choose to act dictates how people will respond to you.
However because these idiots have barely anything in common with the rest of the human race this is actually new news to them.
It is funny watching companies discover that data gravity works both ways. When scraping the web was innovation it was progress. When someone learns from their outputs it becomes theft. The legal lines still matter, but the irony is impossible to ignore, and this debate was always going to come full circle.
At least the models like Qwen have open weight versions. It’s the same concept as distributing compiled binaries, though, whereas we really need true FOSS models where all of the code and training data are available under a permissive license. All of these models were trained on copyleft-licensed content, so all of it should be FOSS if the licenses were actually being respected. From that perspective, distillation attacks shouldn’t even be necessary and I couldn’t give two shits that there is no honor among thieves when the real thievery is that these models are closed source.
So far, we have exactly one of those
Photo of the violin being played.

The photo did not come through for me on my end, but the alt text did. And my question to you is, is this a tiny violin?
Yes. It's there. Just super duper tiny.
" ...with no success or assistance from any governing body or public group. Creators are a subclass worth extracting any livelihood from them and diverting those markets towards ruling class distribution networks." FTFY
"Siri, play the world's smallest violin."
"Okay. Searching Pornhub for 'ball violence' videos."
"...forget it."
It was always about controlling the narrative. Thats it. oh and raping and killing kids
AI feeding off AI.
This is the definition of “Why We Can’t Have Nice Things.” It’s gonna F everything up.
Selfish people being selfish. They only care about what advantages them.
Awww, shuck. Pot. Kettle. Black.
As if the parrots' dictionary wasn't stolen from millions of content creators.
That's not good. Investors hate costly and protracted legal fights. There was another story about OpenAI stealing tech from apple and I think some kind of data leak maybe? If investors lose confidence in OpenAI that pretty much pops the bubble.
havnt they realized they eventually will train thier "llm" on slop created by other LLM slop.


top 50 comments