▲ 81 ▼ Cable can't compete with 5G home internet, so it's cheating (www.spacebar.news) submitted 2 years ago by corbin@infosec.pub to c/technology@beehaw.org 60 comments fedilink hide all child comments
[–] onlinepersona@programming.dev 18 points 2 years ago (2 children) The US, wow... what a place to live in as the 99%. CC BY-NC-SA 4.0 permalink fedilink source hideshow 4 child comments replies: [–] acastcandream@beehaw.org 32 points 2 years ago (2 children) CC BY-NC-SA 4.0 Why are you putting a CC license on your comments? permalink fedilink source parent hideshow 4 child comments replies: [–] BotCheese@beehaw.org 21 points 2 years ago (5 children) From what I understand it is some thing for AI, to stop them from harvesting or to poison the data, by having it repeating therefore more likely to show up. permalink fedilink source parent hideshow 10 child comments replies: [–] beefcat@beehaw.org 59 points 2 years ago (1 child) Sounds an awful lot like that thing boomers used to do on Facebook where they would post a message on their wall rescinding Facebook's rights to the content they post there. I'm sure it's equally effective. permalink fedilink source parent hideshow 2 child comments replies: [–] Bene7rddso@feddit.de 4 points 2 years ago (1 child) Sure, the fun begins when it starts spitting out copyright notices permalink fedilink source parent hideshow 2 child comments replies: [–] t3rmit3@beehaw.org 2 points 2 years ago That would require a significant number of people to be doing it, to 'poison' the input pool, as it were. permalink fedilink source parent [–] corbin@infosec.pub [S] 41 points 2 years ago It seems pretty well established at this point that AI training models don't respect copyright. permalink fedilink source parent [–] mozz@mbin.grits.dev 20 points 2 years ago* (1 child) I would be extremely extremely surprised if the AI model did anything different with "this comment is protected by CC license so I don't have the legal right to it" as compared with its normal "this comment is copyright by its owner so I don't have the legal right to it hahaha sike snork snork snork I absorb" processing mode. permalink fedilink source parent hideshow 2 child comments replies: [–] Max_P@lemmy.max-p.me 13 points 2 years ago (1 child) No but if they forget to strip those before training the models, it's gonna start spitting out licenses everywhere, making it annoying for AI companies. It's so easily fixed with a simple regex though, it's not that useful. But poisoning the data is theoretically possible. permalink fedilink source parent hideshow 2 child comments replies: [–] t3rmit3@beehaw.org 1 point 2 years ago Only if enough people were doing this to constitute an algorithmically-reducible behavior. If you could get everyone who mentions a specific word or subject to put a CC license in their comment, then an ML model trained on those comments would likely output the license name when that subject was mentioned, but they don't just randomly insert strings they've seen, without context. permalink fedilink source parent [–] peter@feddit.uk 19 points 2 years ago That seems stupid permalink fedilink source parent [–] acastcandream@beehaw.org 12 points 2 years ago Interesting. Feels like that thing people used to add to FB comments back in the day that did nothing but in the case of AI I could see it maybe doing something. I’ll be looking into it - thanks! permalink fedilink source parent [–] conciselyverbose@kbin.social 19 points 2 years ago* (last edited 2 years ago) To turn every comment, no matter how on topic, into obnoxious spam. permalink fedilink source parent [–] Danterious@lemmy.dbzer0.com 17 points 2 years ago* (2 children) You know if you want to do something more effective than just putting copyright at the end of your comments you could try creating an adversarial suffix using this technique. It makes any LLM reading your comment begin its response with any specific output you specify (such as outing itself as a language model or calling itself a chicken). It gives you the code necessary to be able to create it. There are also other data poisoning techniques you could use just to make your data worthless to the AI but this is the one I thought would be the most funny if any LLMs were lurking on lemmy (I have already seen a few). permalink fedilink source parent hideshow 4 child comments replies: [–] dubyakay@lemmy.ca 5 points 2 years ago Thanks for the link. This was a good read. permalink fedilink source parent [–] onlinepersona@programming.dev 2 points 2 years ago That's a neat idea and I've considered it, but would need time to research and test. Time I don't have, so this is the easiest thing I came up with. If there were a bot, plugin, browser extension, or something that did the necessary modifications and kept up to date with new developments in AI, I'd use it. CC BY-NC-SA 4.0 permalink fedilink source parent
[–] acastcandream@beehaw.org 32 points 2 years ago (2 children) CC BY-NC-SA 4.0 Why are you putting a CC license on your comments? permalink fedilink source parent hideshow 4 child comments replies: [–] BotCheese@beehaw.org 21 points 2 years ago (5 children) From what I understand it is some thing for AI, to stop them from harvesting or to poison the data, by having it repeating therefore more likely to show up. permalink fedilink source parent hideshow 10 child comments replies: [–] beefcat@beehaw.org 59 points 2 years ago (1 child) Sounds an awful lot like that thing boomers used to do on Facebook where they would post a message on their wall rescinding Facebook's rights to the content they post there. I'm sure it's equally effective. permalink fedilink source parent hideshow 2 child comments replies: [–] Bene7rddso@feddit.de 4 points 2 years ago (1 child) Sure, the fun begins when it starts spitting out copyright notices permalink fedilink source parent hideshow 2 child comments replies: [–] t3rmit3@beehaw.org 2 points 2 years ago That would require a significant number of people to be doing it, to 'poison' the input pool, as it were. permalink fedilink source parent [–] corbin@infosec.pub [S] 41 points 2 years ago It seems pretty well established at this point that AI training models don't respect copyright. permalink fedilink source parent [–] mozz@mbin.grits.dev 20 points 2 years ago* (1 child) I would be extremely extremely surprised if the AI model did anything different with "this comment is protected by CC license so I don't have the legal right to it" as compared with its normal "this comment is copyright by its owner so I don't have the legal right to it hahaha sike snork snork snork I absorb" processing mode. permalink fedilink source parent hideshow 2 child comments replies: [–] Max_P@lemmy.max-p.me 13 points 2 years ago (1 child) No but if they forget to strip those before training the models, it's gonna start spitting out licenses everywhere, making it annoying for AI companies. It's so easily fixed with a simple regex though, it's not that useful. But poisoning the data is theoretically possible. permalink fedilink source parent hideshow 2 child comments replies: [–] t3rmit3@beehaw.org 1 point 2 years ago Only if enough people were doing this to constitute an algorithmically-reducible behavior. If you could get everyone who mentions a specific word or subject to put a CC license in their comment, then an ML model trained on those comments would likely output the license name when that subject was mentioned, but they don't just randomly insert strings they've seen, without context. permalink fedilink source parent [–] peter@feddit.uk 19 points 2 years ago That seems stupid permalink fedilink source parent [–] acastcandream@beehaw.org 12 points 2 years ago Interesting. Feels like that thing people used to add to FB comments back in the day that did nothing but in the case of AI I could see it maybe doing something. I’ll be looking into it - thanks! permalink fedilink source parent [–] conciselyverbose@kbin.social 19 points 2 years ago* (last edited 2 years ago) To turn every comment, no matter how on topic, into obnoxious spam. permalink fedilink source parent
[–] BotCheese@beehaw.org 21 points 2 years ago (5 children) From what I understand it is some thing for AI, to stop them from harvesting or to poison the data, by having it repeating therefore more likely to show up. permalink fedilink source parent hideshow 10 child comments replies: [–] beefcat@beehaw.org 59 points 2 years ago (1 child) Sounds an awful lot like that thing boomers used to do on Facebook where they would post a message on their wall rescinding Facebook's rights to the content they post there. I'm sure it's equally effective. permalink fedilink source parent hideshow 2 child comments replies: [–] Bene7rddso@feddit.de 4 points 2 years ago (1 child) Sure, the fun begins when it starts spitting out copyright notices permalink fedilink source parent hideshow 2 child comments replies: [–] t3rmit3@beehaw.org 2 points 2 years ago That would require a significant number of people to be doing it, to 'poison' the input pool, as it were. permalink fedilink source parent [–] corbin@infosec.pub [S] 41 points 2 years ago It seems pretty well established at this point that AI training models don't respect copyright. permalink fedilink source parent [–] mozz@mbin.grits.dev 20 points 2 years ago* (1 child) I would be extremely extremely surprised if the AI model did anything different with "this comment is protected by CC license so I don't have the legal right to it" as compared with its normal "this comment is copyright by its owner so I don't have the legal right to it hahaha sike snork snork snork I absorb" processing mode. permalink fedilink source parent hideshow 2 child comments replies: [–] Max_P@lemmy.max-p.me 13 points 2 years ago (1 child) No but if they forget to strip those before training the models, it's gonna start spitting out licenses everywhere, making it annoying for AI companies. It's so easily fixed with a simple regex though, it's not that useful. But poisoning the data is theoretically possible. permalink fedilink source parent hideshow 2 child comments replies: [–] t3rmit3@beehaw.org 1 point 2 years ago Only if enough people were doing this to constitute an algorithmically-reducible behavior. If you could get everyone who mentions a specific word or subject to put a CC license in their comment, then an ML model trained on those comments would likely output the license name when that subject was mentioned, but they don't just randomly insert strings they've seen, without context. permalink fedilink source parent [–] peter@feddit.uk 19 points 2 years ago That seems stupid permalink fedilink source parent [–] acastcandream@beehaw.org 12 points 2 years ago Interesting. Feels like that thing people used to add to FB comments back in the day that did nothing but in the case of AI I could see it maybe doing something. I’ll be looking into it - thanks! permalink fedilink source parent
[–] beefcat@beehaw.org 59 points 2 years ago (1 child) Sounds an awful lot like that thing boomers used to do on Facebook where they would post a message on their wall rescinding Facebook's rights to the content they post there. I'm sure it's equally effective. permalink fedilink source parent hideshow 2 child comments replies: [–] Bene7rddso@feddit.de 4 points 2 years ago (1 child) Sure, the fun begins when it starts spitting out copyright notices permalink fedilink source parent hideshow 2 child comments replies: [–] t3rmit3@beehaw.org 2 points 2 years ago That would require a significant number of people to be doing it, to 'poison' the input pool, as it were. permalink fedilink source parent
[–] Bene7rddso@feddit.de 4 points 2 years ago (1 child) Sure, the fun begins when it starts spitting out copyright notices permalink fedilink source parent hideshow 2 child comments replies: [–] t3rmit3@beehaw.org 2 points 2 years ago That would require a significant number of people to be doing it, to 'poison' the input pool, as it were. permalink fedilink source parent
[–] t3rmit3@beehaw.org 2 points 2 years ago That would require a significant number of people to be doing it, to 'poison' the input pool, as it were. permalink fedilink source parent
[–] corbin@infosec.pub [S] 41 points 2 years ago It seems pretty well established at this point that AI training models don't respect copyright. permalink fedilink source parent
[–] mozz@mbin.grits.dev 20 points 2 years ago* (1 child) I would be extremely extremely surprised if the AI model did anything different with "this comment is protected by CC license so I don't have the legal right to it" as compared with its normal "this comment is copyright by its owner so I don't have the legal right to it hahaha sike snork snork snork I absorb" processing mode. permalink fedilink source parent hideshow 2 child comments replies: [–] Max_P@lemmy.max-p.me 13 points 2 years ago (1 child) No but if they forget to strip those before training the models, it's gonna start spitting out licenses everywhere, making it annoying for AI companies. It's so easily fixed with a simple regex though, it's not that useful. But poisoning the data is theoretically possible. permalink fedilink source parent hideshow 2 child comments replies: [–] t3rmit3@beehaw.org 1 point 2 years ago Only if enough people were doing this to constitute an algorithmically-reducible behavior. If you could get everyone who mentions a specific word or subject to put a CC license in their comment, then an ML model trained on those comments would likely output the license name when that subject was mentioned, but they don't just randomly insert strings they've seen, without context. permalink fedilink source parent
[–] Max_P@lemmy.max-p.me 13 points 2 years ago (1 child) No but if they forget to strip those before training the models, it's gonna start spitting out licenses everywhere, making it annoying for AI companies. It's so easily fixed with a simple regex though, it's not that useful. But poisoning the data is theoretically possible. permalink fedilink source parent hideshow 2 child comments replies: [–] t3rmit3@beehaw.org 1 point 2 years ago Only if enough people were doing this to constitute an algorithmically-reducible behavior. If you could get everyone who mentions a specific word or subject to put a CC license in their comment, then an ML model trained on those comments would likely output the license name when that subject was mentioned, but they don't just randomly insert strings they've seen, without context. permalink fedilink source parent
[–] t3rmit3@beehaw.org 1 point 2 years ago Only if enough people were doing this to constitute an algorithmically-reducible behavior. If you could get everyone who mentions a specific word or subject to put a CC license in their comment, then an ML model trained on those comments would likely output the license name when that subject was mentioned, but they don't just randomly insert strings they've seen, without context. permalink fedilink source parent
[–] acastcandream@beehaw.org 12 points 2 years ago Interesting. Feels like that thing people used to add to FB comments back in the day that did nothing but in the case of AI I could see it maybe doing something. I’ll be looking into it - thanks! permalink fedilink source parent
[–] conciselyverbose@kbin.social 19 points 2 years ago* (last edited 2 years ago) To turn every comment, no matter how on topic, into obnoxious spam. permalink fedilink source parent
[–] Danterious@lemmy.dbzer0.com 17 points 2 years ago* (2 children) You know if you want to do something more effective than just putting copyright at the end of your comments you could try creating an adversarial suffix using this technique. It makes any LLM reading your comment begin its response with any specific output you specify (such as outing itself as a language model or calling itself a chicken). It gives you the code necessary to be able to create it. There are also other data poisoning techniques you could use just to make your data worthless to the AI but this is the one I thought would be the most funny if any LLMs were lurking on lemmy (I have already seen a few). permalink fedilink source parent hideshow 4 child comments replies: [–] dubyakay@lemmy.ca 5 points 2 years ago Thanks for the link. This was a good read. permalink fedilink source parent [–] onlinepersona@programming.dev 2 points 2 years ago That's a neat idea and I've considered it, but would need time to research and test. Time I don't have, so this is the easiest thing I came up with. If there were a bot, plugin, browser extension, or something that did the necessary modifications and kept up to date with new developments in AI, I'd use it. CC BY-NC-SA 4.0 permalink fedilink source parent
[–] dubyakay@lemmy.ca 5 points 2 years ago Thanks for the link. This was a good read. permalink fedilink source parent
[–] onlinepersona@programming.dev 2 points 2 years ago That's a neat idea and I've considered it, but would need time to research and test. Time I don't have, so this is the easiest thing I came up with. If there were a bot, plugin, browser extension, or something that did the necessary modifications and kept up to date with new developments in AI, I'd use it. CC BY-NC-SA 4.0 permalink fedilink source parent