Data contamination expert 👌

ElCanut@jlai.lu · 7 months ago

Data contamination expert 👌

Eager Eagle@lemmy.world · 7 months ago

Good move, but anyone using public data already applies a simple spam filter to reject “dumb” data poisoning. Also, hatred and other negative comments as responses will be penalized in a language model training, so an effective data poisoning takes effort. I’ll just throw some ideas here how poisoning could hypothetically have a tangible negative impact in their results.

The best one can do in terms of data poisoning is make comments that are not easily discernible from usual comments - both for humans and machines - but are either unhelpful or misleading. This is an “in-distribution” data poisoning attack. To be really effective in having any impact whatsoever for training, they need to be mass applied using different user accounts that also upvote each others’ comments in a way that mimics real user interaction: if applied in a simplistic way, a simple graph analysis on these interactions can highlight these fake accounts as a christmas tree.

greenskye@lemm.ee · edit-2 7 months ago

but are either unhelpful or misleading

Honestly that just sounds like a lot of Reddit users in general

Darth_Mew@lemmy.world · 7 months ago

yea we know that’s why he said that because that’s “real” reddit content

Adalast@lemmy.world · 7 months ago

I was contemplating the merits of botting with the current model with slight vectorization offsets so the data becomes prone to overfitting.

I would think it would alao work to post using valid, but non-standard syntax so it muddies the n-gram searches.

Flumpkin@slrpnk.net · 7 months ago

I’m pissed at reddit but I still hate searching for something and finding a post on reddit discussing it, only to find some of the posts being deleted or overwritten.

mods_are_assholes@lemmy.world · 7 months ago

Good, then the protest at least worked somewhat.

FIST_FILLET@lemmy.ml · 7 months ago

if you’re lucky, some posts have been archived on the internet archive’s wayback machine. highly recommend pinning the extension to your toolbar, it’ll show a number badge of how many times the current site has been archived :) https://addons.mozilla.org/en-US/firefox/addon/wayback-machine_new

Poem_for_your_sprog@lemmy.world · 7 months ago

Set up a bot that just constantly posts blatantly wrong information, like “the earth is flat according to encyclopedia Britannica”, “the sky is green because it’s full or chlorophyll according to the UK foundation of science”

Zink@programming.dev · 7 months ago

Or in line with current events, “we are sorry about your experience and will refund you triple.”

jkrtn@lemmy.ml · 7 months ago

You won’t poison the data if the bot is on there just doing the same things as the redditors.

Vilian@lemmy.ca · 7 months ago

we need to make a repository just for that and spam reddit with it, everyone is welcome to contribute, open-source fake news

EmperorHenry@discuss.tchncs.de · 7 months ago

after they announced it would’ve been the time to start poisoning the comments. Then it would’ve been completely justified and moral.

Honestly, keep up the good fight. Start poisoning all open sources that are being scraped by any type of AI.

And I use the term “ai” very, very loosely. Because what’s called ai now isn’t real ai. It’s just an automated data collection tool.

It doesn’t create anything, it plagiarizes real artists.

FIST_FILLET@lemmy.ml · 7 months ago

exactly, ”ai” right now is just a computer parrot. why settle for blurry generic versions of the art that it is digesting and shitting back out?

mods_are_assholes@lemmy.world · edit-2 7 months ago

Nailed it. The whole essence of AI is that it can make images with a variety of colors and styles, but it’s not creative or artistic by definition. At the end of the day, it’s just a bunch of numbers and equations being translated into pixels on a screen.

(This comment pasted from NovelAI with this prompt:

Please write a reply to this interrnet comment: exactly, ”ai” right now is just a computer parrot. why settle for blurry generic versions of the art that it is digesting and shitting back out?)

Fades@lemmy.world · edit-2 7 months ago

That is not the “whole essence” of it all… You are summarizing he whole piece of tech off a single use-case (image generation).

AI is MUCH more than just a picture or generator. As a software engineer I use AI for things like debugging or quickly automating some tasks

benignintervention@lemmy.world · 7 months ago

I wonder how much these models are now learning from spam they were used to generate

Ilovethebomb@lemm.ee · 7 months ago

Thing is, you could use a bot to do nothing but post pop culture references, and it would be indistinguishable from a garden variety Redditor. Reddit is one of the worst places to train an AI.

LordOfTheChia@lemmy.world · edit-2 7 months ago

Johnson! Why the hell is your report the most unintelligible thing I’ve read since nineteen ninety eight when the undertaker threw mankind off hеll in a cell, and plummeted sixteen feet through an announcer’s table.

SatansMaggotyCumFart@lemmy.world · 7 months ago

Haha poop knife jolly rancher broken arms blowfly girl.

theneverfox@pawb.social · 7 months ago

I don’t recognize one of those… I’m not going to say which one, because I’m sure I don’t need to know

Kbin_space_program@kbin.social · 7 months ago

Time to make a lot of wandering dwarf bots on reddit to make variations of various game phrases all over, so the LLM based bots just spout Rock And Stone and This is my favourite store on the Citadel?

THE MASTERMIND@feddit.ch · 7 months ago

All of them

TropicalDingdong@lemmy.world · 7 months ago

I used some tools to corrupt about 10 years of comments and posts of mine.

VaultBoyNewVegas@lemmy.world · 7 months ago

I edited mine via a tool to say fuck Reddit and Steve Huffman is a greedy pig boy.

mp04610@lemm.ee · 7 months ago

While that’s the correct thing to do in my opinion, it would be a mistake to assume that Reddit didn’t store your original comments.

By corrupting their dataset, you may actually be helping them recognize maliciously edited comments.

khannie@lemmy.world · 7 months ago

it would be a mistake to assume that Reddit didn’t store your original comments.

They were fairly specific about not doing that (I’d imagine largely because of GDPR).

I deleted 10 years of “content” before I left and checked their policies. They apparently actually do properly delete from their servers.

It's A Faaaahhkeah!@lemmus.org · 7 months ago

But the GDPR only covers European users tho.

khannie@lemmy.world · 7 months ago

That’s true but it’s far easier to globally implement rather than trying to segment. Very difficult to accurately prove a user isn’t EU resident across an entire userbase.

It's A Faaaahhkeah!@lemmus.org · 7 months ago

That’s probably why they don’t let you access Reddit with a VPN, so they can have some idea of location.

Frozengyro@lemmy.world · 7 months ago

I’ve got a bridge in the desert I’d like to sell you.

joenforcer@midwest.social · edit-2 7 months ago

GDPR is no joke. Storing a handful of comments is not worth the penalty if they get caught.

Note that I speak from experience as part of a company that needs to comply with the regulations. We do it because the risk of violation is 10000000% not worth it no matter how annoying and arduous it is to comply.

flambonkscious@sh.itjust.works · 7 months ago

Mass edits made rapidly are obviously suspect, too… If the same user edits anything more than a dozen comments in, say a minute, you have to ask what’s going on

TropicalDingdong@lemmy.world · 7 months ago

Yeah, I mean I knew that when I was doing it.

Sometimes all you can do is make a symbolic gesture that really does nothing, and even if it does nothing, you should still do it.

Probably leaving and supporting lemmy by paying for some developer fees (i’m on the patreon), posting and commenting, probably 100x more damaging to Reddit.

FeelThePower@lemmy.dbzer0.com · 7 months ago

FWIW, I requested an old reddit accounts data the other day under CCPA and all the contamination was in there. My guess is their backend updates every so often. i guess i made a good call to edit my comments and leave them there to simmer before i deleted them along with the account. perhaps this is the way?

Octopus1348@lemy.lol · 7 months ago

What do you mean by corrupt?

Jo Miran@lemmy.ml · 7 months ago

It replaces them with jibbersish. I did the same for my 12+ years worth.

PlasmaDistortion@lemm.ee · 7 months ago

I used a tool that edited my comments to replace it with gibberish. Supposedly Reddit still retains deleted comments but if you edit them, it only keeps the latest version. So by editing it you make the comments worthless.

Octopus1348@lemy.lol · 7 months ago

I also edited my comments to be basically a Lemmy ad and completely deleted the posts except in a few communities where it could be helpful in the future.

citrusface@lemmy.world · 7 months ago

What tool? I’d like to use it as well.

bobs_monkey@lemm.ee · 7 months ago

I used redact.dev

citrusface@lemmy.world · 7 months ago

Thank you

TropicalDingdong@lemmy.world · 7 months ago

I ran a script over all of my comments (through my browser) to edit them into something about how spez had back stabbed the community. I had tens? hundreds of thousands? of comments.

It took several hours to run, but I did a forward pass (newest to oldest) and a backwards pass (oldest to newest). It bugged out because it had to run so long but I think I got it all.

I’m not sure this will really do anything because you could pretty easily statistically isolate any one who did what I did, and roll their account history back to a prior state in the training data.

Regardless, it was the least I could do on the way out the door.

teamevil@lemmy.world · 7 months ago

I simply got permabanned and my account disappeared.

ElCanut@jlai.lu · 7 months ago

Can’t post a genius idea like this one without posting the links of the tools

TropicalDingdong@lemmy.world · 7 months ago

Its not my idea, but I could probably dig up the tool I used. Dollars to donuts, it doesn’t work any more.

This might have been the tool I used. I dont think so because I overwrote everything with one message, but google around you’ll find similar.

https://github.com/adriantache/YARCO

Maalus@lemmy.world · 7 months ago

If you overwrote with a single message, then your messages are back to what they were.

KnightontheSun@lemmy.world · 7 months ago

Not necessarily true. I overwrote several thousand comments with a different tool and used three different quotes on greed. I have periodically checked and about two dozen came back. I just manually changed them at that point.

RecallMadness@lemmy.nz · edit-2 7 months ago

This would be better if it fed the parent comment into ChatGPT prefixed with “create a plausible but factually incorrect aggressive response to <comment>”

Feed the machine to the machine!

TropicalDingdong@lemmy.world · 7 months ago

be the change you want to see

Sabin10@lemmy.world · 7 months ago

A tool like that would almost definitely require api access to function. If that was still possible, most of us wouldn’t be here having this conversation.

dependencyinjection@discuss.tchncs.de · 7 months ago

The tool I used had an extension for Firefox. You then used that Reddit extension so you could get more scrolling on your post history. Then you pressed a button and it would insert gibberish for all comments and posts. Then you’d go next page and do it again.

TropicalDingdong@lemmy.world · 7 months ago

A tool like that would almost definitely require api access to function. If that was still possible, most of us wouldn’t be here having this conversation.

No it didn’t use the API. You had to run it in browser and be logged in to reddit.

prettybunnys@sh.itjust.works · 7 months ago

Most of them just do webpage stuff via browser extensions.

They just automate it

Ragnarok314159@sopuli.xyz · 7 months ago

I think Reddit caught on to this. I tried destroying my comment history (~7 years with 600k karma) with a few of the available tool on GitHub.

Found my account permabanned next time trying to login. People should attempt to eliminate/poison as much as possible, but Reddit has all the comments and modifications in a database somewhere to sell it all to whatever AI is the highest bidder.

They have to do something to make money after taking away awards. The advertising is absolute shit and not worth the $100 entry fee.

Ensign_Crab@lemmy.world · 7 months ago

“So much for your fucking canoe!”

ILikeBoobies@lemmy.ca · 7 months ago

Why care?

It just seems they were correct in changing api prices

Sylvartas@lemmy.world · 7 months ago

I mean, yeah, but because they fully expected to sell their userbase as training data for LLMs, not because they actually care about people using bots to post wrong informations. Wouldn’t that require them caring about actual people posting wrong infos in the first place ?

Norgur@kbin.social · 7 months ago

we need a bot that deletes comments and replaces them with some faulty grammar yoda-speak.

nickwitha_k (he/him)@lemmy.sdf.org · 7 months ago

I really ought to have done that.

byroon@lemmy.world · 7 months ago

So you’ve contaminated the training data for an LLM by spamming a public forum? Seems like everyone loses

wildginger@lemmy.myserv.one · 7 months ago

I dont lose, I get a good laugh out of watching idiots feed unreliable data to their LLMs because it was cheap

byroon@lemmy.world · 7 months ago

I mean the people using the forum who have to navigate around your spam

wildginger@lemmy.myserv.one · 7 months ago

Theyre on reddit, the spam site. I think theyre okay with a little more spam on their spam.

magnetosphere@kbin.social · 7 months ago

This is the ideal meme format. Pedro’s smile is perfect.

LunaCtld@lemmy.world · 7 months ago

So does that mean, that this time DAN will come pre-installed?

Valmond@lemmy.mindoki.com · 7 months ago

They’ll just find the signal in what you’re doing. Sorry but checkmate, mate.

Flying Squid@lemmy.world · 7 months ago