openmic.social is an uncensored community. You may encounter strong language, controversial opinions, and mature or NSFW material. You must be 18+ to browse. Illegal content is prohibited and removed on sight — please report it. By continuing, you accept that you may see content you personally disagree with.
Na, for it to be effective it needs to be wide spread, but if its wide spread then it can be filtered out of the training material.
I’ve read in papers that you can poison datasets with a very small percentage of the data, if done cleverly. I can fish up the source if you want (but it might take me some time).
edit: here it is.
Emphasis mine. All it takes is 250 poisoned documents.