Recently I set up https://llmsecurity.net to track all the huge mass of research on LLM security. There's a definition there too, if you're in NLP, security, or just curious and don't know much about this area. It's only about a year old, full of serious folks, and a lot of fun!
Remote
Leon Derczynski 🏔️✍🏻🌲☀️
@leon@dair-community.social
scientist: machine learning, language🔹prof with affiliations and an h-index🔹i research online harms - characterisation, identification, and mitigation🔹british, based in copenhagen, living in seattle🔹i 💚 equity and efficiency
0 Followers
0 Following
5 Posts
Joined November 13, 2022
Open post
People talk about "uncensored" LLMs, but are they actually censored in the first place?
Turns out you can use a *very* simple model to guide other LLMs into generating toxicity, and you can do this automatically
https://interhumanagreement.substack.com/p/faketoxicityprompts-automatic-red
2
0
1
0
Open post
Replying to
@transponderings@eldritch.cafe @JesseSkinner@toot.cafe @Bdellar@mastodon.social not convinced birdsite was the origin of/gets to own this behaviour, it used to be found in fax machines often enough!
2
1
1
0
Open post
@aletsi@mastodon.uno they are a Cutie!
0
0
0
0