Subject
2 entries
Jailbreak
Bookmarks
jailbreak_llms: CCS'24 Jailbreak Prompt Dataset
A dataset of 15,140 ChatGPT prompts including 1,405 jailbreak prompts collected from Reddit, Discord, and open-source datasets — published at CCS 2024. The most comprehensive public collection of real-world jailbreak attempts against LLMs.
AI Adversarial Attacks: Automated Jailbreaks via Text Suffixes
Ars Technica covers the Universal Adversarial Attacks paper from CMU/Center for AI Safety — automated adversarial suffixes appended to prompts reliably bypass safety training on GPT-4, Claude, and open-source models. The attacks are transferable and potentially unstoppable with current alignment techniques.
