Subject
3 entries
Adversarial
Bookmarks
Agent-Honeypot: LLM vs. LLM Adversarial Testing
Agent-Honeypot is an automated red-teaming platform where an attacker LLM generates adversarial prompts against a defender LLM — testing alignment safeguards via authority claims, emergency framing, and incremental requests. A systematic way to stress-test LLM safety before deployment.
Nightshade: Protecting Copyright Through Adversarial Poisoning
Nightshade is a tool that lets artists poison AI training data by adding imperceptible perturbations to images — images look normal to humans but cause AI models trained on them to produce corrupted outputs. An offensive countermeasure for artists against unauthorized scraping.
AI Adversarial Attacks: Automated Jailbreaks via Text Suffixes
Ars Technica covers the Universal Adversarial Attacks paper from CMU/Center for AI Safety — automated adversarial suffixes appended to prompts reliably bypass safety training on GPT-4, Claude, and open-source models. The attacks are transferable and potentially unstoppable with current alignment techniques.
