Skip to main content
Ryan Orban

Ryan Orban

Subject
3 entries

Adversarial

Bookmarks

  1. Agent-Honeypot: LLM vs. LLM Adversarial Testing

    Agent-Honeypot is an automated red-teaming platform where an attacker LLM generates adversarial prompts against a defender LLM — testing alignment safeguards via authority claims, emergency framing, and incremental requests. A systematic way to stress-test LLM safety before deployment.

  2. Nightshade: Protecting Copyright Through Adversarial Poisoning

    Nightshade is a tool that lets artists poison AI training data by adding imperceptible perturbations to images — images look normal to humans but cause AI models trained on them to produce corrupted outputs. An offensive countermeasure for artists against unauthorized scraping.

  3. AI Adversarial Attacks: Automated Jailbreaks via Text Suffixes

    Ars Technica covers the Universal Adversarial Attacks paper from CMU/Center for AI Safety — automated adversarial suffixes appended to prompts reliably bypass safety training on GPT-4, Claude, and open-source models. The attacks are transferable and potentially unstoppable with current alignment techniques.

All bookmarks