Skip to main content
Ryan Orban

Ryan Orban

Subject
3 entries

Prompt Injection

Bookmarks

  1. ai-llm-agent-solver: Autonomous Gandalf Challenge Solver

    An LLM-powered agent that autonomously solves the Gandalf AI challenge — a prompt injection security game where you try to extract a secret password from a guarded AI. Uses OpenAI API and agent-based reasoning.

  2. Data Exfiltration from Writer.com via Indirect Prompt Injection

    PromptArmor and Kai Greshake demonstrate data exfiltration from Writer.com via indirect prompt injection — where malicious content in a document the AI assistant processes causes it to leak user data. A concrete case study of the class of attack affecting all document-processing LLM applications.

  3. Prompt Injection Attacks Against GPT-3

    Simon Willison's September 2022 post naming and describing prompt injection attacks against GPT-3 — one of the first clear articulations of the attack class where malicious content in the environment overrides the developer's system prompt. The post that put the term 'prompt injection' into common use.

All bookmarks