Subject
3 entries
Prompt Injection
Bookmarks
ai-llm-agent-solver: Autonomous Gandalf Challenge Solver
An LLM-powered agent that autonomously solves the Gandalf AI challenge — a prompt injection security game where you try to extract a secret password from a guarded AI. Uses OpenAI API and agent-based reasoning.
Data Exfiltration from Writer.com via Indirect Prompt Injection
PromptArmor and Kai Greshake demonstrate data exfiltration from Writer.com via indirect prompt injection — where malicious content in a document the AI assistant processes causes it to leak user data. A concrete case study of the class of attack affecting all document-processing LLM applications.
Prompt Injection Attacks Against GPT-3
Simon Willison's September 2022 post naming and describing prompt injection attacks against GPT-3 — one of the first clear articulations of the attack class where malicious content in the environment overrides the developer's system prompt. The post that put the term 'prompt injection' into common use.
