Get Started
Home
Topics
Search
Library
Research questionHow should security evaluations measure indirect prompt-injection risk when attackers adapt their search and test-time compute?Indirect prompt injection exposes task- and environment-dependent attack surfaces, so fixed attack-success results can miss vulnerabilities discovered through adaptive searching. The attacker’s compute budget may therefore change the apparent security of the victim agent.
AI
AI Agents
Alignment & Safety
Evaluation & Benchmarks
Latest papersRecent research connected to this question, newest first.Rethinking Indirect Prompt Injection as a Test-Time Search ProblemThe evidence covers an agentic attacker that performs environment reconnaissance, reasons over attack strategies, and adaptively evaluates attempts using victim-agent feedback across heterogeneous tasks. It reports improved vulnerability discovery and exploitation with more attacker test-time compute, while explicit strategy management reduces redundant search at larger budgets. The findings are limited to the evaluated tasks and systems.research paper · Sep 3, 2026
Related questions
How should ML vulnerability-detection benchmarks measure practical security capabilities beyond narrow binary function-level tasks?How can language models distinguish trusted instructions from untrusted text to resist prompt injection?How can tool-using agents prevent sensitive conclusions assembled from individually non-revealing tool outputs?How should LLM safety be evaluated when harmful prompts vary in implicitness and sophistication?