Research questionHow can language models distinguish trusted instructions from untrusted text to resist prompt injection?Text alone can make user input, tool output, and instructions appear interchangeable. As a result, malicious instructions embedded in untrusted context can redirect a model’s behavior.