Research questionHow can LLM agents generalize to unseen tasks without directly fine-tuning their policies?LLM agents can perform well on familiar tasks yet fail when they encounter unfamiliar task types. Improving this transfer is difficult because changes that help known tasks may not produce robust behavior on new ones.