Research questionHow can evaluations separate model capability from execution-harness capability?An agent’s task results depend on its execution infrastructure as well as its model weights, so a single harness can obscure where capability comes from. It remains unclear whether agents can create or revise harnesses whose gains persist on held-out tasks and across executing models.