Research questionHow can we reduce per-task LLM-agent evaluation cost without distorting benchmark outcomes?A single LLM-agent benchmark run can be expensive, and iterative development repeats that cost. Reducing the number of tasks does not reduce the execution cost of each task that remains.