Research questionHow can agentic benchmarks be compared and reused across complex environments and bespoke agent integrations?Agentic benchmarks depend on complex environments and bespoke integrations, making them difficult to run consistently across agents. This limits reliable comparison and broad reuse of benchmark results.