Research questionHow should AI agents be benchmarked for environmental geospatial workflows using structured calls to realistic APIs?Environmental scientists can spend substantial effort preparing and wrangling data, yet agent performance on realistic API-driven geospatial workflows remains poorly characterized. Generic GIS benchmarks may not show whether agents can select and execute the structured tool calls required in practical workflows.