Get Started
Research questionHow can agentic benchmarks be compared and reused across complex environments and bespoke agent integrations?Agentic benchmarks depend on complex environments and bespoke integrations, making them difficult to run consistently across agents. This limits reliable comparison and broad reuse of benchmark results.
AI Agents
Evaluation & Benchmarks
Latest papersRecent research connected to this question, newest first.Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic EvaluationThe source covers adapters for more than 80 benchmarks, validated through code review and parity experiments, plus evaluations of 8 models across 54 benchmarks using Terminus-2 and three native harnesses. It also presents Harbor-Index, an 82-task subset drawn from 29 benchmarks and refined through difficulty filtering and human and AI audits; the evidence is limited to these evaluated benchmark, model, and harness configurations.research paper · Sep 10, 2026
Related questions
How should AI agents be benchmarked for environmental geospatial workflows using structured calls to realistic APIs?How can we reduce per-task LLM-agent evaluation cost without distorting benchmark outcomes?How can agent harnesses adapt across tasks and models without manual redesign?How should enterprise decision agents be evaluated when rankings change between fixed-opponent and shared-market settings?
Home
Topics
Search
Library