Get Started
Research questionHow can evaluations separate model capability from execution-harness capability?An agent’s task results depend on its execution infrastructure as well as its model weights, so a single harness can obscure where capability comes from. It remains unclear whether agents can create or revise harnesses whose gains persist on held-out tasks and across executing models.
AI
AI Agents
Evaluation & Benchmarks
Latest papersRecent research connected to this question, newest first.HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?The evidence covers harness creation from a minimal seed and a few cases, followed by harness evolution using downstream execution feedback. Harnesses are assessed on held-out task success and execution-token cost across writing, code, search and research, and machine-learning experimentation; results show substantial domain variation, unstable evolution gains, partial transfer, and strong dependence on the model executing the harness.research paper · Sep 1, 2026
Related questions
How can agent harnesses adapt across tasks and models without manual redesign?How can evaluators distinguish missing knowledge from miscalibrated outputs in language models?How can agents adapt as their tool, skill, and specialist-agent harness evolves without losing existing capabilities?How can agentic benchmarks be compared and reused across complex environments and bespoke agent integrations?
Home
Topics
Search
Library