Get Started
Research questionHow can we assess whether terminal-use agents reliably handle routine, scientific, and engineering workflows?Existing computer-use evaluations often focus on graphical interfaces, while terminal benchmarks tend to emphasize programming and other technical workflows. As a result, agent reliability across routine digital activities and specialized scientific or engineering work remains poorly characterized.
AI Agents
Evaluation & Benchmarks
Latest papersRecent research connected to this question, newest first.TUA-Bench: A Benchmark for General-Purpose Terminal-Use AgentsTUA-Bench provides evidence from 120 manually designed real-world tasks spanning five families, including document editing, email management, live-web information seeking, and scientific or engineering workflows requiring specialized software. Tasks run in deterministic terminal setups and use execution-based scoring; the reported strongest agent achieves 65.8% overall performance, with substantial variation across tracks.research paper · Jun 26, 2026
Related questions
How can computer-use agents efficiently coordinate GUI and CLI actions over shared application state?How can computer-use agents retain, refine, and reliably reuse procedural skills across repeated GUI tasks?How should scientific agents be evaluated on underspecified, attachment-rich requests without ground truth?How should AI agents be benchmarked for environmental geospatial workflows using structured calls to realistic APIs?
Home
Topics
Search
Library