Research questionHow can we assess whether terminal-use agents reliably handle routine, scientific, and engineering workflows?Existing computer-use evaluations often focus on graphical interfaces, while terminal benchmarks tend to emphasize programming and other technical workflows. As a result, agent reliability across routine digital activities and specialized scientific or engineering work remains poorly characterized.