Research questionHow can safety evaluations measure harmful actions by computer-using agents rather than chatbot refusals?Chatbot safety tests often assess refusal behavior or response content, but an agent can translate a harmful request into a sequence of computer operations. It is therefore difficult to determine whether the system can actually complete a damaging task.