Research questionHow can medical AI agents be evaluated for fabricated evidence and incoherent reasoning beyond final-answer correctness?A medically correct answer can still rely on fabricated evidence or an incoherent reasoning path. Final-answer accuracy therefore may not reveal clinically dangerous failures in an agent’s intermediate reasoning.