Research questionHow reliable are single-pass benchmarks at detecting factual errors in long-form medical chatbot responses?Factual errors can be subtle and embedded in responses that are mostly correct, making them easy for a single annotator to miss. Benchmark results can also vary with the expertise and evidence used to judge whether an error is present.