Research questionHow should scientific agents be evaluated on underspecified, attachment-rich requests without ground truth?Conventional scientific-AI benchmarks use fixed questions, reference solutions, or simulators with known structure. Real requests may be underspecified, include attachments, and lack objective answers, making the quality of delivered work and its claims difficult to assess.