Get Started
Research questionHow should scientific agents be evaluated on underspecified, attachment-rich requests without ground truth?Conventional scientific-AI benchmarks use fixed questions, reference solutions, or simulators with known structure. Real requests may be underspecified, include attachments, and lack objective answers, making the quality of delivered work and its claims difficult to assess.
AI
AI Agents
Evaluation & Benchmarks
Latest papersRecent research connected to this question, newest first.K-Bench: measuring model performance on real scientific agent requestsThe evidence covers 1,602 end-to-end runs on first-turn requests sampled from live K-Dense Web traffic, using nine frontier models in identical sandboxes. Three blinded language-model judges applied an eight-dimension rubric; judge rankings differed, and no model cleared the stated acceptance threshold under all three judges. The findings do not establish a universally calibrated absolute score or a definitive top system.research paper · Sep 2, 2026
Related questions
How can scientific agents choose domain-specific procedures that make analyses defensible?How do we evaluate whether scientific agents make justified discoveries from data rather than reproduce known analyses?How can research agents refine multi-constraint answers while keeping evidence verified over long horizons?How can medical AI agents be evaluated for fabricated evidence and incoherent reasoning beyond final-answer correctness?
Home
Topics
Search
Library