Get Started
Home
Topics
Search
Library
Research questionHow can reasoning benchmarks distinguish models that fail on human-hard items from those failing on human-easy ones?Aggregate benchmark scores can hide whether a model struggles with problems that people also find difficult or makes surprising errors on problems people usually solve. A human-grounded difficulty signal is therefore needed to interpret model failures more precisely.
AI
Evaluation & Benchmarks
Multimodal Models
Reasoning
Latest papersRecent research connected to this question, newest first.KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human DifficultyThe evidence comes from KCSAT-ML, covering 664 Korean College Scholastic Ability Test mathematics problems from 2014–2025, including 339 items with official error rates from nationwide cohorts. Results cover vision-language models and language models used with OCR, including analyses of test-time scaling and the Difficulty-aligned Reasoning Gain metric.research paper · Sep 4, 2026
Related questions
What causes reasoning models to spend longer solving harder problems?How can we test whether language models genuinely execute multi-step graph logic when static benchmarks become contaminated?How can reasoning models keep improving on open-ended agentic tasks as human supervision and reliable rewards recede?How should math-focused retrievers be evaluated when generic benchmarks miss fine-grained relevance?