Get Started
Home
Topics
Search
Library
Topic · 44 recaps
Evaluation & Benchmarks
How we measure model capability — designing benchmarks, spotting contamination, judging open-ended outputs, and stress-testing claims about progress.
...
Sort
Newest
Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
Evaluation · Aug 28
0
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
Agents · Aug 28
0
Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
Agents · Aug 28
0
J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data
Evaluation · Aug 27
0
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
Agents · Aug 20
0
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
Agents · Aug 19
0
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Agents · Aug 13
0
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
Agents · Aug 10
0
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
Agents · Aug 7
0
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
Evaluation · Aug 6
0