Get Started
Topic · 81 recaps

Evaluation & Benchmarks

How we measure model capability — designing benchmarks, spotting contamination, judging open-ended outputs, and stress-testing claims about progress.
PostsQuestions
Home
Topics
Search
Library
Sort
Newest
$Φ$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Agents · Sep 9
A Hallucination Score Is Two Different Things
Evaluation · Sep 8 · 11:56
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
Agents · Sep 8
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
Agents · Sep 8
Three Ways "Our Agent's Memory Still Works" Can Be Wrong
Agents · Sep 7 · 12:45
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
Agents · Sep 7
Two Different Failures Look Identical in Your Audio Eval
Audio/Speech · Sep 6 · 14:03
Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning
Evaluation · Sep 6
Two Cache Risks: Replay Failures and Hidden Grounding Loss
Evaluation · Sep 5 · 13:25
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
Evaluation · Sep 5