Get Started
Topic · 81 recaps
Evaluation & Benchmarks
How we measure model capability — designing benchmarks, spotting contamination, judging open-ended outputs, and stress-testing claims about progress.
Play all
...
Posts
Questions
Home
Topics
Search
Library
Sort
Newest
$Φ$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Agents · Sep 9
0
A Hallucination Score Is Two Different Things
Evaluation · Sep 8 · 11:56
0
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
Agents · Sep 8
0
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
Agents · Sep 8
0
Three Ways "Our Agent's Memory Still Works" Can Be Wrong
Agents · Sep 7 · 12:45
0
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
Agents · Sep 7
0
Two Different Failures Look Identical in Your Audio Eval
Audio/Speech · Sep 6 · 14:03
0
Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning
Evaluation · Sep 6
0
Two Cache Risks: Replay Failures and Hidden Grounding Loss
Evaluation · Sep 5 · 13:25
0
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
Evaluation · Sep 5
0