Get Started
Home
Topics
Search
Library
Research questionWhen do compute-saving responsible-AI evaluations preserve full-benchmark conclusions?Changing batching, numerical precision, or benchmark size alters the measurement protocol, so stable aggregate accuracy may not imply stable bias, subgroup, or reasoning conclusions. Reduced subsets can also be sensitive to which benchmark items are retained.
AI
Alignment & Safety
Evaluation & Benchmarks
Inference Optimization
Machine Learning
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark ConclusionsThe evidence covers three dense and mixture-of-experts models evaluated on BBQ and BBQ-V under batching, quantization, benchmark reduction, and combined conditions, compared with a full-benchmark BF16 baseline. It examines accuracy, bias severity and prevalence, reasoning quality, subgroup behavior, subset-membership stability, runtime, and measured GPU energy; generalization beyond these models, datasets, and tested conditions is not established.research paper · Sep 4, 2026
Related questions
How can we reduce per-task LLM-agent evaluation cost without distorting benchmark outcomes?How do we evaluate whether scientific agents make justified discoveries from data rather than reproduce known analyses?How can agentic benchmarks be compared and reused across complex environments and bespoke agent integrations?How can AI-assisted scoring reduce grading workload in national assessments without compromising human-judged scores or pass/fail decisions?