Get Started
Topic · 81 recaps

Evaluation & Benchmarks

How we measure model capability — designing benchmarks, spotting contamination, judging open-ended outputs, and stress-testing claims about progress.
PostsQuestions
Home
Topics
Search
Library
Questions researchers are working onFollow a question through Rcap’s explanations and the latest papers addressing it.
Can automated alignment research mitigate multiple measurable safety failures without sacrificing general model capability?Alignment failures such as deception, sycophancy, and jailbreaks can be measured, but reducing several simultaneously may interfere with a model’s broader capabilities. It is also unclear whether automated researchers can develop effective interventions without extensive human guidance.Can automatic metrics and LLM judges reliably reflect human judgments of multilingual summary quality?Automatic evaluation makes it practical to compare summaries, but its scores may not capture how people judge summary quality. This makes it difficult to know which evaluators can be trusted across languages and criteria.Can black-box LLM judges provide reproducible measurements on shared endpoints?The same request to the same model name may produce different rankings across repeated or later calls on shared infrastructure. This instability can make filtering, scoring, and pass/fail decisions irreproducible even when execution records are complete.Can causal fairness constraints transfer across synthetic-data generators and privacy levels without sacrificing fidelity?Synthetic data releases must suppress unfair causal pathways while retaining enough statistical structure for downstream use. It is unclear whether these controls remain effective when the generator family or formal privacy guarantee changes.Can chain-of-thought monitoring detect consequential computation hidden in semantically irrelevant filler tokens?Language models may gain task performance from semantically irrelevant filler tokens without making the relevant computation interpretable in their visible reasoning. This complicates the use of chain-of-thought as evidence of what a model has computed.Can chain-of-thought monitoring detect preferences received through tools or inferred from raw artifacts?Chain-of-thought monitoring assumes that a model’s reasoning trace reveals the information influencing its answer. Preferences delivered through tool returns or inferred from unprocessed artifacts may affect answers without being clearly verbalized in the trace.Can intermediate LLM activations guide faster jailbreak search without weakening attack effectiveness?Refusal behavior may be represented in transformer activations before the model produces its output. The difficulty is using that signal to reduce the cost of prompt search without losing the effectiveness of the resulting attacks.Can internal photogrammetric validation certify metric accuracy without external survey or control-point measurements?A reconstruction can be internally geometrically consistent while still containing coherent global distortion. This makes metric correctness difficult to establish when external survey or control-point measurements are unavailable.Can language agents maintain hidden state consistently across dialogue branches using only public conversation history?A chat interface exposes conversation history but provides no separate channel for state that must remain hidden. When dialogue branches, the agent must preserve the same secret and answer consistently without revealing or reconstructing it from public text.Can language models infer a verb’s intended semantic frame from context?The same verb can evoke different semantic frames—and different implied knowledge—depending on its context. It is unclear whether language models make this kind of implicit enrichment reliably and in a human-like way.Can language models infer others’ mental states as social interactions evolve under unreliable information?Socially grounded tasks require models to use interaction history, infer what participants know or intend, and distinguish reliable from unreliable information. Existing evaluations often isolate these demands, making performance in changing social environments difficult to characterize.Can language models infer the intended pragmatic function of naturally occurring indirect Chinese comments from conversational context?Indirect and playful Chinese comments can support multiple plausible readings, with their intended social function depending on the surrounding exchange. Models may recognize broad irony or playfulness while misidentifying the particular interactional move.Can large language models reliably perform Arabic morphosyntactic tagging and dependency parsing despite morphological and orthographic ambiguity?Arabic’s rich morphology and orthographic ambiguity make morphological and syntactic interpretation closely interdependent. Performance can also vary with how text is represented and whether relevant annotated examples are available as demonstrations.Can multilingual LLMs maintain mathematical reasoning when equivalent inputs use different word order or voice?A mathematically equivalent prompt can be expressed through reordered constituents or active-passive voice. Models that rely on surface form may change their answers even when the underlying entity-quantity relations remain unchanged.Can multimodal chest-radiograph triage trained on NLP-derived labels reliably match expert severity judgments?Chest-radiograph triage must distinguish urgent examinations from routine ones, but labels extracted from reports may not capture radiologists’ severity judgments. Strong benchmark performance can also coexist with visual explanations that do not localize clinically relevant findings.Can multimodal models match human judgments of facial attractiveness, not merely rank faces correctly?A model can track which faces people prefer while still assigning scores that are systematically too high and too compressed. Agreement in rankings therefore does not establish that its attractiveness ratings reflect human judgments in absolute terms.Can phase-transition counts during fine-tuning predict final test accuracy across architectures and distribution shifts?During fine-tuning, class separability may change in discrete jumps, but the relationship between the number of jumps and eventual test accuracy may depend on architecture and whether evaluation data are i.i.d. or corrupted.Can post-training ternarization make language models smaller without unacceptable capability loss or slower inference?Ultra-low-bit weights can shrink model storage, but nominal bit counts may not reflect the stored representation, uneven task degradation, or actual inference speed. Compression may therefore improve footprint without improving end-to-end deployment performance.Can preprocessing defenses detect adversarial attacks in depthwise-separable edge vision CNNs when they cannot restore predictions?Preprocessing defenses are often assumed to transfer across model architectures, but depthwise-separable CNNs may respond differently to adversarial perturbations than residual or Inception-style networks. Their failure to recover predictions may still produce measurable differences between clean and adversarial inputs, while image-quality scores may not reflect defensive value.Can prompt phrasing reliably improve LLM-derived chemical features for drug-toxicity prediction?Minor changes in prompt phrasing can alter LLM outputs, making it unclear whether prompt optimization produces stable chemical features for toxicity models. This variability complicates the use of LLM-generated features in a costly drug-development process.
Previous
1 / 32
Next