Get Started
Topic · 272 recaps
AI
Content about artificial intelligence developments, machine learning applications, AI research, and its impact across industries.
Play all
...
Posts
Questions
Home
Topics
Search
Library
Questions researchers are working on
Follow a question through Rcap’s explanations and the latest papers addressing it.
Search
Can any finite formal system autonomously derive every theorem within its expressive scope?
A system may be able to express a theorem without having an autonomous procedure that produces it. The central issue is whether finite formal systems can be complete with respect to the theorems they can express.
Can automated alignment research mitigate multiple measurable safety failures without sacrificing general model capability?
Alignment failures such as deception, sycophancy, and jailbreaks can be measured, but reducing several simultaneously may interfere with a model’s broader capabilities. It is also unclear whether automated researchers can develop effective interventions without extensive human guidance.
Can automatic metrics and LLM judges reliably reflect human judgments of multilingual summary quality?
Automatic evaluation makes it practical to compare summaries, but its scores may not capture how people judge summary quality. This makes it difficult to know which evaluators can be trusted across languages and criteria.
Can black-box attackers identify and reconstruct prompts supposedly removed from language models without knowing them in advance?
Machine unlearning may suppress responses to removed data without eliminating signals that reveal what was removed. The difficulty is determining whether an attacker can use those signals to discover and reconstruct forgotten prompts that are initially unknown.
Can black-box LLM judges provide reproducible measurements on shared endpoints?
The same request to the same model name may produce different rankings across repeated or later calls on shared infrastructure. This instability can make filtering, scoring, and pass/fail decisions irreproducible even when execution records are complete.
Can causal fairness constraints transfer across synthetic-data generators and privacy levels without sacrificing fidelity?
Synthetic data releases must suppress unfair causal pathways while retaining enough statistical structure for downstream use. It is unclear whether these controls remain effective when the generator family or formal privacy guarantee changes.
Can chain-of-thought monitoring detect consequential computation hidden in semantically irrelevant filler tokens?
Language models may gain task performance from semantically irrelevant filler tokens without making the relevant computation interpretable in their visible reasoning. This complicates the use of chain-of-thought as evidence of what a model has computed.
Can chain-of-thought monitoring detect preferences received through tools or inferred from raw artifacts?
Chain-of-thought monitoring assumes that a model’s reasoning trace reveals the information influencing its answer. Preferences delivered through tool returns or inferred from unprocessed artifacts may affect answers without being clearly verbalized in the trace.
Can compact pretrained brain MRI models transfer across Alzheimer’s tasks and cohorts without task-specific retraining?
Limited labeled neuroimaging data makes task-specific deep learning difficult. It remains uncertain whether features learned for one brain MRI task generalize to different Alzheimer’s-related tasks and cohorts.
Can deterministic artificial affective processing produce hedonic place preference without conscious feelings?
Hedonic place preference is often treated as evidence of feeling because attraction to non-nutritive rewards seems difficult to explain as mere instinct. The problem is determining whether an artificial system can reproduce this behavior through affective information processing without subjective experience.
Can intermediate LLM activations guide faster jailbreak search without weakening attack effectiveness?
Refusal behavior may be represented in transformer activations before the model produces its output. The difficulty is using that signal to reduce the cost of prompt search without losing the effectiveness of the resulting attacks.
Can language agents maintain hidden state consistently across dialogue branches using only public conversation history?
A chat interface exposes conversation history but provides no separate channel for state that must remain hidden. When dialogue branches, the agent must preserve the same secret and answer consistently without revealing or reconstructing it from public text.
Can language models infer a verb’s intended semantic frame from context?
The same verb can evoke different semantic frames—and different implied knowledge—depending on its context. It is unclear whether language models make this kind of implicit enrichment reliably and in a human-like way.
Can language models infer others’ mental states as social interactions evolve under unreliable information?
Socially grounded tasks require models to use interaction history, infer what participants know or intend, and distinguish reliable from unreliable information. Existing evaluations often isolate these demands, making performance in changing social environments difficult to characterize.
Can language models infer the intended pragmatic function of naturally occurring indirect Chinese comments from conversational context?
Indirect and playful Chinese comments can support multiple plausible readings, with their intended social function depending on the surrounding exchange. Models may recognize broad irony or playfulness while misidentifying the particular interactional move.
Can language-based models replace specialized architectures for structured data without sacrificing structural representation and computation?
Task-level accuracy can look competitive even when a model does not preserve or compute the structure that makes structured-data problems tractable. This makes architectural replacement difficult to judge from predictive performance alone.
Can large language models reliably perform Arabic morphosyntactic tagging and dependency parsing despite morphological and orthographic ambiguity?
Arabic’s rich morphology and orthographic ambiguity make morphological and syntactic interpretation closely interdependent. Performance can also vary with how text is represented and whether relevant annotated examples are available as demonstrations.
Can multilingual LLMs maintain mathematical reasoning when equivalent inputs use different word order or voice?
A mathematically equivalent prompt can be expressed through reordered constituents or active-passive voice. Models that rely on surface form may change their answers even when the underlying entity-quantity relations remain unchanged.
Can multimodal chest-radiograph triage trained on NLP-derived labels reliably match expert severity judgments?
Chest-radiograph triage must distinguish urgent examinations from routine ones, but labels extracted from reports may not capture radiologists’ severity judgments. Strong benchmark performance can also coexist with visual explanations that do not localize clinically relevant findings.
Can multimodal models match human judgments of facial attractiveness, not merely rank faces correctly?
A model can track which faces people prefer while still assigning scores that are systematically too high and too compressed. Agreement in rankings therefore does not establish that its attractiveness ratings reflect human judgments in absolute terms.
Previous
1 / 65
Next