Get Started
Home
Topics
Search
Library
Research questionHow can we compare language models’ conditional behavior and predict the effects of prompt changes?Model-level summaries can obscure how a language model’s response distribution varies with the prompt. This makes it difficult to connect model differences and prompt changes to downstream task performance.
AI
Evaluation & Benchmarks
Machine Learning
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Language Model Maps for Prompt-Response Distributions via Log-Likelihood VectorsThe source studies publicly available language models using log-likelihood and PMI representations over prompt-response pairs. Its experiments examine model relationships, links to attributes and task performance, systematic prompt shifts, and approximations of composite prompt effects when the corresponding vectors are not directly observed.research paper · Sep 2, 2026
Related questions
How should open reasoning language models be selected under prompt and deployment-resource trade-offs?How can we test whether language models causally use scientific mechanisms instead of answer-correlated shortcuts?How should language-model robustness be evaluated when text perturbations affect hidden states and attention differently from outputs?When do natural-language rules outperform examples for LLM in-context learning?