Get Started
Research questionHow reliably do human psychometric questionnaires predict LLM behavior in realistic user interactions?Questionnaire items can contain lexical cues that reveal the trait being measured and encourage socially desirable responses. These cues may make questionnaire scores diverge from an LLM’s behavior on less structured user queries.
AI
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Human Psychometric Questionnaires Mischaracterize LLM BehaviorThe evidence covers eight open-source LLMs, comparing value and personality profiles from Likert self-reports on PVQ-40/21 and BFI-44/10 with generation probabilities for value-laden responses to everyday queries. It also examines whether demographic persona prompts produce consistent shifts across the two settings; the findings do not establish validity beyond the tested models, questionnaires, prompts, and profiling procedures.research paper · Sep 2, 2026
Related questions
How reliably do proxy LLM judges capture users’ perceived helpfulness and privacy in privacy-sensitive scenarios?How can we assess whether LLM-generated personas reproduce culturally conditioned worldviews and moral values?How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?How can we evaluate LLM reasoning quality beyond final-answer accuracy across deployment contexts?
Home
Topics
Search
Library