Get Started
Home
Topics
Search
Library
Research questionHow can we assess whether language models reliably answer or refuse questions grounded in FDA drug labels?FDA drug labels combine heterogeneous clinical and regulatory information, making it difficult for models to retrieve the right evidence, connect facts across sections, and recognize questions they should refuse.
AI
Alignment & Safety
Evaluation & Benchmarks
Health
Information Retrieval
Natural Language Processing
Reasoning
Latest papersRecent research connected to this question, newest first.FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug AssessmentThe source presents FDARxBench, an expert-curated benchmark developed with FDA regulatory assessors from U.S. FDA drug labels. It covers factual, multi-hop, and refusal tasks and reports experiments on proprietary and open-weight models, identifying gaps in factual grounding, long-context retrieval, and safe refusal behavior.research paper · Sep 2, 2026
Related questions
How reliably can language and vision-language models answer veterinary clinical questions with retrieval or supervised adaptation?How can we test whether language models causally use scientific mechanisms instead of answer-correlated shortcuts?How can we reliably detect when an LLM response is unsupported by its reference documents?How can vision-language models express visually grounded answers with context-appropriate information structure?