Get Started
Home
Topics
Search
Library
Research questionHow can compact Persian medical QA models reason reliably and estimate answer confidence on consumer hardware?Persian medical QA remains underserved, and compact models must handle clinical reasoning despite limited language-specific resources. Reliable confidence estimates are also needed to distinguish answers that may be unsafe to use.
AI
Alignment & Safety
Evaluation & Benchmarks
Health
Natural Language Processing
Reasoning
Small / On-device Models
Latest papersRecent research connected to this question, newest first.Gaokerena: A Small Persian Medical Language Model FamilyThe source describes Gaokerena-V and Gaokerena-R, trained on Persian medical data and evaluated on a translated medical MMLU benchmark. Both models use uncertainty heads based on internal hidden states; the reported performance is explicitly insufficient to establish readiness for clinical use.research paper · Sep 3, 2026
Related questions
How can we compare medical VQA models when generation obscures answer signal and biases answer positions?How can language models reason iteratively to diagnose complex clinical cases safely and accurately?How reliably can language and vision-language models answer veterinary clinical questions with retrieval or supervised adaptation?How can we assess whether language models reliably answer or refuse questions grounded in FDA drug labels?