Get Started
Research questionHow can we build reproducible Armenian LLMs without sacrificing knowledge for fluency?Armenian is morphologically rich, but suitable open pretraining data are scarce. News-focused training can improve fluency while reducing knowledge, and overlap between web-derived training and evaluation data can make capabilities appear stronger than they are.
Evaluation & Benchmarks
LLM Pretraining & Post-training
Natural Language Processing
Latest papersRecent research connected to this question, newest first.From Zero to Hero: An Open LLM Ecosystem for ArmenianThe evidence concerns continued pretraining of Gemma-4-E4B with the released ArmWeb Armenian news corpus and ArmSTEM English–Armenian math and science problems, alongside ablations and overlap analysis for public Armenian evaluations. Data, models, and code are released; conclusions are limited to the reported Armenian models and evaluation panels.research paper · Sep 3, 2026
Related questions
How can instruction-tuned LLMs learn corpus-specific knowledge without exhaustive synthetic QA or instruction fine-tuning?How should multilingual LLMs estimate uncertainty and calibrate abstention across languages and model sizes?How can Arabic LLMs generate accurate target dialects from MSA prompts without fine-tuning?How should low-resource LLM fine-tuning use task-level language priors with ambiguous or incomplete data?
Home
Topics
Search
Library