Research questionHow can we build reproducible Armenian LLMs without sacrificing knowledge for fluency?Armenian is morphologically rich, but suitable open pretraining data are scarce. News-focused training can improve fluency while reducing knowledge, and overlap between web-derived training and evaluation data can make capabilities appear stronger than they are.