Get Started
Home
Topics
Search
Library
Research questionHow can large-scale phylogenetic inference avoid manual cognacy annotation while remaining computationally tractable?Global linguistic phylogenies require comparing lexical forms across many languages, but character-based approaches depend on labor-intensive cognacy judgments. Scaling inference to thousands of language varieties also creates substantial computational demands.
AI
Machine Learning
Natural Language Processing
Research Paper
Technology
Latest papersRecent research connected to this question, newest first.Self-Supervised Lexical Representation Learning for Fast, Large-Scale Phylogenetic InferenceThe study concerns raw IPA-transcribed wordlists and global phylogenetic inference across 3,399 language varieties. Its self-supervised representations require no cognacy annotations, alignments, or additional expert input; the reported tree is competitive with several baselines and can be inferred in minutes on a standard notebook GPU. The representations are also evaluated for ranking diachronic concept stability, with results correlated to established rankings.research paper · Sep 4, 2026
Related questions
How can large language models allocate reasoning computation to preserve accuracy under limited training and inference budgets?How can large language models cut training and inference costs without materially harming accuracy?How can missing typological features be predicted with interpretable evidence, including for low-resource languages?Can language-based models replace specialized architectures for structured data without sacrificing structural representation and computation?