Get Started
Research questionHow should Irish tokenization be evaluated for alignment with morphological boundaries?Irish tokenization lacks dedicated resources for determining whether token boundaries correspond to morphological boundaries. Standard tokenization measures may not capture this linguistic alignment or its trade-offs with efficiency.
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.MoirfEolas and CríochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish MorphologyThe evidence concerns Irish words annotated for eclipses, prefixes, and suffixes in a resource of over 35,000 words, and evaluates common tokenization algorithms with morphological-alignment and intrinsic tokenization measures. It reports comparisons using CríochScore and finds differing alignment and efficiency trade-offs, with the Unigram Language Model showing stronger morphological alignment among the evaluated algorithms.research paper · Sep 4, 2026
Related questions
How should subtoken granularity be chosen to reduce masked diffusion language-model training loss?How should phonetic encoders be evaluated for preserving pairwise similarities among IPA transcriptions?Can large language models reliably perform Arabic morphosyntactic tagging and dependency parsing despite morphological and orthographic ambiguity?How can we measure and control an LLM’s reliance on token-frequency priors when context is sparse?
Home
Topics
Search
Library