Get Started
Home
Topics
Search
Library
Research questionHow can authorship attribution remain reliable for long-form LLM text across languages and distribution shifts?Authorship-attribution systems are often tested on English, shorter, or controlled text and can lose accuracy when the text distribution changes. Multilingual transfer may also depend strongly on the language pair and the detector’s underlying signals.
AI
Evaluation & Benchmarks
Machine Learning
Natural Language Processing
Latest papersRecent research connected to this question, newest first.MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution ShiftsThe evidence covers 928 books generated by five recent LLMs in six languages and three scripts, averaging about 59,000 words each. It evaluates domain, author, and language shifts; reported methods show no consistent winner, with transformer-based transfer varying by language pair and statistical or fingerprint-based methods being more language-dependent.research paper · Sep 2, 2026
Related questions
How can watermarking provide trustworthy provenance for LLM text at scale despite transformations and accumulating false positives?How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?How can watermarking preserve reliable provenance detection after LLM text is paraphrased or translated?How can activation steering represent and control multidimensional authorship style without dedicated training?