Research questionHow can we tell whether generated text matches an author's style across pragmatic contexts?Common lexical, embedding-based, and language-model-based metrics may disagree with human judgments and overlook how authorial style changes with pragmatic context. This makes it difficult to determine whether a personalized output genuinely resembles its intended author.