Research questionHow can we reliably evaluate specialised style in LLM-generated Chinese legal text?LLM-generated legal text may be factually sound yet violate implicit conventions of legal writing. Human assessment is expensive, while reference-based metrics and LLM judges may conflate content with style or produce opaque, inconsistent judgments.