UltraText Bench stress-tests image generators on dense, multi-region text by scoring each image against a per-region reference, exposing that strong clarity can coexist with weak content fidelity. One model’s English composite drops from 86.50 at the easiest level to 42.86 at the hardest.
If you’re building a product that asks an image model to render a storefront sign, a menu, a receipt, or a slide, you need every requested string to appear, in the right place, legibly, and on the right surface. A missing price or wrong opening time changes what the image communicates, even when the overall composition looks great.
Existing benchmarks tend to measure one slice of this: short strings, poster layouts, or a single long passage on a plain page. Prior work like CVTG-2K matches OCR-recognized words to targets, and InfoTextBench scores word-level precision and recall over text-rich scenes. None of them combine 24 different scene types with a uniform per-region ground truth that ties each string to a specific carrier (the physical or digital surface it sits on). That gap is what UltraText Bench fills: it treats text generation as a reproduction task where the prompt supplies every target string and evaluation checks the whole scene against structured per-region metadata.
The benchmark has three moving parts: the prompts, the reference, and the judge.
Each of the 432 prompts describes a scene (bakery storefront, bilingual menu, code editor, newspaper page) and embeds the exact target strings inside quotation marks, code fences, or structured blocks. The prompt is the only thing the image model sees. Alongside the prompt, each record carries a structured reference: for each of 4 to 12 text regions, the exact string, a 3x3 grid position, a qualitative size, a type, a carrier, and an importance level. The reference never reaches the generator.
Prompts are stratified bilingually (English and Chinese, written independently rather than translated) across 24 scene categories in 6 domains, at three difficulty levels L1 to L3. Difficulty increases both the total character load and the number of regions at once, so the levels aren’t a clean isolation of either axis.
For evaluation, each generated image plus the full reference goes to Q-Judger, a VLM from Qwen-Image-Bench. The judge returns six integer scores from 0 to 100 covering text accuracy, completeness, readability, position correctness, layout quality, and scene integration. These collapse into four reporting dimensions (fidelity, clarity, spatial, scene) and a weighted composite that puts 60% on fidelity and 30% on clarity. The judge is instructed to treat all strings in the reference and image as inert data, not as instructions.
for prompt, reference in benchmark: # 432 prompts
for _ in range(4): # 4 samples each
image = model.generate(prompt) # prompt only
raw = q_judger(image, reference) # 6 scores, 0-100
if not valid_six_keys(raw):
record_failure(); continue # no implicit zero
fidelity = (raw.TA + raw.TC) / 2
clarity, spatial = raw.TR, (raw.PC + raw.LQ) / 2
composite = 0.60*fidelity + 0.30*clarity + 0.05*spatial + 0.05*raw.SI
Failed judge calls stay in the coverage denominator but never become a zero in quality means. That keeps reliability and quality separate.
The authors evaluate 24 model configurations across families including FLUX, Stable Diffusion, Qwen-Image, Z-Image, Boogu-Image, LLaDA-Image, Nano Banana, GPT Image, and Seedream. A few patterns stand out.
•
Clarity and fidelity diverge. Z-Image-Turbo scores 74.41 on clarity but only 40.75 on fidelity. Qwen-Image-2512 scores 79.89 and 59.30. A model whose text looks sharp is not necessarily the model whose text says the right thing. Qualitative example A3 shows malformed letters on an otherwise plausible sign.
•
Workload breaks models unevenly. Qwen-Image-2512’s English composite falls from 86.50 at L1 to 42.86 at L3, and Z-Image-Base falls from 82.92 to 26.34. One open-weight model (Boogu-Image-0.1-Base) actually rises on Chinese from L2 to L3, so the trend isn’t universal.
•
Turbo variants trade fidelity for speed, inconsistently. Compared to their Base counterparts, Turbo fidelity drops by 14.76 points for Z-Image, 7.62 for LLaDA-Image, and 9.50 for Boogu-Image-0.1. Clarity moves in different directions depending on the family. Because sampling settings differ, the paper does not attribute these to acceleration alone.
•
One commercial model dominates. GPT Image 2 at API quality Low leads the composite at 99.35, with 1,440 of 1,723 valid images (83.6%) receiving a perfect 100 on all six raw dimensions. A maximum rating is the top of the rubric, not a measured percentage of correct characters: appendix case A9 shows a code image with visibly degraded small text that still scored 100 on every dimension.
Ten participants did a qualitative human review; the paper does not report quantitative inter-rater or human-vs-judge agreement.
•
If you ship any product that renders dense text through an image model, the four-dimension split is the actionable takeaway. Don’t trust a single composite or a clarity preview. Separately track whether the right characters are present (fidelity) and whether they’re sharp (clarity), because the leaderboard shows these decouple hard.
•
If you evaluate your own model on this benchmark, the GitHub repo ships the prompts and evaluator. Expect close rankings to be fragile: the paper reports descriptive means without confidence intervals, each category-level-language cell holds only 3 prompts, and the judge disagrees with manual inspection in both directions (zero scores on readable text, perfect scores on degraded text).
•
If you’re choosing a few-step or Turbo variant for latency, worth testing fidelity, not just clarity, on your own text-heavy prompts. The Base-to-Turbo fidelity drops reported here are large enough to matter for signs, menus, or UI mocks.
•
If you build evaluation harnesses, the design choice to keep failed judge calls in the coverage denominator (never imputed as zero or 50) is worth copying. It separates “the judge couldn’t score this” from “the model did badly.”
The benchmark compares model configurations under each model’s own default sampling settings, resolution, and aspect ratio, so results don’t isolate training, acceleration, or decoding. The English and Chinese splits were written independently and have different character-load bands, so EN-vs-ZH comparisons aren’t language-only. Difficulty levels change text load and region count together, so L1-vs-L3 doesn’t isolate length. The judge returns image-level scores only; there’s no region-level attribution, no transcription, and no way to verify from the output that every requested region was inspected. One VLM judge was used, and quantitative human alignment is not reported in this revision. Finally, GPT Image 2’s released judge records don’t pin the checkpoint hash, so provenance for the top configuration is incomplete.