UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
UltraText Bench scores dense bilingual text rendering across 432 prompts and 24 image models.
UltraText Bench is a bilingual benchmark for prompt-only generation of dense visual text, with 432 human-reviewed prompts across 24 scene categories and three difficulty levels, split equally between English and Chinese. Each prompt specifies exact strings for four to twelve text regions plus structured references for content, placement, and visual attributes. Q-Judger scores text fidelity, clarity, spatial quality, and scene quality. Across 24 model configurations, Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base but loses 14.76 fidelity points, and Qwen-Image-2512's English composite falls from 86.50 at level 1 to 42.86 at level 3; ten people checked the automatic scores.
- 432 bilingual prompts cover 24 scene categories and three difficulty levels.
- Each prompt specifies four to twelve exact text regions with placement references.
- Z-Image-Turbo gains 3.81 clarity points but loses 14.76 fidelity points.
- Qwen-Image-2512 English composite falls from 86.50 at L1 to 42.86 at L3.
Full article166 words · extracted from huggingface.co · click to collapse
Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: https://github.com/LINs-lab/UltraText_Bench.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.09823