UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
UltraText Bench scores dense bilingual text rendering across 432 prompts and 24 image models.
UltraText Bench is a bilingual benchmark for prompt-only generation of dense visual text, with 432 human-reviewed prompts across 24 scene categories and three difficulty levels, split equally between English and Chinese. Each prompt specifies exact strings for four to twelve text regions plus structured references for content, placement, and visual attributes. Q-Judger scores text fidelity, clarity, spatial quality, and scene quality. Across 24 model configurations, Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base but loses 14.76 fidelity points, and Qwen-Image-2512's English composite falls from 86.50 at level 1 to 42.86 at level 3; ten people checked the automatic scores.