Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
Iris-3B shows 3B pixel-space pretraining can match latent image models, without better downstream detail.
Researchers pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch on a 256-to-512-to-1024 curriculum and also convert FLUX.2 Klein base 4B into pixel space. Fine-tuning both for monocular depth and 4x DIV2K restoration finds no significant gain over a latent FLUX.2 Klein prior; the converted pixel model trails slightly. Iris-3B still reaches text-to-image quality competitive with latent models, matching Qwen-Image on OneIG at 1024 squared under official evaluators. Weights and training code are released.
- Iris-3B is a 3B pixel-space text-to-image transformer pretrained through a 256-to-1024 curriculum.
- A converted pixel-space FLUX.2 Klein 4B was compared on depth and 4x restoration.
- Pixel-space priors did not beat a latent FLUX.2 Klein fine-tune on those tasks.
- Iris-3B matches Qwen-Image on OneIG at 1024 squared; weights and code are released.
Full article216 words · extracted from huggingface.co · click to collapse
Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters. We test this claim along both routes to a pixel-space backbone. We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a 256to512to1024 curriculum, after first ablating the prediction target and representation alignment at 256^2 to decide what to scale. We also convert a pretrained latent model, FLUX.2 Klein base 4B, to pixel space. We fine-tune both families for monocular depth estimation and for image restoration/super-resolution. We find no significant improvement from using a pixel-space generative prior. Fine-tuned for depth with one matched direct-regression recipe, Iris-3B is level with the latent FLUX.2 Klein and the converted pixel FLUX.2 Klein falls behind it, and on 4times DIV2K restoration neither pixel model beats a latent FLUX.2 Klein fine-tune, the converted one trailing it slightly. We document the recipes, the failure modes and the remaining confounds behind this negative result. Nevertheless, Iris-3B shows that pixel-space pretraining with the pixel-transformer (PiT) head of PixelDiT scales to 3B parameters and to text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at 1024^2. We release its weights and training code in the hope that they help pave the way for further work on pixel-space generation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.09450