Think Before You Paint: Recursive Latent Reasoning for Diffusion Models
Painter-Thinker adds recursive latent reasoning to diffusion, solving 92.5% of hard MNIST Sudoku puzzles.
Painter-Thinker (PaTh) uses a small recursive Thinker that reasons over learned tokens encoding the noisy image and conditioning, refines a latent state inside every denoising step, and steers a frozen diffusion Painter through ControlNet adapters. The Thinker is trained with reconstruction loss alone, without symbolic targets, a solver, or a verifier. PaTh solves 92.5% of hard MNIST Sudoku puzzles versus a prior best of 75%, and 71.2% of extreme puzzles versus 4.1%, with 10M parameters against 82M for a standard diffusion model. It also improves mazes, Queens, and CLEVR scenes, and can recover from injected mistakes that the diffusion model cannot repair.
- A recursive Thinker steers a frozen diffusion Painter through ControlNet adapters.
- Training uses reconstruction loss only, with no symbolic targets, solver, or verifier.
- Hard MNIST Sudoku reaches 92.5% versus 75%; extreme puzzles reach 71.2% versus 4.1%.
- PaTh uses 10M parameters versus 82M for a standard diffusion model.
- It also improves mazes, Queens, and CLEVR scenes with spatial relations.
Full article227 words · extracted from arxiv.org · click to collapse
Diffusion models generate realistic images but often fail on visual reasoning tasks, such as filling in a Sudoku or drawing the path through a maze. When a discrete symbolic representation is available, recursive methods such as the Tiny Recursive Model (TRM) solve even hard instances of these puzzles. We ask how such reasoning can be carried over to pixels, where no symbolic representation is available. We propose Painter-Thinker (PaTh): a small recursive network (the Thinker) reasons over a grid of learned tokens that encode the noisy image and the conditioning, refines a latent state within every denoising step, and steers a frozen diffusion model (the Painter) through ControlNet adapters. The Thinker is trained with the standard reconstruction loss alone, without symbolic targets, a solver, or a verifier. PaTh solves 92.5% of hard MNIST Sudoku puzzles (prior best 75%) and 71.2% of extreme ones (prior best 4.1%), with 10M parameters against 82M for a standard diffusion model. It also improves on mazes, Queens, and CLEVR scenes with specified spatial relations, and its advantage grows with problem size. Diagnostic experiments show that PaTh recovers from injected mistakes that the diffusion model cannot repair, especially when many cells are wrong. Together, these results show that reasoning mechanisms developed for symbolic data can be integrated into pixel-space diffusion without symbolic supervision, opening a path toward generating data under increasingly complex constraints.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.09876