Reasoning with Continuous Latent Diffusion
Latent Flow Reasoning Models generate math and code solutions by iterative denoising in continuous latent space.
The paper introduces Latent Flow Reasoning Models, an ELF-based recipe that generates reasoning solutions by iterative refinement in continuous latent space. Compact representations are learned from multiple layers of an autoregressive teacher, with asynchronous denoising and a staged prompt encoder that replaces the teacher Transformer at inference. With a 638M-parameter denoising backbone, post-NFT LFRM-L reaches 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 steps. Supervised models outperform recent continuous-diffusion baselines at comparable backbone scales.
- LFRMs refine complete reasoning solutions by denoising in latent space.
- A 638M denoiser scores 63.74% on GSM8K and 24.6% on MATH500.
- HumanEval and HumanEval+ reach 32.85% and 30.18% pass@1.
- A learned prompt encoder replaces the teacher Transformer at inference.
Full article178 words · extracted from arxiv.org · click to collapse
Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT LFRM-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: https://github.com/chengxiang/LFRM
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35694