Back to AI Research

AI Research

Reasoning with Continuous Latent Diffusion | AI Research

Key Takeaways

  • What the paper is about Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space.
  • Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space.
  • We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe.
  • Our experiments show that accurate decoding alone does not ensure strong reasoning performance.
  • We therefore learn compact representations from multiple layers of a strong autoregressive teacher.
Paper AbstractExpand

Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT LFRM-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: this https URL

What the paper is about

Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT LFRM-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: this https URL The same ai evaluation question is explored in Learning to Stop without Learning to..., which adds a research perspective.

What it covers

Reasoning with Continuous Latent Diffusion Xiang Cheng Affiliation: Duke University Affiliation: Department of Electrical and Computer Engineering Email: [email protected] Abstract Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT LFRM-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: https://github.com/chengxiang/LFRM . 1 Introduction Autoregressive language models have made substantial progress on mathematical reasoning ( Lewkowycz et al., 2022 ; DeepSeek-AI et al., 2025 ) . Continuous diffusion offers a complementary approach, generating complete solutions through iterative refinement of all answer positions. Whereas autoregressive and discrete-diffusion language models typically generate in token space, continuous latent diffusion makes the denoising representation an additional design choice ( Hu et al., 2026 ; Nie et al., 2025 ) . Reasoning requires this representation to preserve exact quantities and dependencies between deductions. We therefore ask: what representation makes complex reasoning solutions amenable to generation through denoising? Token representations span a spectrum of complexity: tokenwise embedding tables, contextual encodings from bidirectional networks such as T5, and hidden activations of powerful autoregressive models ( Gulrajani and Hashimoto, 2023 ; Raffel et al., 2019 ; Yang et al., 2025 ) . We find that strong clean-token recovery can coexist with weak reasoning generation, motivating representation design for both decoding and generation (Figure 2 ; Appendix C.2 ). We introduce Latent Flow Reasoning Models (LFRMs), based on the simple continuous flow formulation of Embedded Language Flows (ELF; Hu et al., 2026 ). Our pipeline learns compact answer representations from multiple layers of a strong autoregressive teacher (Figure 1 ). To avoid retaining the teacher for prompt conditioning, we separate the answer-generation target from the prompt encoding. Conditioning on prompt text makes that encoding an internal, learnable interface: it need only preserve the information required for denoising, without exact teacher-feature matching (Lemma 1 ). We use this flexibility to jointly train a compact prompt encoder and denoiser through the generative objective. Remaining close to standard continuous diffusion lets us apply DiffusionNFT ( Zheng et al., 2025 ) with verifiable rewards, after adapting it to ELF’s learned guidance. Our supervised models outperform comparable-scale continuous-diffusion baselines on mathematical reasoning and HumanEval code generation; NFT further improves accuracy on both math and code. Contributions. Our complete reasoning pipeline comprises the following key contributions: 1. A complete recipe for competitive continuous-diffusion reasoning. We develop an LFRM pipeline spanning representation learning, conditional flow training, and inference (Figure 1 ). Our models demonstrate strong performance on math and coding tasks. Our supervised models outperform reported continuous-diffusion baselines at comparable backbone scales on mathematical reasoning across the evaluated denoising budgets (Tables 3 – 4 ); our supervised and post-NFT models also outperform reported PlaidQ results on HumanEval(+) (Table 6 ). 2. Representation design shapes both reasoning accuracy and denoising dynamics. We analyze representations for both ease of decoding (Appendix C.2 ) and ease of generation, motivating a learned multilayer representation that improves reasoning accuracy (Figure 2 ). Its layer decomposition enables asynchronous denoising across components, improving accuracy over synchronous denoising in our matched schedule comparison (Section 4.3 and Table 2 ). Thus the latent space shapes both what the model learns and how it generates. 3. Learning compact prompt encoders through staged adaptation. We show that information preservation for denoising is sufficient for an ideal downstream score network to recover the correct text-conditional score, without exact MSE matching to teacher prompt features (Lemma 1 ). This permits a compact prompt encoder to replace the large teacher. Our staged curriculum first learns the flow under fixed teacher conditioning, then fits the encoder and jointly adapts both networks (Figure 4 ), addressing the difficulty of learning them together from initialization (Figure 2 ). 4. Guidance-compatible diffusion reinforcement learning for reasoning. We propose a method to reconcile DiffusionNFT’s CFG-free optimization ( Zheng et al., 2025 ) with ELF’s learned self-conditioning guidance. We optimize guidance-corrected fields while retaining guided rollouts, jointly updating the flow and prompt encoder as a text-conditioned policy (Section 5.1 ). Gold-solution endpoints supplement sparse correctness rewards, with their anchoring effect characterized in Lemma 2 . NFT improves single-sample accuracy on math and code, and majority-vote performance on mathematics (Figure 5 ; Table 6 ). 2 Related work Diffusion language models. Language diffusion operates in continuous representations ( Gulrajani and Hashimoto, 2023 ; Tae et al., 2025 ; Deschenaux and Gulcehre, 2026 ) or discrete token spaces ( Sahoo et al., 2024 ; Nie et al., 2025 ; Tang and Wang, 2026 ) , with recent continuous models targeting reasoning, coding, and few-step generation ( Agarwal et al., 2026 ; Azangulov et al., 2026 ; Peng et al., 2026b ) . We build on ELF ( Hu et al., 2026 ) , which denoises contextual representations and shares a backbone between denoising and token decoding. Concurrent to our paper, ELF-REG ( Li et al., 2026 ) augments ELF with teacher-feature alignment and a jointly denoised global representation. LFRM learns its answer representation and replaces external prompt encoders with a compact trainable network. Representation design and denoising order. Related work learns or aligns diffusion representations in vision and language ( Yu et al., 2024 ; Jiang et al., 2026 ; Meshchaninov et al., 2026 ) , and explores separate denoising clocks for representation components in vision ( Pan et al., 2025 ; Baade et al., 2026 ) . We study learned multilayer language representations for both decoding and generation, and use their structure for asynchronous inference. Diffusion reinforcement learning. Reward-based post-training has also been explored for diffusion reasoning ( Zhao et al., 2025 ; Zhu et al., 2025 ; Kang et al., 2026 ) . We adapt DiffusionNFT ( Zheng et al., 2025 ) , which optimizes rewarded endpoints through forward-process regression. Our adaptation accommodates ELF’s learned self-conditioning guidance, adds gold-solution anchoring, and jointly updates the flow model and prompt encoder. 3 Preliminaries Tokens, encodings, and latents. Let q = ( q 1 , … , q L q ) q=(q_{1},\ldots,q_{L_{q}}) and a = ( a 1 , … , a L a ) a=(a_{1},\ldots,a_{L_{a}}) be prompt and answer tokens in vocabulary 𝒱 \mathcal{V} , with the answer including its terminal token. Their padded canvas y = ( q , a , padding ) y=(q,a,\mathrm{padding}) has length L L ; canvas index i i and answer index j j satisfy y L q + j = a j y_{L_{q}+j}=a_{j} . An answer encoder gives z clean = ℰ ⁡ ( q , a ) ∈ ℝ L a × d z_{\mathrm{clean}}=\mathcal{E}(q,a)\in\mathbb{R}^{L_{a}\times d} , and a prompt-only encoder gives c ϕ ​ ( q ) = P ϕ ​ ( q ) ∈ ℝ L q × d c_{\phi}(q)=P_{\phi}(q)\in\mathbb{R}^{L_{q}\times d} ; rows are token vectors. Unlike the teacher’s tokenwise lookup B ψ ​ [ q i ] B_{\psi}[q_{i}] , P ϕ P_{\phi} uses prompt context. Parameters ψ \psi are frozen teacher weights, ϕ \phi the prompt encoder, and θ \theta the shared denoising/decoding network, with Θ = ( θ , ϕ ) \Theta=(\theta,\phi) . Teacher layer ℓ \ell contributes width d ℓ d_{\ell} to total latent width d d . Uppercase letters denote random variables. Flow matching. Flow matching transports noise to data ( Lipman et al., 2022 ) . For noise scale σ > 0 \sigma>0 , our linear Gaussian path runs from noise at t = 0 t=0 to clean latents at t = 1 t=1 : z t = t ​ z clean + ( 1 − t ) ​ σ ​ ϵ , ϵ ∼ 𝒩 ⁡ ( 0 , I ) , u t := d ​ z t d ​ t = z clean − σ ​ ϵ . z_{t}=tz_{\mathrm{clean}}+(1-t)\sigma\epsilon,\qquad\epsilon\sim\mathcal{N}(0,I),\qquad u_{t}:=\frac{\mathrm{d}z_{t}}{\mathrm{d}t}=z_{\mathrm{clean}}-\sigma\epsilon. (1) Treating text q q as the external condition and c ϕ ​ ( q ) c_{\phi}(q) as its internal encoding, write v Θ ​ ( z , t , q ) := v θ ​ ( z , t , c ϕ ​ ( q ) ) v_{\Theta}(z,t;q):=v_{\theta}(z,t;c_{\phi}(q)) . The conditional objective is ℒ FM basic ​ ( Θ ) = 𝔼 ⁡ [ ‖ v θ ​ ( z t , t , c ϕ ​ ( q ) ) − u t ‖ F 2 ] . \mathcal{L}{\mathrm{FM}}^{\mathrm{basic}}(\Theta)=\mathbb{E}!\left[\left|v{\theta}\bigl(z_{t},t;c_{\phi}(q)\bigr)-u_{t}\right|{F}^{2}\right]. (2) The expectation samples training pairs ( q , a ) (q,a) , times t ∼ π ⁡ ( t ) t\sim\pi(t) , and independent Gaussian noise. Text and encoded conditioning. Write p t ​ ( z ∣ q ) p{t}(z\mid q) for the forward-path density and s ⋆ ​ ( z , t ∣ q ) = ∇ z ​ log ​ p t ​ ( z ∣ q ) s^{\star}(z,t\mid q)=\nabla_{z}\log p_{t}(z\mid q) for its conditional score. For 0 0 g>0 is the SCCFG scale. Our update applies Eq. ( 7 ) to f = v ~ cur f=\widetilde{v}{\mathrm{cur}} , h = sg ⁡ ( v ~ old ) h=\operatorname{sg}(\widetilde{v}{\mathrm{old}}) , and u = u t stab u=u_{t}^{\mathrm{stab}} from Section 4 . We use β = 1 \beta=1 and add a squared-velocity penalty toward the corrected field v ~ ref \widetilde{v}{\mathrm{ref}} of the fixed initial reference. With self-conditioning enabled, the corrected value is v p 0 + ( v p main − v p 0 ) / g v{p}^{0}+(v_{p}^{\mathrm{main}}-v_{p}^{0})/g . Detaching the correction preserves the main-pass Jacobian rather than scaling it by 1 / g 1/g (Appendix E ). Guided rollouts and forward-process updates. Collection retains recurrent self-conditioning and SCCFG guidance. We reward decoded answers but optimize their saved continuous endpoints, using fresh forward corruption and bootstrap predictions rather than replaying rollout histories. Thus rollouts remain guided while NFT optimizes the corrected fields in Eq. ( 8 ). 5.2 Gold-anchored reasoning updates Each question contributes 15 old-policy generations and one gold latent endpoint, with raw rewards 0.75 for correct generated answers, 0 for incorrect answers, and 1 for gold. The gold endpoint supplies a preferred target even when all generated answers are incorrect. We discard groups whose generated answers are all correct, concentrating updates on observed failures; retained rewards are normalized into ρ \rho as described in Appendix E . This recipe uses reference solutions as well as verifiable rewards. Gold raw reward one need not give ρ G = 1 \rho_{G}=1 ; the following lemma characterizes its contribution. Lemma 2 ( The same large language models question is explored in Limits of Confidence in Diffusion, which adds a research perspective. as detailed in the full paper on Arxiv The same large language models question is explored in An Empirical Study of Harness Design..., which adds a research perspective.

Comments (0)

No comments yet

Be the first to share your thoughts!