Back to AI Research

AI Research

One-step image generators can carry a denoising trajectory across network depth

Key Takeaways

  • Intermediate-layer probes reveal denoising-like structure in selected generators, motivating a smaller MeanFlow model.
  • Generating an image in one network call removes the visible sequence of sampling steps.
  • Arnold Caleb Asiimwe and colleagues investigate whether some of that progressive denoising computation remains inside the network's layers.
  • Their [Depth as Time study](https://arxiv.org/abs/2610.03626) reports a denoising-like progression in selected diffusion-based one-step generators.
  • Intermediate-layer decoding exposes coarse-to-fine image structure, but the authors distinguish that empirical observation from a claim that network depth reproduces the original diffusion timeline exactly.

Generating an image in one network call removes the visible sequence of sampling steps. Arnold Caleb Asiimwe and colleagues investigate whether some of that progressive denoising computation remains inside the network's layers.
Their Depth as Time study reports a denoising-like progression in selected diffusion-based one-step generators. Intermediate-layer decoding exposes coarse-to-fine image structure, but the authors distinguish that empirical observation from a claim that network depth reproduces the original diffusion timeline exactly.

Reading intermediate layers without retraining the generator

The researchers apply a model's own output head to the hidden representation after each block. This probe asks what endpoint prediction the current representation already supports, without executing the remaining blocks or adding learned parameters to the generator.
They compare those decoded outputs with the final prediction. A shrinking distance alone would show convergence, not prove denoising, so the study also fits mixtures of the input noise and final generated image to intermediate predictions. The recovered coefficients provide another test of signal-noise organization across depth.
MeanFlow, FLUX-schnell and Shortcut show structure consistent with that interpretation in the reported experiments. A Drifting model provides a contrasting case: its intermediate predictions are less well explained by the same mixture. The finding depends on the model's training objective and the transport problem it is conditioned to solve.

Some layers denoise before adding noise back

A generator can be asked to move between noisy states rather than produce a clean image immediately. For MeanFlow, the researchers find that intermediate layers can first reveal cleaner structure and then add noise back toward the requested endpoint.
This denoise-then-renoise behavior makes the internal progression less simple than assigning each layer one ordinary diffusion timestep. The recovered mixing schedules can include coefficients outside the interval used by standard flow-matching interpolation. The authors describe depth as its own effective denoising-like coordinate.
In distilled generators used for multiple sampling steps, the paper also observes refinement within each network evaluation. Internal depthwise progression and external sampling time can therefore coexist instead of being mutually exclusive descriptions.

Compression changes both parameter count and quality

The team uses the interpretation to train a single time-conditioned block across layers. On ImageNet 256-by-256 generation, the reported MeanFlow SiT-L/2 compression reduces parameters from 459 million to 27.6 million, a factor of 16.6.
The compressed model initially reaches an FID of 11.7. FD-loss post-training improves that to 4.9, compared with 4.0 for the original model. Those numbers describe a quality cost as well as a parameter reduction. They do not establish a 16.6-fold latency improvement or equal image quality.
Applying the same compression to the Drifting model produces an FID of 47.1 in the reported experiment. That contrast limits the method's scope and supports investigating which trained computations can tolerate this kind of sharing.
The paper offers a way to inspect and compress selected one-step generators. It leaves broader claims about editing control, other architectures and production inference speed to separate validation.

Comments