The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page AI · Looped-DiT · arXiv

260M‑parameter image model tops one 6.5 times bigger, authors say

A new arXiv paper says Looped-DiT, with 260 million parameters, tops a 1.7B baseline on five text-to-image benchmarks by rerunning shared transformer blocks inside each denoising step, with 4.9 times less inference compute.

Ten researchers, including Ziwei Liu, Dahua Lin and Gao Huang, posted a paper on arXiv on 30 September describing the Looped Diffusion Transformer. According to the abstract, the approach improves text-to-image quality by repeatedly running shared transformer blocks within each denoising step, rather than by adding parameters or more denoising steps.

The authors say a 260M-parameter model, Looped-DiT B/16, achieves the best results on DPG-Bench, PRISM, T2I-CoReBench and two more suites. They report it outperforms InternVL-U, a 1.7B-parameter baseline, with 6.5 times fewer parameters and 4.9 times lower inference compute. The figures are the authors' own and nobody else has reproduced them.

The design splits the network into three stages. Blocks before and after the loop run once per denoising step, while the looped blocks run several times with shared weights. The paper adds deep supervision, which applies the flow-matching objective at every loop depth, plus gated or exclusive self-attention to stop repeated passes eroding local detail.

The authors say that, under a fixed inference budget, looping beat simply adding denoising steps. They also acknowledge limits: the evaluation covers roughly 260M parameters at 512 by 512 resolution, and their evidence for latent visual reasoning is mostly behavioural. The paper points to a GitHub repository under OpenSenseNova, which returned a 404 when checked on Thursday, so the code could not be inspected.

Sources 2 sources

  1. Source arXiv
  2. Source Aran Komatsuzaki