New training trick cuts image‑generation error nearly 30%
A USC-led team says a technique called FuseReg, which randomly mixes encoder layers during training, cut a diffusion model's image-generation error score by nearly 30 percent.
Researchers led by USC's Physical Superintelligence Lab said a training technique called FuseReg cut a diffusion model's image-generation error score nearly 30 percent. The gFID metric fell from 13.96 to 9.93 on a standard benchmark, the team reported.
Representation autoencoders reuse a pretrained visual encoder's features as latents for reconstruction and generation. Shallow layers preserve pixel detail, while deeper layers suit diffusion better, creating a trade-off. FuseReg trains the decoder and diffusion model on a random mix of encoder layers at each step, instead of one fixed split, the paper said.
The team reported a similar gain on a larger DiT-XL model, where gFID fell from 2.91 to 2.38. Swapping in a FuseReg-trained decoder alone cut a separate run's gFID from 3.01 to 2.21, without retraining the generator. The method adds no cost at inference time, since the mixing happens only during training.
The paper appeared on Hugging Face's daily papers list with 31 upvotes. It lists lead author Hongyang Du of USC's PSI Lab, with co-authors from Brown, Rice, Notre Dame, Maryland and Pennsylvania. The team said the approach generalized to other encoders, including DINOv3-L and SigLIP2-L. Code is posted on GitHub, and trained checkpoints are hosted on Hugging Face under access-controlled terms.
The technique could generalize further, according to the authors, though the results are the team's own and have not been independently reproduced.