arXiv Analytics

Sign in

arXiv:2307.14648 [cs.CV]AbstractReferencesReviewsResources

Spatial-Frequency U-Net for Denoising Diffusion Probabilistic Models

Xin Yuan, Linjie Li, Jianfeng Wang, Zhengyuan Yang, Kevin Lin, Zicheng Liu, Lijuan Wang

Published 2023-07-27Version 1

In this paper, we study the denoising diffusion probabilistic model (DDPM) in wavelet space, instead of pixel space, for visual synthesis. Considering the wavelet transform represents the image in spatial and frequency domains, we carefully design a novel architecture SFUNet to effectively capture the correlation for both domains. Specifically, in the standard denoising U-Net for pixel data, we supplement the 2D convolutions and spatial-only attention layers with our spatial frequency-aware convolution and attention modules to jointly model the complementary information from spatial and frequency domains in wavelet data. Our new architecture can be used as a drop-in replacement to the pixel-based network and is compatible with the vanilla DDPM training process. By explicitly modeling the wavelet signals, we find our model is able to generate images with higher quality on CIFAR-10, FFHQ, LSUN-Bedroom, and LSUN-Church datasets, than the pixel-based counterpart.

Related articles: Most relevant | Search more
arXiv:2307.15988 [cs.CV] (Published 2023-07-29)
RGB-D-Fusion: Image Conditioned Depth Diffusion of Humanoid Subjects
arXiv:2206.11892 [cs.CV] (Published 2022-06-23)
Remote Sensing Change Detection (Segmentation) using Denoising Diffusion Probabilistic Models
arXiv:2209.14828 [cs.CV] (Published 2022-09-29)
Denoising Diffusion Probabilistic Models for Styled Walking Synthesis