arXiv:2408.06070 Abstract | arXiv Analytics

arXiv:2408.06070 [cs.CV]Abstract References Reviews Resources

ControlNeXt: Powerful and Efficient Control for Image and Video Generation

Bohao Peng, Jian Wang, Yuechen Zhang, Wenbo Li, Ming-Chang Yang, Jiaya Jia

Published 2024-08-12Version 1

Diffusion models have demonstrated remarkable and robust abilities in both image and video generation. To achieve greater control over generated results, researchers introduce additional architectures, such as ControlNet, Adapters and ReferenceNet, to integrate conditioning controls. However, current controllable generation methods often require substantial additional computational resources, especially for video generation, and face challenges in training or exhibit weak control. In this paper, we propose ControlNeXt: a powerful and efficient method for controllable image and video generation. We first design a more straightforward and efficient architecture, replacing heavy additional branches with minimal additional cost compared to the base model. Such a concise structure also allows our method to seamlessly integrate with other LoRA weights, enabling style alteration without the need for additional training. As for training, we reduce up to 90% of learnable parameters compared to the alternatives. Furthermore, we propose another method called Cross Normalization (CN) as a replacement for Zero-Convolution' to achieve fast and stable training convergence. We have conducted various experiments with different base models across images and videos, demonstrating the robustness of our method.

Comments: controllable generation

Categories: cs.CV

Keywords: video generation, efficient control, controlnext, base model, substantial additional computational resources

Related articles: Most relevant | Search more

arXiv:2410.22979 [cs.CV] (Published 2024-10-30)

LumiSculpt: A Consistency Lighting Control Network for Video Generation

Yuxin Zhang, Dandan Zheng, Biao Gong, Jingdong Chen, Ming Yang, Weiming Dong, Changsheng Xu

arXiv:2406.02540 [cs.CV] (Published 2024-06-04, updated 2024-06-30)

ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation

Tianchen Zhao et al.

arXiv:2404.13026 [cs.CV] (Published 2024-04-19)

PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation

Tianyuan Zhang et al.

arXiv Analytics

arXiv:2408.06070 [cs.CV]Abstract References Reviews Resources

ControlNeXt: Powerful and Efficient Control for Image and Video Generation

Links

Toolbox

arXiv:2408.06070 [cs.CV]AbstractReferencesReviewsResources

ControlNeXt: Powerful and Efficient Control for Image and Video Generation

Links

Toolbox

arXiv:2408.06070 [cs.CV]Abstract References Reviews Resources