TY - GEN
T1 - Unknown Pixel Mask Based Fine-tuning of 2D Inpainting Models for Unbounded 3D Scene Generation from a Single Image
AU - Zheng, Dezhi
AU - Deng, Kaijun
AU - Hou, Xianxu
AU - Wang, Jinbao
AU - Wang, Xiaoqin
AU - Shen, Linlin
N1 - Publisher Copyright:
© 2025 ACM.
PY - 2025/10/27
Y1 - 2025/10/27
N2 - Conventional 2D inpainting models are trained using masks confined to 2D scenarios, resulting in meaningless content when applied to 3D-specific masks. These 3D-specific masks, termed Unknown Pixels (UP) masks, represent unseen pixels from novel viewpoints that remain obscured in the original input image. Existing methods attempt to mitigate this issue by employing post-processing techniques to transform UP masks into 2D equivalents, frequently suffering from unnatural distortions. To address these issues, we investigate the efficacy of directly training 2D inpainting models with UP masks to circumvent such distortions. In this paper, we introduce a novel framework designed to generate unbounded 3D scenes from a single image, guided by textual descriptions. Our approach leverages fine-tuned inpainting models that iteratively reconstruct incomplete images originating from pure projection. The generated points are then seamlessly integrated into the original point cloud via pixel-wise depth alignment. Extensive evaluations demonstrate that our framework outperforms existing methods in scene quality, processing speed, and memory efficiency.
AB - Conventional 2D inpainting models are trained using masks confined to 2D scenarios, resulting in meaningless content when applied to 3D-specific masks. These 3D-specific masks, termed Unknown Pixels (UP) masks, represent unseen pixels from novel viewpoints that remain obscured in the original input image. Existing methods attempt to mitigate this issue by employing post-processing techniques to transform UP masks into 2D equivalents, frequently suffering from unnatural distortions. To address these issues, we investigate the efficacy of directly training 2D inpainting models with UP masks to circumvent such distortions. In this paper, we introduce a novel framework designed to generate unbounded 3D scenes from a single image, guided by textual descriptions. Our approach leverages fine-tuned inpainting models that iteratively reconstruct incomplete images originating from pure projection. The generated points are then seamlessly integrated into the original point cloud via pixel-wise depth alignment. Extensive evaluations demonstrate that our framework outperforms existing methods in scene quality, processing speed, and memory efficiency.
KW - inpainting models
KW - multimodal generative models
KW - point cloud generation
KW - unbounded 3d scene generation
UR - https://www.scopus.com/pages/publications/105024079383
U2 - 10.1145/3746027.3754743
DO - 10.1145/3746027.3754743
M3 - Conference contribution
AN - SCOPUS:105024079383
T3 - MM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025
SP - 9306
EP - 9315
BT - MM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025
PB - Association for Computing Machinery, Inc
T2 - 33rd ACM International Conference on Multimedia, MM 2025
Y2 - 27 October 2025 through 31 October 2025
ER -