T2LF: LLM-Guided Multimodal Diffusion for Text-to-Light Field Synthesis

Citations

SCOPUS

0

초록

We present a novel text-driven approach for light field (LF) synthesis. Existing methods typically generate LFs from given images, requiring users to find reference images, which makes it difficult to construct the desired scene directly and limits scene diversity. Moreover, existing methods are mainly designed for limited baselines from training datasets, making it difficult to implement various viewpoint changes and consequently limiting the flexibility of motion. In contrast, our method directly synthesizes LFs from user-provided text descriptions by leveraging the scene understanding capabilities of a multi-modal large language model (LLM) and the generative power of a diffusion model. Given a text prompt describing the desired LF, the multimodal LLM extracts relevant information for LF synthesis, which then guides a diffusion model to produce diverse scenes and motions. This approach enables LF synthesis even with a pre-trained model not initially designed for this purpose, requiring only minimal fine-tuning. The proposed framework enables visually diverse LF synthesis with only text input. Experimental results demonstrate that the synthesized LFs exhibit geometric consistency and achieve advanced synthesis quality compared to existing methods. © 2026 IEEE.

키워드

image synthesislight fieldllmvideo diffusion
제목
T2LF: LLM-Guided Multimodal Diffusion for Text-to-Light Field Synthesis
저자
Yoon, SoyoungAhn, NamhyukPark, In Kyu
DOI
10.1109/WACV61042.2026.00707
발행일
2026
유형
Conference paper
저널명
Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
페이지
7322 ~ 7332