BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Hong_Kong
X-LIC-LOCATION:Asia/Hong_Kong
BEGIN:STANDARD
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
TZNAME:HKT
DTSTART:19911015T033000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20251218T030656Z
LOCATION:Meeting Room S423+S424\, Level 4
DTSTART;TZID=Asia/Hong_Kong:20251217T104000
DTEND;TZID=Asia/Hong_Kong:20251217T105000
UID:siggraphasia_SIGGRAPH Asia 2025_sess135_papers_1405@linklings.com
SUMMARY:VideoFrom3D: 3D Scene Video Generation via Complementary Image and
  Video Diffusion Models
DESCRIPTION:Geonung Kim, Janghyeok Han, and Sunghyun Cho (POSTECH)\n\nIn t
 his paper, we propose VideoFrom3D, a novel framework for synthesizing high
 -quality 3D scene videos from coarse geometry, a camera trajectory, and a 
 reference image. Our approach streamlines the 3D graphic design workflow, 
 enabling flexible design exploration and rapid production of deliverables.
  A straightforward approach to synthesizing a video from coarse geometry m
 ight condition a video diffusion model on geometric structure. However, ex
 isting video diffusion models struggle to generate high-fidelity results f
 or complex scenes due to the difficulty of jointly modeling visual quality
 , motion, and temporal consistency. To address this, we propose a generati
 ve framework that leverages the complementary strengths of image and video
  diffusion models. Specifically, our framework consists of a Sparse Anchor
 -view Generation (SAG) and a Geometry-guided Generative Inbetweening (GGI)
  module. The SAG module generates high-quality, cross-view consistent anch
 or views using an image diffusion model, aided by Sparse Appearance-guided
  Sampling. Building on these anchor views, GGI module faithfully interpola
 tes intermediate frames using a video diffusion model, enhanced by flow-ba
 sed camera control and structural guidance. Notably, both modules operate 
 without any paired dataset of 3D scene models and natural images, which is
  extremely difficult to obtain. Comprehensive experiments show that our me
 thod produces high-quality, style-consistent scene videos under diverse an
 d challenging scenarios, outperforming simple and extended baselines.\n\nR
 egistration Category: Full Access, Full Access Supporter\n\nSession Chair:
  Or Patashnik (Tel Aviv University, Snap Research)\n\n
END:VEVENT
END:VCALENDAR
