BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Hong_Kong
X-LIC-LOCATION:Asia/Hong_Kong
BEGIN:STANDARD
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
TZNAME:HKT
DTSTART:19911015T033000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20251218T030653Z
LOCATION:Meeting Room S423+S424\, Level 4
DTSTART;TZID=Asia/Hong_Kong:20251217T113400
DTEND;TZID=Asia/Hong_Kong:20251217T114500
UID:siggraphasia_SIGGRAPH Asia 2025_sess135_papers_1747@linklings.com
SUMMARY:Voyager: Long-Range and World-Consistent Video Diffusion for Explo
 rable 3D Scene Generation
DESCRIPTION:Tianyu Huang (Harbin Institute of Technology, City University 
 of Hong Kong); Wangguandong Zheng and Tengfei Wang (Tencent); Yuhao Liu (T
 encent, City University of Hong Kong); Zhenwei Wang, Junta Wu, and Jie Jia
 ng (Tencent); Hui Li (Harbin Institute of Technology); Rynson Lau (City Un
 iversity of Hong Kong); Wangmeng Zuo (Harbin Institute of Technology); and
  Chunchao Guo (Tencent)\n\nReal-world applications like video gaming and v
 irtual reality often demand the ability to model 3D scenes that users can 
 explore along custom camera trajectories. While significant progress has b
 een made in generating 3D objects from text or images, creating long-range
 , 3D-consistent, explorable 3D scenes remains a complex and challenging pr
 oblem. In this work, we present Voyager, a novel video diffusion framework
  that generates world-consistent 3D point-cloud sequences from a single im
 age with user-defined  camera path. Unlike existing approaches, Voyager ac
 hieves end-to-end scene generation and reconstruction with inherent consis
 tency across frames, eliminating the need for 3D reconstruction pipelines 
 (e.g., structure-from-motion or multi-view stereo). Our method integrates 
 three key components:  1) World-Consistent Video Diffusion: A unified arch
 itecture that jointly generates aligned RGB and depth video sequences, con
 ditioned on existing world observation to ensure global coherence 2) Long-
 Range World Exploration: An efficient world cache with point culling and a
 n auto-regressive inference with smooth video sampling for iterative scene
  extension with context-aware consistency, and 3) Scalable Data Engine: A 
 video reconstruction pipeline that automates camera pose estimation and me
 tric depth prediction for arbitrary videos, enabling large-scale, diverse 
 training data curation without manual 3D annotations. Collectively, these 
 designs result in a clear improvement over existing methods in visual qual
 ity  and geometric accuracy, with versatile applications.\n\nRegistration 
 Category: Full Access, Full Access Supporter\n\nSession Chair: Or Patashni
 k (Tel Aviv University, Snap Research)\n\n
END:VEVENT
END:VCALENDAR
