BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Hong_Kong
X-LIC-LOCATION:Asia/Hong_Kong
BEGIN:STANDARD
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
TZNAME:HKT
DTSTART:19911015T033000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20251218T030656Z
LOCATION:Meeting Room S423+S424\, Level 4
DTSTART;TZID=Asia/Hong_Kong:20251217T105000
DTEND;TZID=Asia/Hong_Kong:20251217T110100
UID:siggraphasia_SIGGRAPH Asia 2025_sess135_papers_1416@linklings.com
SUMMARY:SS4D: Native 4D Generative Model via Structured Spacetime Latents
DESCRIPTION:Zhibing Li (The Chinese University of Hong Kong); Mengchen Zha
 ng (Zhejiang University, Shanghai Artificial Intelligence Laboratory); Ton
 g Wu (Stanford University); Jing Tan (Chinese University of Hong Kong); Ji
 aqi Wang (Shanghai AI Laboratory); and Dahua Lin (The Chinese University o
 f Hong Kong)\n\nWe present SS4D, a native 4D generative model that synthes
 izes dynamic 3D objects directly from monocular video. Unlike prior approa
 ches that construct 4D representations by optimizing over 3D or video gene
 rative models, we train a generator directly on 4D data, achieving high fi
 delity, temporal coherence, and structural consistency. At the core of our
  method is a compressed set of structured spacetime latents. Specifically,
  (1) To address the scarcity of 4D training data, we build on a pre-traine
 d single-image-to-3D model, preserving strong spatial consistency. (2) Tem
 poral consistency is enforced by introducing dedicated temporal layers tha
 t reason across frames. \n(3) To support efficient training and inference 
 over long video sequences, we compress the latent sequence along the tempo
 ral axis using factorized 4D convolutions and temporal downsampling blocks
 . In addition, we employ a carefully designed training strategy to enhance
  robustness against occlusion and motion blur, leading to high-quality gen
 eration. Extensive experiments show that {\ourmethod} produces spatio-temp
 orally consistent 4D objects with superior quality and efficiency, signifi
 cantly outperforming state-of-the-art methods on both synthetic and real-w
 orld datasets.\n\nRegistration Category: Full Access, Full Access Supporte
 r\n\nSession Chair: Or Patashnik (Tel Aviv University, Snap Research)\n\n
END:VEVENT
END:VCALENDAR
