BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Hong_Kong
X-LIC-LOCATION:Asia/Hong_Kong
BEGIN:STANDARD
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
TZNAME:HKT
DTSTART:19911015T033000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20251218T030657Z
LOCATION:Meeting Room S426+S427\, Level 4
DTSTART;TZID=Asia/Hong_Kong:20251215T154400
DTEND;TZID=Asia/Hong_Kong:20251215T155500
UID:siggraphasia_SIGGRAPH Asia 2025_sess110_papers_1364@linklings.com
SUMMARY:SeqTex: Generate Mesh Textures in Video Sequence
DESCRIPTION:Ze Yuan, Xin Yu, and Yangtian Sun (University of Hong Kong); Y
 uan-Chen Guo, Yan-Pei Cao, and Ding Liang (VAST); and Xiaojuan Qi (Univers
 ity of Hong Kong)\n\nTraining native 3D texture generative models remains 
 a fundamental yet challenging problem, largely due to the limited availabi
 lity of large-scale, high-quality 3D texture datasets. This scarcity hinde
 rs generalization to real-world scenarios. To address this, most existing 
 methods finetune foundation image generative models to exploit their learn
 ed visual priors. However, these approaches typically generate only multi-
 view images and rely on post-processing to produce UV texture maps-- an es
 sential representation in modern graphics pipelines. Such two-stage pipeli
 nes often suffer from error accumulation and spatial inconsistencies acros
 s the 3D surface.   In this paper, we introduce SeqTex, a novel end-to-end
  framework that leverages the visual knowledge encoded in pretrained video
  foundation models to directly generate complete UV texture maps. Unlike p
 revious methods that model the distribution of UV textures in isolation, S
 eqTex reformulates the task as a sequence generation problem, enabling the
  model to learn the joint distribution of multi-view renderings and UV tex
 tures. This design effectively transfers the consistent image-space priors
  from video foundation models into the UV domain. To further enhance perfo
 rmance, we propose several architectural innovations: a decoupled multi-vi
 ew and UV branch design, geometry-informed attention to guide cross-domain
  feature alignment, and adaptive token resolution to preserve fine texture
  details while maintaining computational efficiency. Together, these compo
 nents allow SeqTex to fully utilize pretrained video priors and synthesize
  high-fidelity UV texture maps without the need for post-processing. Exten
 sive experiments show that SeqTex achieves state-of-the-art performance on
  both image-conditioned and text-conditioned 3D texture generation tasks, 
 with superior 3D consistency, texture-geometry alignment, and real-world g
 eneralization.\n\nRegistration Category: Full Access, Full Access Supporte
 r\n\nSession Chair: Beibei Wang (Nanjing University)\n\n
END:VEVENT
END:VCALENDAR
