BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Hong_Kong
X-LIC-LOCATION:Asia/Hong_Kong
BEGIN:STANDARD
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
TZNAME:HKT
DTSTART:19911015T033000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20251218T030657Z
LOCATION:Meeting Room S421\, Level 4
DTSTART;TZID=Asia/Hong_Kong:20251215T165100
DTEND;TZID=Asia/Hong_Kong:20251215T170200
UID:siggraphasia_SIGGRAPH Asia 2025_sess112_papers_1586@linklings.com
SUMMARY:HRM^2Avatar: High-Fidelity Real-Time Mobile Avatars from Monocular
  Phone Scans
DESCRIPTION:Chao Shi (Alibaba Group); Shenghao Jia (Shanghai Jiao Tong Uni
 versity, Alibaba Group); Jinhui Liu, Yong Zhang, Liangchao Zhu, Zhonglei Y
 ang, and Jinze Ma (Alibaba Group); Chaoyue Niu (Shanghai Jiao Tong Univers
 ity); and Chengfei Lyu (Alibaba Group)\n\nWe present HRM^2Avatar, a novel 
 framework for creating high-fidelity avatars from monocular phone scans, w
 hich can be rendered and animated in real-time on mobile devices. Monocula
 r capture with commodity smartphones provides a low-cost, pervasive altern
 ative to studio-grade multi-camera rigs, making avatar digitization access
 ible to non-expert users. Reconstructing\nhigh-fidelity avatars from singl
 e-view video sequences poses significant challenges due to deficient visua
 l and geometric data relative to multi-camera setups. To address these lim
 itations, at the data level, our method leverages two types of data captur
 ed with smartphones: static pose sequences for\ndetailed texture reconstru
 ction and dynamic motion sequences for learning pose-dependent deformation
 s and lighting changes. At the representation level, we employ a lightweig
 ht yet expressive representation to reconstruct high-fidelity digital huma
 ns from sparse monocular data. First, we extract explicit garment meshes f
 rom monocular data to model clothing deformations\nmore effectively. Secon
 d, we attach illumination-aware Gaussians to the mesh surface, enabling hi
 gh-fidelity rendering and capturing pose-dependent lighting changes. This 
 representation efficiently learns high-resolution and dynamic information 
 from our tailored monocular data, enabling the creation of detailed avatar
 s. At the rendering level, real-time performance\nis critical for renderin
 g and animating high-fidelity avatars in AR/VR, social gaming, and on-devi
 ce creation, demanding sub-frame responsiveness. Our fully GPU-driven rend
 ering pipeline delivers 120 FPS on mobile devices and 90 FPS on standalone
  VR devices at 2K resolution, over 2.7X faster than representative mobile-
 engine baselines. Experiments show that HRM^2Avatar delivers superior visu
 al realism and real-time interactivity at high resolutions, outperforming 
 state-of-the-art monocular methods.\n\nRegistration Category: Full Access,
  Full Access Supporter\n\nSession Chair: Sebastian Starke (Meta)\n\n
END:VEVENT
END:VCALENDAR
