BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Hong_Kong
X-LIC-LOCATION:Asia/Hong_Kong
BEGIN:STANDARD
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
TZNAME:HKT
DTSTART:19911015T033000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20251218T030657Z
LOCATION:Meeting Room S221\, Level 2
DTSTART;TZID=Asia/Hong_Kong:20251215T164000
DTEND;TZID=Asia/Hong_Kong:20251215T165100
UID:siggraphasia_SIGGRAPH Asia 2025_sess115_papers_1360@linklings.com
SUMMARY:PhySIC: Physically Plausible 3D Human-Scene Interaction and Contac
 t from a Single Image
DESCRIPTION:Pradyumna Yalandur-Muralidhar and Yuxuan Xue (University of Tu
 ebingen); Xianghui Xie (University of Tuebingen, Max Planck Institute for 
 Informatics); Margaret Kostyrko (University of Tuebingen); and Gerard Pons
 -Moll (University of Tuebingen, Max Planck Institute for Informatics)\n\nR
 econstructing metrically accurate humans and their surrounding scenes from
  a single image is crucial for virtual reality, robotics, and comprehensiv
 e 3D scene understanding. However, existing methods struggle with depth am
 biguity, occlusions, and physically inconsistent contacts. To address thes
 e challenges, we introduce PhySIC, a unified framework for physically plau
 sible Human–Scene Interaction and Contact reconstruction. PhySIC recovers 
 metrically consistent SMPL-X human meshes, dense scene surfaces, and verte
 x-level contact maps within a shared coordinate frame, all from a single R
 GB image. Starting from coarse monocular depth and parametric body estimat
 es, PhySIC performs occlusion-aware inpainting, fuses visible depth with u
 nscaled geometry for a robust initial metric scene scaffold, and synthesiz
 es missing support surfaces like floors. A confidence-weighted optimizatio
 n subsequently refines body pose, camera parameters, and global scale by j
 ointly enforcing depth alignment, contact priors, interpenetration avoidan
 ce, and 2D reprojection consistency. Explicit occlusion masking safeguards
  invisible body regions against implausible configurations. PhySIC is high
 ly efficient, requiring only 9 seconds for a joint human-scene optimizatio
 n and less than 27 seconds for end to end reconstruction process. Moreover
 , the framework naturally handles multiple humans, enabling reconstruction
  of diverse human scene interactions. Empirically, PhySIC substantially ou
 tperforms single-image baselines, reducing mean per-vertex scene error fro
 m 641 mm to 227 mm, halving the pose-aligned mean per-joint position error
  (PA-MPJPE) to 42 mm, and improving contact F1-score from 0.09 to 0.51. Qu
 alitative results demonstrate that PhySIC yields realistic foot-floor inte
 ractions, natural seating postures, and plausible reconstructions of heavi
 ly occluded furniture. By converting a single image into a physically plau
 sible 3D human-scene pair, PhySIC advances accessible and scalable 3D scen
 e understanding. Code and evaluation scripts will be publicly released upo
 n publication.\n\nRegistration Category: Full Access, Full Access Supporte
 r\n\nSession Chair: Bo Ren (TMCC, College of Computer Science, Nankai Univ
 ersity)\n\n
END:VEVENT
END:VCALENDAR
