BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Hong_Kong
X-LIC-LOCATION:Asia/Hong_Kong
BEGIN:STANDARD
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
TZNAME:HKT
DTSTART:19911015T033000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20251218T030656Z
LOCATION:Meeting Room S221\, Level 2
DTSTART;TZID=Asia/Hong_Kong:20251218T112300
DTEND;TZID=Asia/Hong_Kong:20251218T113400
UID:siggraphasia_SIGGRAPH Asia 2025_sess153_papers_1359@linklings.com
SUMMARY:HOMA: Towards Generic Human-Object Interaction in Multimodal Drive
 n Human Animation with Weak Conditions
DESCRIPTION:Ziyao Huang (University of Chinese Academy of Sciences); Zixia
 ng Zhou (Tencent); Juan Cao (University of Chinese Academy of Sciences); Y
 ifeng Ma and Yi Chen (Tencent); Zejing Rao (University of Chinese Academy 
 of Sciences); Zhiyong Xu, Hongmei Wang, Qin Lin, Yuan Zhou, and Qinglin Lu
  (Tencent); and Fan Tang (University of Chinese Academy of Sciences)\n\nWh
 ile recent advances in human-object interaction (HOI) video generation sho
 wcase promising capabilities for synthesizing coordinated human-object dyn
 amics, existing methods remain constrained by their reliance on meticulous
 ly curated motion sequences and actor-specific data, thereby limiting prac
 tical scalability and user accessibility. Furthermore, generalization to n
 ovel object appearances and interaction scenarios remains understudied. To
  address these limitations, we propose HOMA, a weakly conditioned multimod
 al-driven HOI video generation framework that introduces sparse, decoupled
  motion guidance to enhance controllability and reduce dependency on strin
 gent input conditions. Our approach encodes appearance and motion signals 
 into the dual input space of a multimodal diffusion transformer (MMDiT), f
 using them within a shared context space to enable temporally consistent a
 nd physically plausible interactions. To optimize learning efficiency and 
 feature injection accuracy, we introduce a parameter-space HOI adapter ini
 tialized with pretrained MMDiT weights to preserve prior knowledge while e
 nabling efficient adaptation. Additionally, we design a facial cross-atten
 tion adapter for audio-driven lip synchronization, ensuring anatomically a
 ccurate speech animation. Extensive experiments demonstrate that HOMA achi
 eves state-of-the-art performance in interaction naturalness and generaliz
 ation under weak supervision, outperforming existing methods by significan
 t margins. We further illustrate HOMA's versatility through diverse applic
 ations, including text-conditioned generation and interactive object manip
 ulation, facilitated by a user-friendly demo interface.\n\nRegistration Ca
 tegory: Full Access, Full Access Supporter\n\nSession Chair: Kai Wang (Sim
 on Fraser University)\n\n
END:VEVENT
END:VCALENDAR
