BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Hong_Kong
X-LIC-LOCATION:Asia/Hong_Kong
BEGIN:STANDARD
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
TZNAME:HKT
DTSTART:19911015T033000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20251218T030656Z
LOCATION:Meeting Room S221\, Level 2
DTSTART;TZID=Asia/Hong_Kong:20251216T172400
DTEND;TZID=Asia/Hong_Kong:20251216T173500
UID:siggraphasia_SIGGRAPH Asia 2025_sess131_papers_2494@linklings.com
SUMMARY:Learning to Ball: Composing Policies for Long-Horizon Basketball M
 oves
DESCRIPTION:Pei Xu, Zhen Wu, Ruocheng Wang, Vishnu Sarukkai, and Kayvon Fa
 tahalian (Stanford University); Ioannis Karamouzas (University of Californ
 ia Riverside); Victor Zordan (Roblox); and C. Karen Liu (Stanford Universi
 ty)\n\nLearning a control policy for a multi-phase, long-horizon task, suc
 h as basketball maneuvers, remains challenging for reinforcement learning 
 approaches due to the need for seamless policy composition and transitions
  between skills. A long-horizon task typically consists of distinct subtas
 ks with well-defined goals, separated by transitional subtasks with unclea
 r goals but critical to the success of the entire task. Existing methods l
 ike the mixture of experts and skill chaining struggle with tasks where in
 dividual policies do not share significant commonly explored states or lac
 k well-defined initial and terminal states between different phases. In th
 is paper, we introduce a novel policy integration framework to enable the 
 composition of drastically different motor skills in multi-phase long-hori
 zon tasks with ill-defined intermediate states. Based on that, we further 
 introduce a high-level soft router to enable seamless and robust transitio
 ns between the subtasks. We evaluate our framework on a set of fundamental
  basketball skills and challenging transitions. Policies trained by our ap
 proach can effectively control the simulated character to interact with th
 e ball and accomplish the long-horizon task specified by real-time user co
 mmands, without relying on ball trajectory references.\n\nRegistration Cat
 egory: Full Access, Full Access Supporter\n\nSession Chair: Jungdam Won (S
 eoul National University)\n\n
END:VEVENT
END:VCALENDAR
