BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Hong_Kong
X-LIC-LOCATION:Asia/Hong_Kong
BEGIN:STANDARD
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
TZNAME:HKT
DTSTART:19911015T033000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20251218T030655Z
LOCATION:Meeting Room S423+S424\, Level 4
DTSTART;TZID=Asia/Hong_Kong:20251218T135300
DTEND;TZID=Asia/Hong_Kong:20251218T140400
UID:siggraphasia_SIGGRAPH Asia 2025_sess155_papers_1976@linklings.com
SUMMARY:DreamO: A Unified Framework for Image Customization
DESCRIPTION:Chong Mou (Bytedance, Peking University); Yanze Wu, Wenxu Wu, 
 Zinan Guo, Pengze Zhang, Yufeng Cheng, Yiming Luo, Fei Ding, Shiwen Zhang,
  Xinghui Li, Mengtian Li, Mingcong Liu, Yunsheng Jiang, Shaojin Wu, and So
 ngtao Zhao (Bytedance); Jian Zhang (Peking University); and Qian He and Xi
 nglong Wu (Bytedance)\n\nRecently, extensive research on image customizati
 on (e.g., identity, subject, style, background, etc.) demonstrates strong 
 customization capabilities in large-scale generative models. However, most
  approaches are designed for specific tasks, restricting their generalizab
 ility to combine different types of condition. Developing a unified framew
 ork for image customization remains an open challenge. In this paper, we p
 resent DreamO, an image customization framework designed to support a wide
  range of tasks while facilitating seamless integration of multiple condit
 ions. Specifically, DreamO utilizes a diffusion transformer (DiT) framewor
 k to uniformly process input of different types. During training, we const
 ruct a large-scale training dataset that includes various customization ta
 sks, and we introduce a feature routing constraint to facilitate the preci
 se querying of relevant information from reference images. Additionally, w
 e design a placeholder strategy that associates specific placeholders with
  conditions at particular positions, enabling control over the placement o
 f conditions in the generated results. Moreover, we employ a progressive t
 raining strategy consisting of three stages: an initial stage focused on s
 imple tasks with limited data to establish baseline consistency, a full-sc
 ale training stage to comprehensively enhance the customization capabiliti
 es, and a final quality alignment stage to correct quality biases introduc
 ed by low-quality data. Extensive experiments demonstrate that the propose
 d DreamO can effectively perform various image customization tasks with hi
 gh quality and flexibly integrate different types of control conditions.\n
 \nRegistration Category: Full Access, Full Access Supporter\n\nSession Cha
 ir: Ali Mahdavi-Amiri (Simon Fraser University, MARZ VFX)\n\n
END:VEVENT
END:VCALENDAR
