BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Hong_Kong
X-LIC-LOCATION:Asia/Hong_Kong
BEGIN:STANDARD
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
TZNAME:HKT
DTSTART:19911015T033000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20251218T030657Z
LOCATION:Meeting Room S423+S424\, Level 4
DTSTART;TZID=Asia/Hong_Kong:20251218T131000
DTEND;TZID=Asia/Hong_Kong:20251218T132000
UID:siggraphasia_SIGGRAPH Asia 2025_sess155_papers_1039@linklings.com
SUMMARY:In-Context Brush: Zero-shot Customized Subject Insertion with Cont
 ext-Aware Latent Space Manipulation
DESCRIPTION:Yu Xu, Fan Tang, You Wu, and Lin Gao (Institute of Computing T
 echnology, Chinese Academy of Sciences); Oliver Deussen (University of Kon
 stanz); Hongbin Yan (University of Chinese Academy of Sciences); Jintao Li
  and Juan Cao (Institute of Computing Technology, Chinese Academy of Scien
 ces); and Tong-Yee Lee (National Cheng-Kung University)\n\nRecent advances
  in diffusion models have enhanced multimodal-guided visual generation, en
 abling customized subject insertion that seamlessly "brushes" user-specifi
 ed objects into a given image guided by textual prompts. However, existing
  methods often struggle to insert customized subjects with high fidelity a
 nd align results with the user's intent through textual prompts.  In this 
 work, we propose "In-Context Brush", a zero-shot framework for customized 
 subject insertion by reformulating the task within the paradigm of in-cont
 ext learning.  Without loss of generality, we formulate the object image a
 nd the textual prompts as cross-modal demonstrations, and the target image
  with the masked region as the query. The goal is to inpaint the target im
 age with the subject aligning textual prompts without model tuning. Buildi
 ng upon a pretrained MMDiT-based inpainting network, we perform test-time 
 enhancement via dual-level latent space manipulation:  intra-head latent f
 eature shifting within each attention head that dynamically shifts attenti
 on outputs to reflect the desired subject semantics and inter-head attenti
 on reweighting across different heads that amplifies prompt controllabilit
 y through differential attention prioritization. Extensive experiments and
  applications demonstrate that our approach achieves superior identity pre
 servation, text alignment, and image quality compared to existing state-of
 -the-art methods, without requiring dedicated training or additional data 
 collection.\n\nRegistration Category: Full Access, Full Access Supporter\n
 \nSession Chair: Ali Mahdavi-Amiri (Simon Fraser University, MARZ VFX)\n\n
END:VEVENT
END:VCALENDAR
