BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Hong_Kong
X-LIC-LOCATION:Asia/Hong_Kong
BEGIN:STANDARD
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
TZNAME:HKT
DTSTART:19911015T033000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20251218T030657Z
LOCATION:Meeting Room S423+S424\, Level 4
DTSTART;TZID=Asia/Hong_Kong:20251218T133100
DTEND;TZID=Asia/Hong_Kong:20251218T134200
UID:siggraphasia_SIGGRAPH Asia 2025_sess155_papers_1683@linklings.com
SUMMARY:ConsistEdit: Highly Consistent and Precise Training-free Visual Ed
 iting
DESCRIPTION:Zixin Yin (Hong Kong University of Science and Technology); Li
 ng-Hao Chen (Tsinghua University, International Digital Economy Academy); 
 Lionel Ni (Hong Kong University of Science and Technology, Guangzhou; Hong
  Kong University of Science and Technology); and Xili Dai (Hong Kong Unive
 rsity of Science and Technology, Guangzhou)\n\nRecent advances in training
 -free attention control methods have enabled flexible and efficient text-g
 uided editing capabilities for existing image and video generation models.
  However, current approaches struggle to simultaneously deliver strong edi
 ting strength while preserving consistency with the source. For instance, 
 in color-editing tasks, they struggle to maintain structural consistency i
 n edited regions while preserving the rest intact. This limitation becomes
  particularly critical in multi-round and video editing, where visual erro
 rs can accumulate over time. Moreover, most existing methods enforce globa
 l consistency, which limits their ability to modify individual attributes 
 such as texture while preserving others, thereby hindering fine-grained ed
 iting. Recently, the architectural shift from U-Net to Multi-Modal Diffusi
 on Transformers (MM-DiT) has brought significant improvements in generativ
 e performance and introduced a novel mechanism for integrating text and vi
 sion modalities. These advancements pave the way for overcoming challenges
  that previous methods failed to resolve. Through an in-depth analysis of 
 MM-DiT, we identify three key insights into its attention mechanisms. Buil
 ding on these, we propose ConsistEdit, a novel attention control method sp
 ecifically tailored for MM-DiT. ConsistEdit incorporates vision-only atten
 tion control, mask-guided pre-attention fusion, and differentiated manipul
 ation of the query, key, and value tokens to produce consistent, prompt-al
 igned edits. Extensive experiments demonstrate that ConsistEdit achieves s
 tate-of-the-art performance across a wide range of image and video editing
  tasks, including both structure-consistent and structure-inconsistent sce
 narios. Unlike prior methods, it is the first approach to perform editing 
 across all inference steps and attention layers without handcraft, signifi
 cantly enhancing reliability and consistency, which enables robust multi-r
 ound and multi-region editing. Furthermore, it supports progressive adjust
 ment of structural consistency, enabling finer control. ConsistEdit repres
 ents a significant advancement in generative model editing and unlocks the
  full editing potential of MM-DiT architectures.\n\nRegistration Category:
  Full Access, Full Access Supporter\n\nSession Chair: Ali Mahdavi-Amiri (S
 imon Fraser University, MARZ VFX)\n\n
END:VEVENT
END:VCALENDAR
