PhD Proposal: Toward Interpretable and Controllable Image Generation and Editing

Talk
Arman Zarei
Time: 
07.30.2026 14:00 to 15:30

As image generation and editing models become increasingly capable, a central challenge is ensuring that their outputs faithfully reflect user intent while making their internal behavior understandable and amenable to precise intervention. This work develops methods to improve the interpretability and controllability of these models, with a focus on compositionality, knowledge localization, and fine-grained controllability.For compositional generation, we first identify limitations in CLIP-based text conditioning that impair attribute–object binding and introduce lightweight projection-based interventions that improve compositional fidelity while preserving image quality. We then develop an agentic framework that constructs compositionally contrasted training examples and uses distance-aware preference optimization to improve complex prompt following.We next study how semantic knowledge is represented within diffusion transformers. By tracing textual information across transformer blocks, we localize components responsible for specific visual concepts and validate their causal roles through targeted interventions. These findings enable efficient, selective personalization and knowledge unlearning while preserving unrelated capabilities.Finally, for instruction-based image editing, we develop a framework that disentangles multi-part instructions into continuous controls, allowing users to precisely adjust the strength of each individual edit and iteratively reach their desired output—an essential capability for interactive editing interfaces and fine-grained user control.Together, these works connect mechanistic understanding, targeted model adaptation, and intuitive user control to advance image generation and editing models that are more faithful, interpretable, efficient, and controllable.