PhD Defense: Localize, Edit, Unlearn: Structural Interventions and Data-Level Control in Foundation Models
IRB-4107 https://umd.zoom.us/j/9719759070
As foundation models continue to transform the landscape of artificial intelligence, demonstrating impressive capabilities across vision and language tasks, the need for interpretability, transparency, and control becomes increasingly critical. We develop methods to understand AI models by studying their representation spaces, the role of their internal architectural components, and the role of training data in their test-time behavior. Our work spans three main areas: interpretability of vision models, where we propose methods for mapping internal representations to human-understandable concepts and explaining failure modes; knowledge localization and editing in text-to-image generative models, where we propose techniques to identify and modify the layers responsible for specific concepts; and understanding the impact of data on models' training trajectories through the problem of machine unlearning, where we introduce benchmarks for data-level unlearning and new algorithms that improve unlearning efficacy, particularly through the use of intermediate checkpoints. Through these efforts, we contribute to building more transparent, controllable, and adaptable AI systems.