PhD Defense: Embodied Action Understanding

Talk
Eadom Dessalene
Time: 
09.03.2026 11:00 to 12:30
Location: 

IRB-4105

Much of what we learn to do, we first learn by interpreting the actions of others. Despite the central role of action understanding in perception, much of the literature on action understanding has approached this problem by copying paradigms originally developed for interpreting scenes and objects towards the interpretation of actions. This is consistent with the classical sense–think–act view of cognition, in which perception, cognition, and action are treated as separate processes. Embodied cognition instead suggests that perception and action are tightly coupled and rely on shared representations. I refer to action understanding built around this idea as Embodied Action Understanding.The first part of this dissertation is centered on contact. Contact Anticipation Maps predict where and when contact will occur, while Next Active Object segmentation identifies the next object of interaction. Egocentric Object Manipulation Graphs organize these predictions into hand-object interaction sequences for action anticipation, achieving 1st place in the EPIC-Kitchens Action Anticipation Challenge.Therbligs introduce a vocabulary of reusable motion primitives for video understanding, allowing actions to be represented as compositions of sub-actions. LEAP develops this idea further by leveraging multiple modalities to produce action programs consisting of perceptual functions, motoric functions, and control flow. In this representation, perceptual and motoric structure are no longer modeled separately, but are bound together within a single representation.We next enhance visual representations with signals that reveal the physical structure of action. We release FEEL, a dataset pairing egocentric video with force sensing. FEEL is the first video learning approach to generate contact supervision without any need for manual annotation. Using these contact pseudo-labels we train models to predict when and where contact occurs from video alone.MotorSense goes further by pairing video with bimanual EMG, providing a dense, occlusion-free measurement of the muscle activity underlying hand action. We then learn cross-modal video–EMG alignment, improving action recognition and 3D hand-object reconstruction. CoVE extends this idea by jointly using vision and EMG motor signals to recover hands, objects, and their physical contacts in 3D. We achieve state-of-the-art 3D hand reconstruction, particularly under hand-object occlusion.