PhD Proposal: Music as Structured Prior over Human Motion: Inference, Generation, and Control

Talk
Seong Yoo
Time: 
08.27.2026 10:00 to 11:30

One of the unique properties of humans is our musical behavior. Group dancing is a universal phenomenon across different cultures, ethnicities, and eras. We have a strong innate desire to move our bodies and synchronize with others through musical activity. Simple beats make people dance, and songs let us share emotion and transmit ancestral knowledge across generations. Understanding human behavior in musical contexts bridges the gap between our instinct and intelligence. This thesis argues that music is not merely a conditioning signal for models of human motion, but a structured prior over body kinematics, a temporally organized constraint on which movements are plausible at each instant. I make this claim at three levels: Inference, Generation, and Control.
Inference. Visual data captures low-frequency motion well but often fails to capture high-frequency detail due to occlusion, pixel saturation, and low sampling rates. Audio signals carry exactly that missing detail, yet they are inefficient for estimating low-frequency information. VioPose exploits this complementarity through hierarchical audiovisual inference, using audio to compensate for the limitations of visual data and recover 4D pose in violin playing scenarios.
Generation. Existing literature on dance generation relies primarily on audio to generate correlated motion. While this approach yields high fidelity results, it offers no room for choreographers to control the artistic output. To close this gap, STREAM composes conditioning signals as additive energies through energy-based cross attention. Generation becomes editable, composable, and controllable, and edits remain local and semantically addressable.
Control. I propose to move control from the output of the model into its generative process. Variance Path Diffusion (VPD) makes the noise schedule of a diffusion model an explicit random latent, a Gamma process variance path with a closed-form posterior. A Gamma-Dirichlet duality separates the total noise budget, which an observation determines from its allocation, which it does not, so the allocation becomes a design variable. Human motion supplies the groups over which that budget is spent, since limbs and successive phrases do not all require the same freedom, and music supplies the condition that allocates it, since music constrains the body at every instant in a way that text and class labels cannot. Conditioning the path on music leaves the denoiser output untouched, which makes this a control axis orthogonal by construction to conditioning and complementary to it.
Together these establish music as a structured prior that operates at every level of the generative stack. It serves as evidence for inference, as condition for generation, and as control over the generative process itself.