PhD Proposal: Segmentation Before Structure: Real-Time Moving-Object Segmentation and Egomotion Estimation with Event Cameras
IRB-3137
Interpreting a dynamic world from a moving camera requires answering two entangled questions: how is the observer moving, and what else is moving independently? Classical structure-from-motion, whether frame-based or event-based, typically resolves this chicken-and-egg problem in a fixed order: estimate egomotion first under a static-scene assumption, then explain away independently moving objects as outliers of that model. This ordering is fragile exactly where autonomous systems need vision the most—in cluttered, fast, dynamic scenes—and it has been inherited, largely unexamined, by the event-camera literature. This thesis proposes to reverse the pipeline. Building on the observation, both biological and geometric, that independent motion is detectable before and without egomotion, we develop four tightly connected components: a universal motion signal: per-event normal flow with calibrated confidence, strictly local so that it transfers across scenes and sensors; moving-object segmentation solved first, from events alone, without egomotion priors; egomotion estimation as a learned minimizer that keeps hard geometric constraints inside the loop, learning how to optimize rather than what scenes look like; and temporal integration that lifts instantaneous estimates to multi-frame odometry through event-surface, event-rate, and inertial constraints.The completed work establishes the pivot of this pipeline, segmentation first, through three projects. A graph-convolutional network operating on raw event clouds demonstrated that extended temporal context, together with local geometry rather than semantic appearance, resolves motion segmentation; a classical successor isolated moving objects through residual analysis of a normal-flow field, without a prior egomotion estimate; and our border-ownership network completed the argument: trained entirely in simulation, it segments moving objects faster than real time, transfers zero-shot to real cameras, and matches or surpasses methods supervised with real data.Ongoing work supplies the stages on either side of segmentation: two complementary normal-flow estimators, one dense and real-time, the other asynchronous and geometry-exact, which classify events by reliability as a byproduct; and a physics-in-the-loop recurrent egomotion estimator that recovers the translation direction in real time without pose supervision. The proposed research completes the pipeline by closing the real-time gap for asynchronous estimation, extending segmentation-aware egomotion to all six degrees of freedom, and building the temporal-integration layer: a tightly coupled event-inertial backend that fuses event-surfaces, event-rate, and inertial constraints to anchor rotation, gravity, and metric scale. The result will be a principled, efficient, biologically motivated alternative to both end-to-end learning and classical optimization for event-based dynamic scene understanding.