I am a final-year PhD candidate at the University of Maryland, College Park, advised by Abhinav Shrivastava. My research focuses on video reasoning through explicit motion modeling: understanding actions means capturing how the world moves, not just how it looks.

Earlier in my PhD, I developed trajectory-based video representations that model motion directly for few-shot action recognition. I now carry these ideas into systems that reason and act: post-training video-language models for agentic video understanding, and learning robotic manipulation policies from motion.

I am currently a research intern at NVIDIA, working on post-training Nemotron models for video reasoning, and was previously a student researcher at Google Research. Before graduate school, I was a senior data scientist at ParallelDots. I did my bachelors in Information Technology at NSIT, New Delhi.

News

MemStream was accepted at NeurIPS 2026.
VidParse was accepted at ECCV 2026.
Started as a research intern at NVIDIA, working on post-training Nemotron models for video reasoning and agentic video understanding.
Passed my preliminary exam and became a PhD candidate.
Trokens was accepted at ICCV 2025.
Earlier newsShow less
TATs was accepted at ECCV 2024.
One paper accepted at a CVPR 2024 workshop.
XINC was accepted at CVPR 2024.
Co-organizing the OBJ-DISC and FMDC challenges at the VPLOW workshop at CVPR 2023.
Co-organizing the second DNOW workshop at WACV 2023.
Co-organizing the OBJ-DISC challenge at the VPLOW workshop at CVPR 2022.
Accepted to intern with the Visual Dynamics team at Google Research for summer 2022.
Co-organizing the Dealing with Novelty in Open Worlds (DNOW) workshop at WACV 2022.
Joined the PhD program at UMD.
Finished my MS in Computer Science at UMD.

Selected publications

Older researchHide older research