Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
Semantic sampling of query points for tracking, with explicit motion modeling, improves few-shot video action recognition.
I am a final-year PhD candidate at the University of Maryland, College Park, advised by Abhinav Shrivastava. My research focuses on video reasoning through explicit motion modeling: understanding actions means capturing how the world moves, not just how it looks.
Earlier in my PhD, I developed trajectory-based video representations that model motion directly for few-shot action recognition. I now carry these ideas into systems that reason and act: post-training video-language models for agentic video understanding, and learning robotic manipulation policies from motion.
I am currently a research intern at NVIDIA, working on post-training Nemotron models for video reasoning, and was previously a student researcher at Google Research. Before graduate school, I was a senior data scientist at ParallelDots. I did my bachelors in Information Technology at NSIT, New Delhi.
Semantic sampling of query points for tracking, with explicit motion modeling, improves few-shot video action recognition.
A dynamic KV-cache memory with adaptive token selection and training-free retrieval mixture-of-experts enables streaming video understanding at high token budgets.
Training-free, online parsing of egocentric procedural videos that tracks which step of a task is happening, using manipulation-anchored features and task-graph constraints.
Point tracking and DINO features combine into trajectory-aligned tokens that capture motion and semantics for few-shot action recognition.