PhD Proposal: Response Behavior in Large Language Models: Data Selection, Gradient Analysis, and Reasoning Evaluation
ttps://umd.zoom.us/j/2762491074?omn=91264661808
Post-training shapes how large language models follow instructions, learn from training data, and reason through difficult problems. This proposal investigates the relationship between post-training data and model behavior at three levels: data selection, gradient analysis, and reasoning evaluation. Model-aware data selection identifies examples that provide useful learning signals and enables efficient filtering with smaller proxy models. Gradient analysis reveals how reasoning detail, response relevance, and data quality produce distinct patterns in the magnitude and structure of model updates. Behavioral evaluation further exposes failure modes of extended reasoning, including overthinking on underspecified problems, and represents long reasoning traces as sequences of functional episodes. Building on these findings, the proposed research uses episode-level behavioral signals to select reasoning data and construct training preferences, testing whether behavior-aware post-training can improve the organization of reasoning while preserving correctness and efficiency.