I study how large language models perform multi-step reasoning, and how training and post-training methods can improve their reliability, efficiency, and scalability.
My work focuses on the post-training stack for LLMs: supervised fine-tuning (SFT), preference optimization, reinforcement learning methods such as RLVR and agentic RL, and inference-time compute strategies that improve reasoning without requiring larger models — including RL environments and sandboxed setups for training and evaluating tool-using and coding agents.
I'm also interested in the interpretability of reasoning models: understanding the internal mechanisms behind multi-step reasoning, and diagnosing failures such as shortcut reasoning, reward hacking, and unfaithful chain-of-thought.
Currently researching mechanistic interpretability of reasoning models at IISc, and building and open-sourcing reasoning-focused post-training pipelines and evaluation systems.



