Abhijeet Sinha

PhD researcher at the National University of Singapore working on reinforcement learning and generative AI.

profile_pic.png

Abhijeet Sinha

National University of Singapore

Singapore

I am a PhD researcher in Machine Learning at the National University of Singapore, where I work in the Artificial Scientific Intelligence Lab under the supervision of Dr. Dianbo Liu. My research focuses on reinforcement learning, generative modeling, LLM post-training and alignment, diversity, creativity, and interpretable AI.

My current work studies outcome-level mode collapse in reinforcement learning and reward-proportional sampling. I develop methods that encourage diverse, high-reward generation while preserving alignment and task performance. I also work on interpretable generative systems, including reinforcement-learning agents that edit discrete latent representations.

Previously, I completed a dual Bachelor’s and Master’s degree in Biotechnology at IIT Madras and worked on brain-inspired visual attention, neurological healthcare applications, document understanding, and enterprise AI systems.

selected publications

  1. ICML
    Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability Scaling
    Abhijeet Sinha, Sundari Elango, and Dianbo Liu
    In International Conference on Machine Learning, 2026
  2. AAAI
    Density-Aware Reward Scaling: A Reinforcement Learning Framework for Reward-Proportional Sampling
    Abhijeet Sinha, Ria Shekhawat, and Dianbo Liu
    Submitted to AAAI 2027
  3. CVPR
    How Does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
    Yiming Tang, Abhijeet Sinha, and Dianbo Liu
    In CVPR Workshop on Explainable AI for Computer VisionSpotlight, CVPR XAI4CV Workshop 2026 , 2026
  4. ICONIP
    Brain-Inspired Attention Model for Object Counting
    Abhijeet Sinha, Sweta Kumari, and V. Srinivasa Chakravarthy
    In International Conference on Neural Information Processing, 2023