CV
Education, research, industry experience, publications, projects, patents, and technical skills.
Contact Information
| Name | Abhijeet Sinha |
| Professional Title | PhD Researcher in Machine Learning |
| abhijeet@nus.edu.sg |
Professional Summary
Researcher in reinforcement learning, generative modeling, LLM post-training and alignment, diversity, creativity, and interpretable AI.
Experience
-
2024 - present Singapore
Research Assistant
Artificial Scientific Intelligence Lab, National University of Singapore
- Develop research on reward-proportional reinforcement learning and outcome-level mode collapse, resulting in a first-author ICML 2026 publication and an AAAI 2027 submission.
- Designed density-aware reward scaling for continuous outcome spaces and integrated it with GRPO for multimodal navigation and language-model-based text-to-image prompt adaptation.
- Developed methods for LLM post-training and alignment with emphasis on reward design, diverse generation, and reasoning-oriented evaluation.
- Developed an RL agent that edits discrete VQ-VAE latent codes using classifier-confidence and discriminator-realism feedback.
-
2023 - 2023 San Francisco, CA, USA
Machine Learning Engineer
VAO Labs
- Developed LayoutLM and OCR document-understanding pipelines for structured information extraction.
- Fine-tuned T5-family models for document question answering and answer retrieval.
- Built Llama-based applications for querying large enterprise datasets with natural language.
-
2022 - 2023 Chennai, India
Research Assistant
Computational Neuroscience Lab, IIT Madras
- Designed a brain-inspired visual attention model that learned saccadic scanning policies with Q-learning and recurrent neural networks.
- Developed computer-vision and reinforcement-learning systems for neurological healthcare applications.
- Invented A System for Monitoring Parkinson’s Disease, granted as an Indian patent to IIT Madras.
-
2021 - 2021 Mumbai, India
Software Developer
IIFL
- Developed full-stack software using ASP.NET, SQL Server, Angular, JavaScript, and jQuery.
- Integrated, tested, and documented changes using TFS and Azure in an agile workflow.
Education
-
2024 - present Singapore
PhD
National University of Singapore
Machine Learning
- Thesis: Diversity and Creativity in Generative AI
- Supervisor: Dr. Dianbo Liu
-
2016 - 2021 Chennai, India
Dual Degree (Bachelor's and Master's)
Indian Institute of Technology Madras
Biotechnology
Publications
-
2026 Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability Scaling
International Conference on Machine Learning (ICML)
Explains outcome-level mode collapse induced by expected-return objectives and introduces inverse-probability scaling to recover diverse high-reward outcomes.
-
2026 Density-Aware Reward Scaling: A Reinforcement Learning Framework for Reward-Proportional Sampling
Submitted to AAAI 2027
Introduces DARS, a critic-free method that scales rewards by local outcome density while retaining the standard policy-gradient objective.
-
2026 How Does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
Spotlight, CVPR XAI4CV Workshop
Introduces hierarchical sparse features for coarse-to-fine interpretation of physical-plausibility failures across generative models.
-
2023 Brain-Inspired Attention Model for Object Counting
International Conference on Neural Information Processing (ICONIP)
Developed a Q-learning-based sequential attention model with recurrent visual glimpses, achieving 92.1% object-counting accuracy.
Patents
Projects
-
2024 - present Editing Discrete Latent Variables with a Reinforcement Learning Agent
An RL framework that edits discrete VQ-VAE codes toward a target class while minimizing codebook changes.
- Combines classifier-confidence and discriminator-realism rewards for controlled latent-space editing.
- Produces multi-step trajectories that reveal interpretable, task-dependent latent navigation.