I am a Member of Technical Staff at Microsoft AI (MAI), working on the final climb post-training of the MAI-Thinking Vision Language Model. Previously, I was a Senior Research Scientist at Microsoft Research AI Frontiers, where I was a core member of the Phi-series models — contributing to pre-training with compact world models and building out the RL post-training stack.

I completed my PhD in Machine Learning at Mila – Quebec AI Institute and McGill University, supervised by Doina Precup and Joelle Pineau, with Yoshua Bengio and John Langford as co-advisors. My thesis was on Latent State Discovery for Robust, Reusable and Reliable Deep RL.

Research Interests: I build reasoning models and systems for math, coding and vision tasks, with a focus on long horizon agentic AI, world dynamics models, latent representation learning, goal-directed planning and exploration capabilities. My recent work focuses on Language Model pre-training, thinking in the latent space with faster rollout capabilities; self-verification, curriculum learning, scalable RL post-training recipe and improving long horizon reasoning. I care about reliable and reproducible AI, ensuring robustness of large scale AI systems.

Industry Research Experience

  • Microsoft AI (MAI), New York Member of Technical Staff — RL Post-Training | Final Climb Vision-Language Models Collaborators: Ali Farhadi, John Langford, Ranjay Krishna
  • Microsoft Research AI Frontiers, New York Senior Research Scientist — RL Post-Training | Pre-Training | World Models Collaborators: John Langford, Alex Lamb, Kwangjun Ahn, Tim Pearce, Manan Tomar
  • SDAIA, HUMAIN, Riyadh Senior Research Scientist — Post-Training and Safety Alignment Collaborators: M Saiful Bari, Sheikh Jubair, Ehsan Hoque, Pedro Moreno
  • DreamFold AI, Montreal Research Scientist — RL & Flow Matching for Protein Structure Design Collaborators: Alex Tong, Joey Bose, Jarrid Rector-Brooks
  • Microsoft Research, Montreal Student Researcher — Factorized Latent Representations in RL Collaborators: R. Tachet des Combes, R. Laroche, H. van Seijen, A. Lamb, J. Langford, D. Misra, A. Krishnamurthy
  • Apple MLR, San Jose Research Scientist Intern — Representation Learning for World Models in RL Collaborators: Devon Hjelm, Samy Bengio
  • Microsoft Research, New York Research Scientist Intern — Provable RL for Latent State Discovery Collaborators: Alex Lamb, John Langford, Akshay Krishnamurthy, Dylan Foster, Yonathan Efroni, Dipendra Misra

Academic Research Experience

  • Microsoft Research, Montreal Student Researcher Intern — Reproducibility in Deep RL, Active Imitation Learning Collaborators: Phillip Bachman, Alessandro Sordoni, John Langford
  • Caltech and NASA JPL Improved State Estimation for a Resilient Spacecraft Executive Supervisors: R. Murray, C. McGhan (SURF Program)
  • University of Cambridge Active Learning for Images using Bayesian Deep Learning Uncertainty Supervisors: Z. Ghahramani, Y. Gal
  • UCL Improving Convergence of Deterministic Policy Gradient Algorithms in RL Supervisors: J. Shawe-Taylor, G. Lever; and D. Silver
  • Johns Hopkins University Cost-Sensitive Decision Tree of Classifiers for Medical Data Supervisor: Suchi Saria (SRE Program)

Education

  • McGill University & Mila – Quebec AI Institute, Montreal PhD in Machine Learning Supervisors: Doina Precup, Joelle Pineau; Co-Advisors: Yoshua Bengio, John Langford. Thesis: Latent State Discovery for Robust, Reusable and Reliable Deep RL.
  • University of Cambridge, St John's College MPhil, Machine Learning, Speech & Language Technology Supervisors: Zoubin Ghahramani, Yarin Gal. Thesis: Active Learning with Image Data using Bayesian Uncertainty in Deep Learning.
  • University College London BEng, Electronic & Electrical Engineering (ML Major) Supervisors: John Shawe-Taylor, Guy Lever; Co-Advisor: David Silver. Thesis: Analysing Convergence of Deep Policy Optimization in RL; Deterministic Policy Gradient Methods.

Selected Publications

For a complete list of publications, see my Google Scholar profile.

Journals and Patents

  • Alex Lamb*, Riashat Islam*, et al. "Guaranteed Discovery of Controllable Latent States." Patent: Under Submission with Microsoft Research
  • Vincent Francois-Lavet, Peter Henderson, Riashat Islam, Marc G. Bellemare, Joelle Pineau. "An Introduction to Deep Reinforcement Learning." Foundations and Trends in Machine Learning, 2018