Games & multi-agent RL
Where self-improvement is literal: iteratively expanding a strategy pool with best responses to its equilibrium, converging to global Nash equilibrium sample-efficiently.
Hi, and welcome — I'm Xiaohang. I'm currently a Research Scientist Intern at Google DeepMind, and a PhD candidate at UCL advised by Ilija Bogunovic. I have also been fortunate to work closely with Yaodong Yang. Before the PhD, I obtained my master's degree in Data Science at UCL and a bachelor's degree in Math at Sun Yat-sen University (SYSU). My research aims to build self-improving systems — agents and models for decision-making, reasoning, and generation — adapting to a wide range of important problems, from multi-agent systems and robust control to LLM alignment and post-training.
My research spectrum: from games to post-training
Where self-improvement is literal: iteratively expanding a strategy pool with best responses to its equilibrium, converging to global Nash equilibrium sample-efficiently.
Worst-case-aware agents for adversarial environments: sequence models that condition on returns against imagined adversaries.
Self-play carried into alignment: a policy improved against itself to approximate the Nash equilibrium of the preference optimization game.
The newest instantiation: post-training a denoiser on signals it already generates, without unstable likelihood-ratio estimates.