Xiaohang Tang

Hi, and welcome — I'm Xiaohang. I'm currently a Research Scientist Intern at Google DeepMind, and a PhD candidate at UCL advised by Ilija Bogunovic. I have also been fortunate to work closely with Yaodong Yang. Before the PhD, I obtained my master's degree in Data Science at UCL and a bachelor's degree in Math at Sun Yat-sen University (SYSU). My research aims to build self-improving systems — agents and models for decision-making, reasoning, and generation — adapting to a wide range of important problems, from multi-agent systems and robust control to LLM alignment and post-training.

My research spectrum: from games to post-training

Self-improvement act · evaluate · update Games & multi-agent systems self-play improving against its own best responses Robust Decision-Making imagined self-play planning against imagined adversary LLM post-training self-play alignment optimizing preferences against its older self Diffusion LLM post-training self-distillation distilling its own reward guided denoiser
Selected publications
arXivDiffusion LLM
GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models Xiaohang Tang, Keyue Jiang, Che Liu, Qifang Zhao, Xiaoxiao Xu, Sangwoong Yoon, Ilija Bogunovic · arXiv 2605.29398
ICML 2026Alignment
RSPO: Regularized Self-Play Alignment of Large Language Models Xiaohang Tang, Sangwoong Yoon, Seongho Son, Huizhuo Yuan, Quanquan Gu, Ilija Bogunovic
ICLR 2026Diffusion LLM
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models Xiaohang Tang, Rares Dolga, Sangwoong Yoon, Ilija Bogunovic
ICLR 2026Alignment
Robust Multi-Objective Controlled Decoding of Large Language Models Seongho Son, William Bankes, Sangwoong Yoon, Shyam Sundhar Ramesh, Xiaohang Tang, Ilija Bogunovic
NeurIPS 2024Offline RL
Adversarially Robust Decision Transformer Xiaohang Tang, Afonso Marques, Parameswaran Kamalaruban, Ilija Bogunovic · invited talk, DeepMind Game Theory Group
NeurIPS 2025Reliable ML Workshop
ICML 2023Game theory
Regret-Minimizing Double Oracle for Extensive-Form Games Xiaohang Tang, Le Cong Dinh, Stephen Marcus McAleer, Yaodong Yang
JMLRGame theory
Sample-Efficient Regret-Minimizing Double Oracle for Extensive-Form Games Xiaohang Tang, Chiyuan Wang, Chengdong Ma, Ilija Bogunovic, Stephen McAleer, Yaodong Yang · accepted, minor revision
IJCAI 2021RL theory
Average-Reward Reinforcement Learning with Trust Region Methods Xiaoteng Ma, Xiaohang Tang, Li Xia, Jun Yang, Qianchuan Zhao