Wenhao Yang @ LAMDA, NJU-AI

ywh2024.JPG 
(Filmed at Gosau Lake in Vienna; by Ling-Hao CHEN.)

杨文昊
Wenhao Yang
Ph.D. student, LAMDA Group
Email: yangwh@lamda.nju.edu.cn
Google Scholar
Supervisor: Professor Lijun Zhang
Major: Computer Science and Technology
School of Artificial Intelligence
National Key Laboratory for Novel Software Technology
Nanjing University, Nanjing 210023, China

I am on the 2026–2027 job market and expect to graduate in June 2027. Please feel free to reach out!

Biography

Currently I am a final year Ph.D. student of School of Artificial Intelligence in Nanjing University and a member of LAMDA Group, led by professor Zhi-Hua Zhou.

I received my B.Sc. degree from School of Electronic Engineering, Xidian University in June 2022. In the same year, I was admitted to pursue Ph.D. degree in Nanjing University without entrance examination.

Research Interests

My current research focuses on the post-training of Large Language Models (LLMs), dedicated to developing scalable reinforcement learning algorithms to advance their capabilities, particularly in coding. Previously, I worked on online learning and optimization.

Internship Experience

Qwen Team, Alibaba Group, Beijing, Jul. 2026 - Present.

Project: Improving the performance of reinforcement learning algorithms in hybrid training scenarios.

Multimodal RL, Tencent Hunyuan (青云计划), Beijing, Apr. 2026 - Jul. 2026.

Project: Built a Claude Code-based creative game agent with long-horizon multi-agent workflows for game design, asset generation, Phaser development, and automated visual verification and debugging.

AI Business, Alibaba Group, Hangzhou, Sep. 2025 - Apr. 2026.

Project: Developed a post-training pipeline for Think with Images capabilities in MLLMs using verifiable tool-use data, cold-start SFT, and RL to improve multi-step visual reasoning and tool planning.

AliExpress, Alibaba Group, Hangzhou, Jul. 2023 - Oct. 2023.

Project: Developed region-aware modeling for cross-domain recommendation across multiple countries and markets, leading to publications at WWW 2024 and AAAI 2025.

Preprints

  1. SceneActBench: Can Agents Act on the 3D Scenes They See? [PDF][Website][GitHub][HuggingFace]
    Y. Zhao*, X. Zhou*, W. Yang*, J. Tang*, P. Jian*, H. Yao*, J. Yao*, H. Lin, Z. Chen, W. Lyu, J. Ma, X. Wang, W. Zhu, and T. Pang (* indicates equal contribution)

  2. Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images [arXiv][HuggingFace]
    W. Yang, Y. Xia, J. Huang, S. Lu, Q. Chen, Z. Xu, W. Luo, K. Zhang, Y. Wan, and L. Zhang

  3. Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization [arXiv]
    W. Yang, Y. Xia, J. Huang, S. Lu, Q. Chen, Z. Xu, W. Luo, K. Zhang, Y. Zhou, X. Xia, Y. Wan, L. Zhang, and Tat-seng Chua

  4. Dual Adaptivity: Universal Algorithms for Minimizing the Adaptive Regret of Convex Functions [arXiv]
    L. Zhang, W. Yang, G. Wang, W, Jiang, and Z.-H. Zhou (♣ indicates my supervisor)

  5. Improved Analysis for Sign-based Methods with Momentum Updates [arXiv]
    W. Jiang, D. Yu, S. Yang, W. Yang, and L. Zhang

Publications

  1. Logarithmic Switching Regret for Online Convex Optimization
    W. Yang, Y. Wang, Y. Wan, and L. Zhang
    In Proceedings of the 43rd International Conference on Machine Learning (ICML 2026), to appear, 2026.

  2. Decentralized Online Convex Optimization with Efficient Communication: Improved Algorithm and Lower Bounds [arXiv]
    S. Yang, W. Yang, W. Jiang, and L. Zhang
    In Proceedings of the 43rd International Conference on Machine Learning (ICML 2026), to appear, 2026.

  3. Convergence Analysis of the Lion Optimizer in Centralized and Distributed Settings
    W. Jiang, M. Xu, W. Yang, Y. Wang, Z. Li, and L. Zhang
    In Proceedings of the 43rd International Conference on Machine Learning (ICML 2026), to appear, 2026.

  4. Discounted Online Convex Optimization: Uniform Regret Across a Continuous Interval [arXiv]
    W. Yang, S. Yang, and L. Zhang
    The 14th International Conference on Learning Representations (ICLR 2026), to appear, 2026.

  5. Smoothed Online Convex Optimization with Delayed Feedback
    S. Yang, W. Yang, W. Jiang, Y. Wan, and L. Zhang
    In Proceedings of the 34th International Joint Conference on Artificial Intelligence (IJCAI 2025), to appear, 2025.

  6. Revisiting Differentially Private Algorithms for Decentralized Online Learning
    X. Wang, W. Yang, C. Yao, M. Song, and Y. Wan
    In Proceedings of the 42nd International Conference on Machine Learning (ICML 2025), to appear, 2025.

  7. Towards Unbiased Information Extraction and Adaptation in Cross-Domain Recommendation [PDF, Bibtex]
    Y. Wang, Y. Jian, W. Yang, S. Lu, L. Shen, B. Wang, X. Zeng, and L. Zhang
    In Proceedings of the 39th AAAI Conference on Artificial Intelligence (AAAI 2025), pages 12757--12765, 2025.

  8. Universal Online Convex Optimization with 1 Projection per Round [PDF, arXiv, Bibtex]
    W. Yang, Y. Wang, P. Zhao, and L. Zhang
    In Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pages 31438 -- 31472, 2024.

  9. Online Composite Optimization Between Stochastic and Adversarial Environments [PDF, Bibtex]
    Y. Wang, S. Chen, W. Jiang, W. Yang, Y. Wan and L. Zhang
    In Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pages 94808 -- 94850, 2024.

  10. Efficient Sign-Based Optimization: Accelerating Convergence via Variance Reduction [PDF, arXiv, Bibtex]
    W. Jiang, S. Yang, W. Yang, and L. Zhang
    In Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pages 33891 -- 33932, 2024.

  11. Small-loss Adaptive Regret for Online Convex Optimization [PDF, Bibtex]
    W. Yang, W. Jiang, Y. Wang, P. Yang, Y. Hu, and L. Zhang
    In Proceedings of the 41st International Conference on Machine Learning (ICML 2024), pages 56156--56195, 2024.

  12. Projection-Free Variance Reduction Methods for Stochastic Constrained Multi-Level Compositional Optimization [PDF, Bibtex]
    W. Jiang, S. Yang, W. Yang, Y. Wang, Y. Wan, and L. Zhang
    In Proceedings of the 41st International Conference on Machine Learning (ICML 2024), pages 21962--21987, 2024.

  13. Not All Embeddings are Created Equal: Towards Robust Cross-domain Recommendation via Contrastive Learning [PDF, Bibtex]
    W. Yang, Y. Jian, Y. Wang, S. Lu, L. Shen, B. Wang, H. Tang, and L. Zhang
    In Proceedings of the ACM Web Conference 2024 (WWW 2024), pages 3195--3206, 2024.

  14. Non-stationary Projection-Free Online Learning with Dynamic and Adaptive Regret Guarantees [PDF, arXiv, Bibtex]
    Y. Wang, W. Yang, W. Jiang, S. Lu, B. Wang, H. Tang, Y. Wan, and L. Zhang
    In Proceedings of the 38th AAAI Conference on Artificial Intelligence (AAAI 2024), pages 15671--15679, 2024.

  15. Pluralistic Image Completion with Gaussian Mixture Models [PDF, Supplementary, Code, Bibtex]
    X. Xia*, W. Yang*, J. Ren, Y. Li, Y. Zhan, B. Han and T. Liu (* indicates equal contribution)
    In Advances in Neural Information Processing Systems 35 (NeurIPS 2022), pages 24087--24100, 2022.

Awards & Honors

Academic Service

Teaching Assistant

Correspondence