Han Zheng

Han Zheng

I'm a Ph.D. student at MIT, affiliated with the Laboratory for Information and Decision Systems (LIDS), advised by Prof. Cathy Wu. Previously I was at the University of Michigan working with the late Prof. Huei Peng, and I hold dual B.S./B.E. degrees from UIUC and Zhejiang University.

Research

I'm interested in scalable planning and decision-making for multi-agent systems, combining reinforcement learning with classical search and combinatorial optimization. More recently, I've been exploring self-evolving agents that improve through interaction, accumulated experience, and open-ended discovery.

H. Zheng, Y. Ma, K. Gunasekaran, B. Balaji, Z. Du, S. Vitaladevuni, C. Wu
arXiv, 2026
Introduced METIS, a self-evolving framework that internalizes curriculum judgment as a native LLM capability for RFT, letting the policy learn what to learn next and accelerating convergence by up to 67%.
Z. Qi, H. Su, A. Qu, C. Wang, Y. Yao, H. Zheng, K. Chattopadhyay, G. Xu, Z. Wang, W. Ye, V. J. Reddi, J. Li, P. P. Liang, H. Lakkaraju, S. Kakade, Y. Du
arXiv, 2026
A decentralized multi-agent system inspired by Hayek's economics, where agents compete via auctions and evolve through economic selection. Simple market signals turn populations of weak partial agents into collective intelligence that outperforms stronger monolithic baselines across five agentic tasks.
A. Qu*, H. Zheng*, Z. Zhou*, Y. Yan, Y. Tang, S. Y. Ong, F. Hong, K. Zhou, C. Jiang, M. Kong, J. Zhu, X. Jiang, S. Li, C. Wu, B. K. H. Low, J. Zhao, P. P. Liang
COLM, 2026
First framework for autonomous multi-agent evolution on open-ended problems. CORAL uses long-running LLM agents that explore, reflect, and collaborate through shared persistent memory and asynchronous execution, achieving 3–10× higher improvement rates over fixed evolutionary baselines across mathematical, algorithmic, and systems optimization tasks.
H. Zheng, Y. Ma, B. Araki, J. Chen, C. Wu
JAIR, 2026
Developed RL-RH-PP, combining deep reinforcement learning with search-based prioritized planning for lifelong MAPF. Demonstrated superior throughput and strong generalization across agent densities and planning horizons.
M. Kong, A. Qu, X. Guo, W. Ouyang, C. Jiang, H. Zheng, Y. Ma, D. Zhuang, Y. Tang, J. Li, H. Wang, C. Wu, J. Zhao
KDD, 2026 INFORMS Data Mining, 2025
Self-improving LLM framework that builds an experience library for formulating optimization programs across diverse problem domains.
Y. Tang, Z. Wang, A. Qu, Y. Yan, Z. Wu, D. Zhuang, J. Kai, K. Hou, X. Guo, H. Zheng, T. Luo, J. Zhao, Z. Zhao, W. Ma
EMNLP, 2024
Integrated spatial optimization with LLMs for open-domain urban itinerary planning, generating personalized routes from natural-language user requests.
H. Zheng, Z. Yan, C. Wu
IROS, 2024
Pioneered BK-PBS, integrating offline-trained behavior prediction with multi-agent pathfinding to coordinate connected and human-driven vehicles under realistic kinematic constraints on highways.
Z. Yan, H. Zheng, C. Wu
ICRA, 2024
Proposed OBS-KATS for signal-free intersection coordination. Proved soundness, completeness, and optimality of the crossing order search with significant throughput improvements.
Industry
Research Intern — Amazon AGI 2026
Bellevue, WA
Agentic post-training.
Research Intern — Symbotic May – Sep 2025
Wilmington, MA
Learning-guided multi robot path planning.
Education
2023 –
Massachusetts Institute of Technology
Ph.D. in Mechanical Engineering & Computational Science and Engineering
2021 – 2023
University of Michigan – Ann Arbor
Ph.D. in Mechanical Engineering (transferred to MIT)
2017 – 2021
University of Illinois at Urbana-Champaign
B.S. Mechanical Engineering with Highest Honors
2017 – 2021
Zhejiang University
B.E. in Mechanical Engineering