I am currently a second-year Ph.D. student (2024.06-2027.06 expected) at University of Science and Technology of China (USTC), jointly trained with Shanghai Artificial Intelligence Laboratory, co-supervised by Dr. Jiangmiao Pang, Prof. Feng Zhao, and Prof. Dahua Lin. Prior to my Ph.D., I spent two years in the graduate program at USTC under the supervision of Prof. Yongdong Zhang. I received my B.Eng. with honors from Huazhong University of Science and Technology (HUST) in 2022.
My research focuses on Robotics and Embodied AI, particularly Vision-Language-Action (VLA) models and World-Action Models (WAMs). I am also interested in VLM Brains and Embodied World Models. My experience spans model architecture design, large-scale training, data pipeline construction, benchmark evaluation, and real-world robotic system deployment in Embodied AI.
🔥 News
- 2026.07: 🎉 InternVLA-A1.5 was released, a new unified foundation model that integrates world modeling and improve compositional generalization in embodied manipulation.
- 2026.06: 💪 EBench is officially released! We invite VLA and WAM models to undergo comprehensive evaluation with EBench. Check out the project repository!
- 2026.05: 💪 We propose RoboInter1.5, which extends intermediate representations to World Modeling, unlocking a new dimension for structured physical world prediction.
- 2026.04: 🎉 Robo3R got accepted to RSS 2026, congratulations to Sizhe Yang!
- 2026.03: 💪 FutureVLA was released.
- 2026.02: 💪 Data, Benchmark, Model of RoboInter were open source, and Robo3R was released.
- 2026.02: 🎉 RoboInter and InstructVLA got accepted to ICLR 2026.
- 2026.01: 🎉 We get Second Place Award (2/62) of RoCo Challenge @ AAAI 2026: Robotic Collaborative Assembling for Human-Centered Manufacturing.
- 2025.11: 🎉 CronusVLA got accepted to AAAI 2026 as a oral presentation.
- 2025.10: 💪 Code of InstructVLA, CronusVLA and SimX-OR were open source.
- 2025.09: 💪 InternVLA-M1 was released.
- 2025.06: 💪 CronusVLA was released.
- 2025.03: 🎉 Our GenManip and RoboGround got accepted to CVPR 2025.
- 2024.06: 🌟 I joined Shanghai AI Lab.
- 2024.01: 🎉 Our system for the Efficient and Controllable Text-to-Image Generation in the 2nd International Algorithm Case Competition (IACC) of the Greater Bay Area was awarded Second Prize in the Grand Finals (2/599, prize ¥200,000), where I served as the first contributor, and video of my presentation was released.
📝 Publications
(* :equal contribution; ‡: project leader; †: corresponding author)

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
Hao Li*, Ziqin Wang*, Weijun Wang, Junhao Cai, Jia Zeng, Yilun Chen, Jiangmiao Pang, Si Liu,


RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation
Hao Li, Ziqin Wang, Zi-han Ding, Shuai Yang, Yilun Chen‡, Yang Tian, Xiaolin Hu, Tai Wang, Dahua Lin, Feng Zhao†, Si Liu†, Jiangmiao Pang†

CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
Hao Li, Shuai Yang, Yilun Chen†, Xinyi Chen, Xiaoda Yang, Yang Tian, Hanqing Wang, Tai Wang, Dahua Lin, Feng Zhao†, Jiangmiao Pang†

InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
Shuai Yang*, Hao Li*, Bing Wang, Yilun Chen‡, Yang Tian, Tai Wang, Hanqing Wang, Feng Zhao, Yiyi Liao, Jiangmiao Pang

InternVLA-M1: A Spatially Grounded Foundation Framework for Generalist Robot Policy

FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model.
Xiaoxu Xu, Hao Li*‡, Bing Wang, Jinhui Ye, Yilun Chen, Jia Zeng, Xinyi Chen, Linning Xu, Dahua Lin, Weixin Li, Jiangmiao Pang

Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute Editing
Hao Li, Mengqi Huang, Lei Zhang, Bo Hu, Yu Liu, Zhengdong Mao

Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
Sizhe Yang, Linning Xu†, Hao Li, Juncheng Mu, Jia Zeng, Dahua Lin, Jiangmiao Pang†

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
Haifeng Huang, Xinyi Chen, Yilun Chen, Hao Li, Xiaoshen Han, Zehan Wang, Tai Wang, Jiangmiao Pang, Zhou Zhao,

ER-SAN: Enhanced-Adaptive Relation Self-Attention Network for Image Captioning
Jingyu Li, Zhendong Mao†, Shancheng Fang, Hao Li
🎖 Honors and Awards
- Second Place Award (2/62) of RoCo Challenge @ AAAI 2026: Robotic Collaborative Assembling for Human-Centered Manufacturing
- Second Prize (2/599, ¥200,000) of the 2nd International Algorithm Case Competition of Greater Bay Area, Track of Efficient and Controllable Text-to-Image Generation
- National Scholarship (Top 2%), Ministry of Education of China, 2024
- The First Prize Scholarship, USTC, 2025, 2024, 2023, 2022
- Outstanding Graduate, HUST, 2022
- Outstanding Undergraduate in Terms of Academic Performance (Top 1%), HUST, 2020
📖 Experience
Educations
- Ph.D. in Control Science and Engineering of University of Science and Technology of China & Shanghai Artificial Intelligence Laboratory, 2024.06 - 2027.06 (expected)
- Graduate program in Information and Communication Engineering of University of Science and Technology of China, 2022.09 - 2024.06
- B.S. in School of Electronic Information and Communications of Huazhong University of Science and Technology, 2018.09 - 2022.06
- Electronic Information Engineering (Mathematical Improvement Experimental Class)
Internships
- Research intern in Shanghai AI Lab.
🧐 Community Services
Reviewer
- AAAI 2025–2026, CVPR 2024–2025, ICLR 2025, ICML 2025, etc.
💻 Links
- Personal Email: lihaohn@mail.ustc.edu.cn
- WeChat: lihaohn_Jeas
- RedNote(Xiaohongshu): 5236901971