I am currently a second-year Ph.D. student (2024.06-2027.06 expected) at University of Science and Technology of China (USTC), jointly trained with Shanghai Artificial Intelligence Laboratory, co-supervised by Dr. Jiangmiao Pang, Prof. Feng Zhao, and Prof. Dahua Lin. Prior to my Ph.D., I spent two years in the graduate program at USTC under the supervision of Prof. Yongdong Zhang. I received my B.Eng. with honors from Huazhong University of Science and Technology (HUST) in 2022.

My research focuses on Robotics and Embodied AI, particularly Vision-Language-Action (VLA) models and World-Action Models (WAMs). I am also interested in VLM Brains and Embodied World Models. My experience spans model architecture design, large-scale training, data pipeline construction, benchmark evaluation, and real-world robotic system deployment in Embodied AI.

🔥 News

  • 2026.07: 🎉 InternVLA-A1.5 was released, a new unified foundation model that integrates world modeling and improve compositional generalization in embodied manipulation.
  • 2026.06: 💪 EBench is officially released! We invite VLA and WAM models to undergo comprehensive evaluation with EBench. Check out the project repository!
  • 2026.05: 💪 We propose RoboInter1.5, which extends intermediate representations to World Modeling, unlocking a new dimension for structured physical world prediction.
  • 2026.04: 🎉 Robo3R got accepted to RSS 2026, congratulations to Sizhe Yang!
  • 2026.03: 💪 FutureVLA was released.
  • 2026.02: 💪 Data, Benchmark, Model of RoboInter were open source, and Robo3R was released.
  • 2026.02: 🎉 RoboInter and InstructVLA got accepted to ICLR 2026.
  • 2026.01: 🎉 We get Second Place Award (2/62) of RoCo Challenge @ AAAI 2026: Robotic Collaborative Assembling for Human-Centered Manufacturing.
  • 2025.11: 🎉 CronusVLA got accepted to AAAI 2026 as a oral presentation.
  • 2025.10: 💪 Code of InstructVLA, CronusVLA and SimX-OR were open source.
  • 2025.09: 💪 InternVLA-M1 was released.
  • 2025.06: 💪 CronusVLA was released.
  • 2025.03: 🎉 Our GenManip and RoboGround got accepted to CVPR 2025.
  • 2024.06: 🌟 I joined Shanghai AI Lab.
  • 2024.01: 🎉 Our system for the Efficient and Controllable Text-to-Image Generation in the 2nd International Algorithm Case Competition (IACC) of the Greater Bay Area was awarded Second Prize in the Grand Finals (2/599, prize ¥200,000), where I served as the first contributor, and video of my presentation was released.

📝 Publications

(* :equal contribution; ‡: project leader; †: corresponding author)

Technical Report 2026
sym

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

Hao Li*, Ziqin Wang*, Weijun Wang, Junhao Cai, Jia Zeng, Yilun Chen, Jiangmiao Pang, Si Liu,

[Project]  [Paper]  [Code]  [Data]

Technical Report 2026
sym

InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

Hao Li*(core contributor), Team of InternVLA-A1.5.

[Project]  [Paper]  [Code]  [Model]

Technical Report 2026 & ICLR 2026
sym

RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation

Hao Li, Ziqin Wang, Zi-han Ding, Shuai Yang, Yilun Chen, Yang Tian, Xiaolin Hu, Tai Wang, Dahua Lin, Feng Zhao, Si Liu, Jiangmiao Pang

[Project]  [Paper]  [Code]  [Data]

AAAI 2026 (Oral)
sym

CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling

Hao Li, Shuai Yang, Yilun Chen, Xinyi Chen, Xiaoda Yang, Yang Tian, Hanqing Wang, Tai Wang, Dahua Lin, Feng Zhao, Jiangmiao Pang

[Project]  [Paper]  [Code of Models]  [Code of Benchmark]

ICLR 2026
sym

InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation

Shuai Yang*, Hao Li*, Bing Wang, Yilun Chen, Yang Tian, Tai Wang, Hanqing Wang, Feng Zhao, Yiyi Liao, Jiangmiao Pang

[Project]  [Paper]  [Code]

Technical Report 2025
sym

InternVLA-M1: A Spatially Grounded Foundation Framework for Generalist Robot Policy

InternVLA-M1 Team

[Project]  [Paper]  [Code]

Preprint 2026
sym

FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model.

Xiaoxu Xu, Hao Li*, Bing Wang, Jinhui Ye, Yilun Chen, Jia Zeng, Xinyi Chen, Linning Xu, Dahua Lin, Weixin Li, Jiangmiao Pang

[Project]  [Paper]  [Code]

AAAI
sym

Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute Editing

Hao Li, Mengqi Huang, Lei Zhang, Bo Hu, Yu Liu, Zhengdong Mao

[Project]  [Paper]  [Code]

RSS 2026
sym

Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction

Sizhe Yang, Linning Xu, Hao Li, Juncheng Mu, Jia Zeng, Dahua Lin, Jiangmiao Pang

[Project]  [Paper]  [Code]

CVPR 2025
sym

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors

Haifeng Huang, Xinyi Chen, Yilun Chen, Hao Li, Xiaoshen Han, Zehan Wang, Tai Wang, Jiangmiao Pang, Zhou Zhao,

[Project]  [Paper]  [Code]

IJCAI Long Oral
sym

ER-SAN: Enhanced-Adaptive Relation Self-Attention Network for Image Captioning

Jingyu Li, Zhendong Mao, Shancheng Fang, Hao Li

[Project]  [Paper]  [Code]

🎖 Honors and Awards

  • Second Place Award (2/62) of RoCo Challenge @ AAAI 2026: Robotic Collaborative Assembling for Human-Centered Manufacturing
  • Second Prize (2/599, ¥200,000) of the 2nd International Algorithm Case Competition of Greater Bay Area, Track of Efficient and Controllable Text-to-Image Generation
  • National Scholarship (Top 2%), Ministry of Education of China, 2024
  • The First Prize Scholarship, USTC, 2025, 2024, 2023, 2022
  • Outstanding Graduate, HUST, 2022
  • Outstanding Undergraduate in Terms of Academic Performance (Top 1%), HUST, 2020

📖 Experience

Educations

Internships

  • Research intern in Shanghai AI Lab.

🧐 Community Services

Reviewer

  • AAAI 2025–2026, CVPR 2024–2025, ICLR 2025, ICML 2025, etc.

💻 Links

  • Personal Email: lihaohn@mail.ustc.edu.cn
  • WeChat: lihaohn_Jeas
  • RedNote(Xiaohongshu): 5236901971